anycapanycap
Capabilities

Generate

Image GenerationCreate and edit images from prompts or references.Video GenerationCreate motion outputs from text and image inputs.Music GenerationProduce music tracks through one runtime.Audio GenerationGenerate speech, dialogue, sound effects, and complete audio scenes from text, audio, or image input.

Understand

Image UnderstandingRead screenshots, diagrams, and visual references.Video AnalysisInspect recordings and extract structured details.Audio UnderstandingTranscribe and analyze voice and audio files.

Retrieve

Web SearchSearch the web from the same agent workflow.Grounded Web SearchReturn synthesized answers with live citations.Web CrawlFetch pages and convert them into clean content.Social Media LookupRead Instagram and X profiles, posts, feeds, and X search results.

Store

DriveStore outputs, organize assets, and create public URLs.
Equip Agents
Claude CodeCursorCodexDeepSeek HarnessManus
Resources

Explore

GuidesDecision guides for building reliable agent workflows.Context EngineeringUnderstand how prompts, files, and workspace state shape agent behavior.Agent SkillsSee how reusable skills package workflows and capability usage for agents.

Evaluate

Compare AnyCapBrowse comparison pages for adjacent agent tooling, media APIs, and tradeoffs.GlossaryA shared vocabulary for agent capabilities, tools, and workflows.
Docs ↗Pricing
I'm Agent
I'm Agent
  1. Home
  2. Learn
  3. Product
  4. CLI

CLI · By AnyCap Team

One CLI for the capabilities your agent still needs.

The agent can plan the workflow. The missing layer is usually execution: one command surface for image generation, video generation, image understanding, and video analysis. AnyCap CLI gives that layer one install path, one auth flow, and one interface across Claude Code, Cursor, Codex, and similar agent products.

Install once

  1. 1. Install the CLI

    $
    curl -fsSL https://anycap.ai/install.sh | sh

    Prefer a package manager? npm install -g @anycap/cli works on macOS, Linux, and Windows.

  2. 2. Log in once

    $
    anycap login

    Opens a browser, or prints a device code on a headless machine. Every command below reuses the same session.

  3. 3. Check the connection

    $
    anycap status

    Prints JSON with the server status, so the agent can confirm it is ready before the first real task.

The first commands most agents need

  • Image Generation →

    Generate and edit visuals with Seedream 5, Nano Banana Pro, and more.

    $
    anycap image generate --model seedream-5 --prompt "ceramic mug on a walnut desk, soft morning light"
  • Video Generation →

    Generate walkthroughs, clips, and motion output with Veo 3.1.

    $
    anycap video generate --model veo-3.1 --prompt "slow dolly shot across a tidy desk at sunrise"
  • Image Understanding →

    Analyze screenshots, diagrams, OCR, and visual references through one runtime.

    $
    anycap actions image-read --file ./screenshot.png --instruction "List every visible error message"
  • Video Analysis →

    Inspect recordings, summarize scenes, and extract structured video intelligence.

    $
    anycap actions video-read --file ./recording.mp4 --instruction "Summarize the key events"

Why one CLI matters

Keep the command surface stable
Without a unified CLI, every new capability becomes a new SDK, dashboard, or shell script. AnyCap keeps the execution layer consistent.
Log in once
Authentication happens once and carries across image, video, and vision workflows instead of fragmenting across providers.
Move across agents without re-learning the runtime
The same commands can sit under Claude Code, Cursor, Codex, and similar agent environments without forcing a new mental model each time.

Available across agent products: Claude Code, Cursor, Codex, OpenCode, OpenClaw

Reference docs

Every command, flag and output field, per-platform install notes, and credentials for CI.

CLI reference ↗Install AnyCap ↗Configuration ↗

Understand the rest of the stack

  • Image Generation →

    Go deeper on text-to-image, image editing, supported models, and the real CLI workflow.

  • Video Generation →

    See how text-to-video and image-to-video fit into the same agent command surface.

  • What agents can't do →

    The gaps a coding agent hits first, each with the page that closes it.

  • How skills fit →

    How a skill file tells the agent when to call the CLI and what to do with the output.

  • MCP vs Skills →

    When to reach for an MCP server and when a skill plus a CLI is simpler.

  • Context Engineering →

    When the agent should call a capability instead of reasoning inside the prompt.

Get StartedView on GitHubFor Claude Code

Capabilities

  • Overview
  • Image Generation
  • Video Generation
  • Music Generation
  • Image Understanding
  • Video Analysis
  • Audio Understanding
  • Web Search
  • Grounded Web Search
  • Web Crawl
  • Social Media Lookup
  • Drive

Equip Agents

  • Overview
  • Start here
  • Claude Code
  • Cursor
  • Codex
  • Manus

Resources

  • Overview
  • Context Engineering
  • Agent Skills
  • What Agents Can't Do
  • Compare agents and tools

Product

  • Product overview
  • Models
  • Install AnyCap
  • Install the Agent Skill

Documentation

  • Docs overview
  • Install AnyCap
  • MCP setup
  • CLI reference

Company

  • About
  • Contact
  • Privacy
  • Terms
anycap
Join the AnyCap Discord