CLI · By AnyCap Team
One CLI for the capabilities your agent still needs.
The agent can plan the workflow. The missing layer is usually execution: one command surface for image generation, video generation, image understanding, and video analysis. AnyCap CLI gives that layer one install path, one auth flow, and one interface across Claude Code, Cursor, Codex, and similar agent products.
Install once
1. Install the CLI
curl -fsSL https://anycap.ai/install.sh | shPrefer a package manager? npm install -g @anycap/cli works on macOS, Linux, and Windows.
2. Log in once
anycap loginOpens a browser, or prints a device code on a headless machine. Every command below reuses the same session.
3. Check the connection
anycap statusPrints JSON with the server status, so the agent can confirm it is ready before the first real task.
The first commands most agents need
Image Generation →
Generate and edit visuals with Seedream 5, Nano Banana Pro, and more.
anycap image generate --model seedream-5 --prompt "ceramic mug on a walnut desk, soft morning light"Video Generation →
Generate walkthroughs, clips, and motion output with Veo 3.1.
anycap video generate --model veo-3.1 --prompt "slow dolly shot across a tidy desk at sunrise"Image Understanding →
Analyze screenshots, diagrams, OCR, and visual references through one runtime.
anycap actions image-read --file ./screenshot.png --instruction "List every visible error message"Video Analysis →
Inspect recordings, summarize scenes, and extract structured video intelligence.
anycap actions video-read --file ./recording.mp4 --instruction "Summarize the key events"
Why one CLI matters
- Keep the command surface stable
- Without a unified CLI, every new capability becomes a new SDK, dashboard, or shell script. AnyCap keeps the execution layer consistent.
- Log in once
- Authentication happens once and carries across image, video, and vision workflows instead of fragmenting across providers.
- Move across agents without re-learning the runtime
- The same commands can sit under Claude Code, Cursor, Codex, and similar agent environments without forcing a new mental model each time.
Available across agent products: Claude Code, Cursor, Codex, OpenCode, OpenClaw
Understand the rest of the stack
Image Generation →
Go deeper on text-to-image, image editing, supported models, and the real CLI workflow.
Video Generation →
See how text-to-video and image-to-video fit into the same agent command surface.
What agents can't do →
The gaps a coding agent hits first, each with the page that closes it.
How skills fit →
How a skill file tells the agent when to call the CLI and what to do with the output.
MCP vs Skills →
When to reach for an MCP server and when a skill plus a CLI is simpler.
Context Engineering →
When the agent should call a capability instead of reasoning inside the prompt.