anycapanycap
Capabilities

Generate

Image GenerationCreate and edit images from prompts or references.Video GenerationCreate motion outputs from text and image inputs.Music GenerationProduce music tracks through one runtime.Audio GenerationGenerate speech, dialogue, sound effects, and complete audio scenes from text, audio, or image input.

Understand

Image UnderstandingRead screenshots, diagrams, and visual references.Video AnalysisInspect recordings and extract structured details.Audio UnderstandingTranscribe and analyze voice and audio files.

Retrieve

Web SearchSearch the web from the same agent workflow.Grounded Web SearchReturn synthesized answers with live citations.Web CrawlFetch pages and convert them into clean content.Social Media LookupRead Instagram and X profiles, posts, feeds, and X search results.

Store

DriveStore outputs, organize assets, and create public URLs.
Equip Agents
Claude CodeCursorCodexDeepSeek HarnessManus
Resources

Explore

GuidesDecision guides for building reliable agent workflows.Context EngineeringUnderstand how prompts, files, and workspace state shape agent behavior.Agent SkillsSee how reusable skills package workflows and capability usage for agents.

Evaluate

Compare AnyCapBrowse comparison pages for adjacent agent tooling, media APIs, and tradeoffs.GlossaryA shared vocabulary for agent capabilities, tools, and workflows.
Docs ↗Pricing
I'm Agent
I'm Agent
  1. Home
  2. Equip Agents
  3. Codex

For Codex · Production field test · Tested September 29, 2026 · AnyCap team

Can Codex generate or analyze video?

Not on its own. With the AnyCap CLI installed, Codex made the clip below, reviewed it, and animated a still from it: 3 tasks, 44 credits, about 3 minutes of recorded CLI time. Both clips are the original production outputs.

Generated by Codex with Kling 3.0: 5 s, 1920×1080, 17 credits. The box stays blank, as requested.

Paste this into Codex

Codex follows the install guide, adds the skill and CLI, and asks you to log in once. The manual setup is further down.

Set up AnyCap using anycap.ai/install.txt, then show me what you can do with video.
Get startedSee pricing

Can Codex generate video?

Yes, through the AnyCap CLI. Codex checked Kling 3.0's production schema and submitted the job through the CLI. The clip at the top of this page is what came back.

What we asked
“A 5-second slow push-in on a product box on a desk.” We reused Kling 3.0 and the same settings as the earlier trial.
The call, shortened
$
anycap --endpoint https://api.anycap.ai video generate --model kling-3.0 --mode text-to-video --prompt "a slow camera push-in on a blank product box on a desk" --param duration=5 --param aspect_ratio=16:9 --param resolution=1080p --param format=mp4 --param generate_audio=false --no-wait -o codex-demo.mp4
What came back
A 1920×1080, 24 fps MP4 saved as codex-demo.mp4. The box grows larger in the frame during a smooth push-in and remains blank, matching the no-text instruction.
What it cost
17 credits and 1 min 1 s, of which 1 min was spent in the wait command, including polling and download.

How video generation works →

Can Codex analyze video?

Codex has no video input, so it can't watch a file by itself. One video-read call gave it a written review it could act on. About the clip above, it reported:

  • Camera: a slow push-in from 0:00 to 0:05; the box does not move.
  • Text: no visible text or logos; the analysis correctly describes the box as blank.
  • Model-reported artifact: slight shifting of the top flap and seam from about 0:02 to 0:05. It reported no pronounced flicker or jitter.
What we asked
“What does the camera do, is any text legible, and what artifacts are there?”
The call, shortened
$
anycap --endpoint https://api.anycap.ai actions video-read --file ./codex-demo.mp4 --instruction "Respond in English. Describe the camera movement, any legible text without guessing, and visible artifacts with timestamps."
What came back
English notes returned as JSON. Extracted frames confirm the closer framing and blank packaging. The subtle temporal-artifact finding is the analysis model’s observation.
What it cost
10 credits and 18 s.

How video analysis works →

Can Codex turn an image into video?

Yes. Codex pulled the first frame with ffmpeg, checked that Kling 3.0 accepts a start image, and asked for a slow orbit around the box.

Started from the first frame of the production clip above. The box remains blank as the viewpoint moves around it.
What we asked
“Take the first frame of the clip and orbit slowly around the box.”
The call, shortened
$
anycap --endpoint https://api.anycap.ai video generate --model kling-3.0 --mode image-to-video --prompt "slow clockwise orbit around the blank box" --param images=./product-frame.png --param duration=5 --param aspect_ratio=16:9 --param resolution=1080p --param format=mp4 --param generate_audio=false --no-wait -o animated-frame.mp4
What came back
A 5.04-second animated-frame.mp4, preserved without trimming. The first frame matches the reference, and the right-hand face becomes visible later in the clip. These frame checks do not measure the requested 20-degree orbit precisely.
What it cost
17 credits and 1 min 48 s.

Image-to-video guide →

How much does it cost?

44 credits for all 3 tasks. Times sum the recorded CLI commands, including polling and downloads. Setup, page editing, and pauses between commands are excluded.

TaskModelTimeCredits
Text to videoKling 3.01 min 1 s17
Video analysisvideo-read18 s10
Image to videoKling 3.01 min 48 s17
Totalabout 3 min44

Generation cost 17 credits per 5-second 1080p kling-3.0 clip in this test; other models and settings are priced differently; every successful video-read request costs a fixed 10 credits. See pricing →

How do you set it up?

Three commands, run once, in the environment where Codex runs shell commands.

  1. 1. Add the skill

    $
    npx -y skills add anycap-ai/anycap -a codex -y

    Teaches Codex when to call AnyCap and how.

  2. 2. Install the CLI

    $
    curl -fsSL https://anycap.ai/install.sh | sh

    One binary with no runtime dependencies.

  3. 3. Log in

    $
    anycap login && anycap status

    Run status from the same working directory and environment that will submit the task.

Recorded separately: Codex installing AnyCap from a one-line prompt and generating its first image.

What breaks inside Codex?

One setup issue appeared in this production run. These checks also help prevent duplicate jobs and unexpected spend.

  1. 1.

    A repository .env points at dev

    Our repository loaded dev credentials even when we selected the production URL. We moved the test to a clean working directory and verified that status reported prod before generating.

    $
    anycap --endpoint https://api.anycap.ai status
  2. 2.

    The executing environment needs its own access

    This run used Codex desktop on macOS with working network and credential access. A different sandbox or cloud environment needs those checked separately; a successful login elsewhere is not proof.

    $
    anycap status
  3. 3.

    The session ends before the video does

    The wait command took 1 min for the first clip. If a session ends early, the job keeps running on AnyCap; resume it by ID instead of submitting a duplicate:

    $
    anycap video tasks wait <task_id> -o codex-demo.mp4
  4. 4.

    Codex picks the model unless you do

    Name the model and settings when cost or style matters. We reused Kling 3.0 with 5-second, 1080p, silent output and recorded each returned charge.

When shouldn't you use it?

When the job doesn't need a video model.

  1. —

    Cuts, trims, captions, overlays

    Codex can use local tools such as ffmpeg for deterministic edits. This run used ffmpeg to extract the reference frame and inspect the outputs.

  2. —

    Text that has to be exact

    Generated video invents lettering. Render titles and UI text in code and composite them.

  3. —

    No network access

    AnyCap is a hosted API. An offline sandbox or cloud task can't reach it.

  4. —

    Large batches without a budget

    One 5-second 1080p clip cost 17 credits here. Pick the model before asking for dozens of variants.

Which other models can Codex use?

This run used Kling 3.0. To request a different supported model, name it in the prompt and inspect its current schema:

  • Seedream 5 · Image model

    A polished first-pass image from a prompt, when Codex needs one good result.

  • Nano Banana 2 · Image model

    Faster iteration when Codex needs more drafts or variants.

  • Seedance 2.5 · Video model

    Accepts text, image, first/last-frame, and multimodal reference inputs in one model.

FAQ

Can Codex analyze video?

+

Not natively. Codex takes images (codex -i screenshot.png) but not video files. With the AnyCap CLI installed, it can send a local file to anycap actions video-read and get a written analysis back. In our test that was one command and 10 credits for a 5-second clip.

Can Codex generate video?

+

Not with its own tools. With the AnyCap skill and CLI it lists the video models, reads the chosen model's schema, and runs anycap video generate. Our 5-second 1080p kling-3.0 clip cost 17 credits and the wait command took 1 min, including polling and download.

Does Codex have vision?

+

For images, yes: attach a screenshot or diagram with -i, or paste it in the app, and the model reads it. For video there is no input path, which is the gap the AnyCap video-read action fills.

Can Codex watch my screen recordings?

+

Yes, through the same video-read command with --file. The CLI uploads local files up to 100 MB. For longer recordings, have Codex cut the relevant part with ffmpeg first, or pass a public link with --url.

How much did the test cost?

+

44 credits: 17 for the text-to-video clip, 10 for the analysis, 17 for the image-to-video clip. Generation cost depends on the model, duration, and resolution; video-read is a flat 10 credits per successful request.

Which video model does Codex use?

+

Whichever it picks from anycap video models, unless you name one. We reused kling-3.0 in this production run. Seedance 2.5 is the broader option when you need multimodal references.

Does this work in Codex cloud or the IDE extension?

+

This production run used Codex desktop on macOS. The requirements are the same anywhere: the environment must be able to run the anycap binary, reach api.anycap.ai, and hold a credential. Cloud tasks need network access enabled for their environment.

Can Codex generate or edit images?

+

Codex takes images as input but needs a tool to create or edit them. anycap image generate covers both; we did not re-test images for this page.

Official documentation

Set up AnyCap in Codex.

Use the official Docs for Agent Skill installation, MCP setup, and the video capability reference.

Agent Skill ↗MCP setup ↗Video capability ↗

Also available for

Claude CodeCursorManus

Tested September 29, 2026 in production with Codex desktop (runtime 0.158.0-alpha.2.1) and AnyCap CLI 0.6.3

Create your first image in Codex

A generated image saved in your project, with the actual task charge to review.

Install the CLI and sign in once. If you are already configured, continue with the access and cost checks.

$
curl -fsSL https://anycap.ai/install.sh | sh
anycap login

Add the skill, then reopen your agent session if needed.

$
npx -y skills add anycap-ai/anycap -a codex -y

Run anycap status, inspect the current model schema, and check pricing and your balance. Generation is billed; model access and rates can change.

Paste this prompt into your configured agent. Inspect the delivered file and actual charge. If a video is pending, resume its task instead of submitting a duplicate.

Use AnyCap to create a small olive-green desk lamp on an off-white background, soft studio light, no text. Check my login and the current image models and schema first. Use Seedream 5 text-to-image and save the result as ./desk-lamp.png. Before generation, explain the applicable pricing and ask me to approve the billed task. Report the saved file and actual charge after it succeeds.

Read the related workflow. For setup help, see installation and login.

Capabilities

  • Overview
  • Image Generation
  • Video Generation
  • Music Generation
  • Image Understanding
  • Video Analysis
  • Audio Understanding
  • Web Search
  • Grounded Web Search
  • Web Crawl
  • Social Media Lookup
  • Drive

Equip Agents

  • Overview
  • Start here
  • Claude Code
  • Cursor
  • Codex
  • Manus

Resources

  • Overview
  • Context Engineering
  • Agent Skills
  • What Agents Can't Do
  • Compare agents and tools

Product

  • Product overview
  • Models
  • Install AnyCap
  • Install the Agent Skill

Documentation

  • Docs overview
  • Install AnyCap
  • MCP setup
  • CLI reference

Company

  • About
  • Contact
  • Privacy
  • Terms
anycap
Join the AnyCap Discord