For Codex · Production field test · Tested September 29, 2026 · AnyCap team
Can Codex generate or analyze video?
Not on its own. With the AnyCap CLI installed, Codex made the clip below, reviewed it, and animated a still from it: 3 tasks, 44 credits, about 3 minutes of recorded CLI time. Both clips are the original production outputs.
Paste this into Codex
Codex follows the install guide, adds the skill and CLI, and asks you to log in once. The manual setup is further down.
Set up AnyCap using anycap.ai/install.txt, then show me what you can do with video.Can Codex generate video?
Yes, through the AnyCap CLI. Codex checked Kling 3.0's production schema and submitted the job through the CLI. The clip at the top of this page is what came back.
- What we asked
- “A 5-second slow push-in on a product box on a desk.” We reused Kling 3.0 and the same settings as the earlier trial.
- The call, shortened
anycap --endpoint https://api.anycap.ai video generate --model kling-3.0 --mode text-to-video --prompt "a slow camera push-in on a blank product box on a desk" --param duration=5 --param aspect_ratio=16:9 --param resolution=1080p --param format=mp4 --param generate_audio=false --no-wait -o codex-demo.mp4- What came back
- A 1920×1080, 24 fps MP4 saved as codex-demo.mp4. The box grows larger in the frame during a smooth push-in and remains blank, matching the no-text instruction.
- What it cost
- 17 credits and 1 min 1 s, of which 1 min was spent in the wait command, including polling and download.
Can Codex analyze video?
Codex has no video input, so it can't watch a file by itself. One video-read call gave it a written review it could act on. About the clip above, it reported:
- Camera: a slow push-in from 0:00 to 0:05; the box does not move.
- Text: no visible text or logos; the analysis correctly describes the box as blank.
- Model-reported artifact: slight shifting of the top flap and seam from about 0:02 to 0:05. It reported no pronounced flicker or jitter.
- What we asked
- “What does the camera do, is any text legible, and what artifacts are there?”
- The call, shortened
anycap --endpoint https://api.anycap.ai actions video-read --file ./codex-demo.mp4 --instruction "Respond in English. Describe the camera movement, any legible text without guessing, and visible artifacts with timestamps."- What came back
- English notes returned as JSON. Extracted frames confirm the closer framing and blank packaging. The subtle temporal-artifact finding is the analysis model’s observation.
- What it cost
- 10 credits and 18 s.
Can Codex turn an image into video?
Yes. Codex pulled the first frame with ffmpeg, checked that Kling 3.0 accepts a start image, and asked for a slow orbit around the box.
- What we asked
- “Take the first frame of the clip and orbit slowly around the box.”
- The call, shortened
anycap --endpoint https://api.anycap.ai video generate --model kling-3.0 --mode image-to-video --prompt "slow clockwise orbit around the blank box" --param images=./product-frame.png --param duration=5 --param aspect_ratio=16:9 --param resolution=1080p --param format=mp4 --param generate_audio=false --no-wait -o animated-frame.mp4- What came back
- A 5.04-second animated-frame.mp4, preserved without trimming. The first frame matches the reference, and the right-hand face becomes visible later in the clip. These frame checks do not measure the requested 20-degree orbit precisely.
- What it cost
- 17 credits and 1 min 48 s.
How much does it cost?
44 credits for all 3 tasks. Times sum the recorded CLI commands, including polling and downloads. Setup, page editing, and pauses between commands are excluded.
| Task | Model | Time | Credits |
|---|---|---|---|
| Text to video | Kling 3.0 | 1 min 1 s | 17 |
| Video analysis | video-read | 18 s | 10 |
| Image to video | Kling 3.0 | 1 min 48 s | 17 |
| Total | about 3 min | 44 |
Generation cost 17 credits per 5-second 1080p kling-3.0 clip in this test; other models and settings are priced differently; every successful video-read request costs a fixed 10 credits. See pricing →
How do you set it up?
Three commands, run once, in the environment where Codex runs shell commands.
1. Add the skill
npx -y skills add anycap-ai/anycap -a codex -yTeaches Codex when to call AnyCap and how.
2. Install the CLI
curl -fsSL https://anycap.ai/install.sh | shOne binary with no runtime dependencies.
3. Log in
anycap login && anycap statusRun status from the same working directory and environment that will submit the task.
What breaks inside Codex?
One setup issue appeared in this production run. These checks also help prevent duplicate jobs and unexpected spend.
A repository .env points at dev
Our repository loaded dev credentials even when we selected the production URL. We moved the test to a clean working directory and verified that status reported prod before generating.
anycap --endpoint https://api.anycap.ai statusThe executing environment needs its own access
This run used Codex desktop on macOS with working network and credential access. A different sandbox or cloud environment needs those checked separately; a successful login elsewhere is not proof.
anycap statusThe session ends before the video does
The wait command took 1 min for the first clip. If a session ends early, the job keeps running on AnyCap; resume it by ID instead of submitting a duplicate:
anycap video tasks wait <task_id> -o codex-demo.mp4Codex picks the model unless you do
Name the model and settings when cost or style matters. We reused Kling 3.0 with 5-second, 1080p, silent output and recorded each returned charge.
When shouldn't you use it?
When the job doesn't need a video model.
Cuts, trims, captions, overlays
Codex can use local tools such as ffmpeg for deterministic edits. This run used ffmpeg to extract the reference frame and inspect the outputs.
Text that has to be exact
Generated video invents lettering. Render titles and UI text in code and composite them.
No network access
AnyCap is a hosted API. An offline sandbox or cloud task can't reach it.
Large batches without a budget
One 5-second 1080p clip cost 17 credits here. Pick the model before asking for dozens of variants.
Which other models can Codex use?
This run used Kling 3.0. To request a different supported model, name it in the prompt and inspect its current schema:
- Seedream 5 · Image model
A polished first-pass image from a prompt, when Codex needs one good result.
- Nano Banana 2 · Image model
Faster iteration when Codex needs more drafts or variants.
- Seedance 2.5 · Video model
Accepts text, image, first/last-frame, and multimodal reference inputs in one model.
FAQ
Can Codex analyze video?
Not natively. Codex takes images (codex -i screenshot.png) but not video files. With the AnyCap CLI installed, it can send a local file to anycap actions video-read and get a written analysis back. In our test that was one command and 10 credits for a 5-second clip.
Can Codex generate video?
Not with its own tools. With the AnyCap skill and CLI it lists the video models, reads the chosen model's schema, and runs anycap video generate. Our 5-second 1080p kling-3.0 clip cost 17 credits and the wait command took 1 min, including polling and download.
Does Codex have vision?
For images, yes: attach a screenshot or diagram with -i, or paste it in the app, and the model reads it. For video there is no input path, which is the gap the AnyCap video-read action fills.
Can Codex watch my screen recordings?
Yes, through the same video-read command with --file. The CLI uploads local files up to 100 MB. For longer recordings, have Codex cut the relevant part with ffmpeg first, or pass a public link with --url.
How much did the test cost?
44 credits: 17 for the text-to-video clip, 10 for the analysis, 17 for the image-to-video clip. Generation cost depends on the model, duration, and resolution; video-read is a flat 10 credits per successful request.
Which video model does Codex use?
Whichever it picks from anycap video models, unless you name one. We reused kling-3.0 in this production run. Seedance 2.5 is the broader option when you need multimodal references.
Does this work in Codex cloud or the IDE extension?
This production run used Codex desktop on macOS. The requirements are the same anywhere: the environment must be able to run the anycap binary, reach api.anycap.ai, and hold a credential. Cloud tasks need network access enabled for their environment.
Can Codex generate or edit images?
Codex takes images as input but needs a tool to create or edit them. anycap image generate covers both; we did not re-test images for this page.
Tested September 29, 2026 in production with Codex desktop (runtime 0.158.0-alpha.2.1) and AnyCap CLI 0.6.3
Create your first image in Codex
A generated image saved in your project, with the actual task charge to review.
Install the CLI and sign in once. If you are already configured, continue with the access and cost checks.
curl -fsSL https://anycap.ai/install.sh | sh
anycap loginAdd the skill, then reopen your agent session if needed.
npx -y skills add anycap-ai/anycap -a codex -yRun anycap status, inspect the current model schema, and check pricing and your balance. Generation is billed; model access and rates can change.
Paste this prompt into your configured agent. Inspect the delivered file and actual charge. If a video is pending, resume its task instead of submitting a duplicate.
Use AnyCap to create a small olive-green desk lamp on an off-white background, soft studio light, no text. Check my login and the current image models and schema first. Use Seedream 5 text-to-image and save the result as ./desk-lamp.png. Before generation, explain the applicable pricing and ask me to approve the billed task. Report the saved file and actual charge after it succeeds.Read the related workflow. For setup help, see installation and login.