If your production stack runs on GPT-5.5, Terra is the direct upgrade. Same capability tier, better benchmark scores, half the cost, and 16% fewer output tokens per task according to early adopters including Notion's co-founder.
That's not a small efficiency gain — at production token volumes, a 2× price reduction and 16% output reduction compounds into meaningful savings. Terra isn't positioned as the "budget Sol"; it's positioned as the replacement for GPT-5.5 that costs less than what you're already paying.
Terra at a Glance
| Specification | Value |
|---|---|
| Input price | $2.50 per 1M tokens |
| Output price | $15.00 per 1M tokens |
| Positioning | Balanced — GPT-5.5-competitive at lower cost |
| Availability | API, ChatGPT (Plus, Pro, Business, Enterprise), Codex |
| Default for | Free and Go users in ChatGPT Work and Codex |
Terra sits at the center of the GPT-5.6 family. It costs exactly half what Sol costs, and according to Simon Last (Co-Founder at Notion), "Many agents running GPT-5.5 perform just as well on Terra for half the cost and 16% fewer tokens."

Terra's Benchmark Results
All scores sourced from OpenAI's official GPT-5.6 launch data (July 9, 2026):
| Benchmark | Terra | GPT-5.5 | Sol | Luna |
|---|---|---|---|---|
| Agents' Last Exam | 50.4% | 46.9% | 52.7% | 50.3% |
| Coding Agent Index | 77.4 | 76.4 | 80 | 74.6 |
| BrowseComp | 87.5% | 84.4% | 90.4% | 83.3% |
| Terminal-Bench 2.1 | 87.4% | 85.6% | 88.8% | 84.7% |
| GPQA Diamond | 92.9% | 93.6% | 94.6% | 92.3% |
| DeepSWE v1.1 | 69.6% | 67.0% | 72.7% | 67.2% |
| OSWorld 2.0 | 50.2% | 47.5% | 62.6% | 45.6% |
| HealthBench Professional | 57.7% | 49.5% | 60.5% | 55.7% |
| BrowseComp | 87.5% | 84.4% | 90.4% | 83.3% |
Terra outperforms GPT-5.5 across every benchmark listed. On Agents' Last Exam — the most demanding real-world agentic evaluation — Terra scores 50.4% against GPT-5.5's 46.9%. On the Coding Agent Index, Terra (77.4) edges Claude Fable 5 (77.2) by a narrow margin.
Why Terra Is the Default Production Choice
It exceeds GPT-5.5 across the board
Terra was explicitly positioned as the "GPT-5.5-compatible" tier — the drop-in upgrade for teams that built on GPT-5.5 and want better results without the cost increase. In practice, Terra does not just match GPT-5.5; it consistently outperforms it.
For any workflow built on GPT-5.5, Terra is the immediate upgrade path: same cost bracket, better performance.
It handles the majority of real-world production tasks
The benchmark categories where Sol's lead over Terra is largest — OSWorld 2.0 (computer use: 62.6% vs 50.2%) and long-context recall — represent a subset of production use cases. For the core workloads that most agents run — coding, research, content generation, multi-step analysis — Terra's gap from Sol is narrow enough that the 2× cost difference is hard to justify.
It's cost-sustainable for continuous operation
Production agent workflows often run continuously: monitoring data feeds, processing incoming requests, generating regular reports. At Sol's pricing, high-volume continuous operation becomes expensive quickly. Terra's pricing makes it feasible to run sophisticated agent workflows at production scale without the cost becoming the primary constraint.
Where Terra Wins Over Luna
Terra's meaningful advantage over Luna concentrates in two areas:
Long-context reliability. On OpenAI's MRCR v2 benchmark at 512K–1M context (8-needle retrieval), Terra scores 72.5% against Luna's 41.3%. For agents that process long documents, maintain extended session histories, or need to reliably retrieve information from large contexts, Terra is substantially more reliable.
Complex multi-step reasoning. Terra's lead over Luna on coding benchmarks (77.4 vs 74.6 on Coding Agent Index) and OSWorld (50.2% vs 45.6%) suggests it handles longer reasoning chains with greater reliability. For agents whose execution path involves multiple dependent steps, Terra's consistency reduces compounding error risk.
Where Sol Is Worth the Premium Over Terra
Sol's most pronounced benchmark advantages over Terra appear in:
- OSWorld 2.0: 62.6% vs 50.2% — a 12.4-point gap that matters for agents interacting with real GUI environments
- BrowseComp: 90.4% vs 87.5% — relevant for complex agentic web research tasks
- ARC-AGI-3: 7.78% vs 0.80% — for tasks requiring novel abstract reasoning
If your agent's primary task involves extended computer use, complex GUI interaction, or genuinely novel reasoning problems, Sol's premium is justified. For most other production workloads, Terra's gap from Sol is narrow enough that the 2× cost difference is the deciding factor.
Terra in Multi-Tier Architectures
Terra often plays the "main execution" role in mixed-tier agent stacks:
- Luna handles initial triage, routing, and parallel coordination — fast and cheap
- Terra handles the main reasoning, coding, and content generation path
- Sol handles the steps where errors are expensive or the task genuinely exceeds Terra's capability
This architecture captures the efficiency gains from tiering without sacrificing quality where it matters. Terra's reliability and cost profile make it the natural backbone of the stack.
Practical Migration from GPT-5.5
For teams currently running GPT-5.5 workloads, the migration path to Terra is straightforward:
- Run the same prompts through Terra. On most tasks, output quality will match or exceed GPT-5.5 with no prompt changes required.
- Measure output token counts. Early adopters report 10–16% fewer output tokens per task on Terra — meaning the effective cost reduction is slightly larger than the headline 2× pricing difference.
- Identify tasks that still benefit from Sol. After migrating the bulk of workloads to Terra, the remaining cases where Sol's higher capability matters become clearer and more identifiable.
The Capability Layer Terra Needs
Terra handles reasoning, synthesis, and coding. The tasks it cannot do natively — generating images, searching the live web, producing video, publishing output to a hosted URL — require a capability layer outside the model itself.
That is where AnyCap fits: a single CLI that adds image generation, video production, web search, web crawl, persistent storage, and page hosting to any agent workflow. The agent does its reasoning with Terra; AnyCap handles the execution steps that go beyond what a language model can do on its own.
For production content pipelines running on Terra, the combination covers the full workflow: Terra drafts and refines the content; AnyCap generates the visuals, pulls live data when needed, and publishes the finished output.
All benchmark data sourced from OpenAI's official GPT-5.6 announcement (July 9, 2026). Pricing is per 1 million tokens at standard API rates.