Last updated: July 30, 2026
If you work anywhere near tech, you probably noticed something odd about OpenAI's biggest release of the summer: it didn't show up when everyone expected it to.
GPT-5.6 was supposed to arrive in late June. Instead, OpenAI held it back for nearly two weeks at the request of the U.S. government, which reviewed the model under a Trump administration executive order that asks AI companies to voluntarily submit their most powerful systems for a 30-day safety check before public release. Only a small group of "trusted partners" got early access while that review played out.
Then, on July 9, 2026, the wait ended. OpenAI didn't just ship an update — it restructured its entire flagship lineup into three distinct tiers (Sol, Terra, and Luna) and, on the same day, launched a brand-new agent called ChatGPT Work that's meant to do actual finished work, not just answer questions.
This piece breaks down what's actually confirmed about GPT-5.6 and ChatGPT Work, what the real benchmark numbers say against Claude and Gemini, what it costs, and how people are actually reacting three weeks in — not the marketing version.
1. Why GPT-5.6 Got Delayed in the First Place
Here's the part most coverage buries: GPT-5.6's release wasn't a normal product launch. OpenAI first previewed the Sol, Terra, and Luna lineup in late June, but at the U.S. government's request, it initially limited access to a small group of partners whose involvement had been coordinated with federal officials, citing an ongoing AI-cybersecurity review process. OpenAI was fairly candid that it disliked this arrangement, stating plainly that this kind of government access gatekeeping shouldn't become permanent policy.
The company's own account frames the review as tied to a broader early-June AI cybersecurity order, which asks AI developers to voluntarily present frontier models to the Department of Commerce's Center for AI Standards and Innovation roughly a month before public release. After OpenAI sent technical staff to Washington to work through the review, the government cleared a full public launch for July 9 — about two weeks after the original preview, rather than the full 30 days.
Worth noting: this is not unique to OpenAI. Anthropic's Claude Fable 5 and Claude Mythos 5 went through a similar suspension-and-restoration cycle around the same export-control review process in June, which is part of why several of the comparison sources below reference Fable 5 as "suspended" during parts of July.
2. Meet the Family: Sol, Terra, and Luna
The biggest structural change in GPT-5.6 isn't a smarter model — it's that OpenAI stopped shipping one flagship and started shipping three, differentiated by capability and price rather than by a single "best" release:
- Sol — the flagship. OpenAI positions it as its strongest coding and agentic model to date, and has directly compared it to Anthropic's Claude Fable 5 in its own marketing.
- Terra — the mid-tier, "everyday" model. OpenAI says it performs close to the previous flagship, GPT-5.5, at roughly half the cost.
- Luna — the budget tier, built for high-volume and lower-latency workloads.
All three share the same underlying GPT-5.6 generation and a roughly 1.05-million-token context window (with a 128K output cap) — they differ in reasoning depth, speed, and price, not in knowledge cutoff or context size. Sam Altman has publicly claimed Sol is meaningfully more token-efficient on coding tasks than prior models, though that specific efficiency figure comes from OpenAI itself rather than an independent lab.
3. Pricing Breakdown (Including the July 30 Price Cut)
API Pricing
| Model | Input (per 1M tokens) | Output (per 1M tokens) | Context Window |
|---|---|---|---|
| GPT-5.6 Sol | $5.00 | $30.00 | ~1.05M tokens |
| GPT-5.6 Terra | $2.50 → cut ~20% (as of July 30) | $15.00 → cut ~20% | ~1.05M tokens |
| GPT-5.6 Luna | $1.00 → cut ~80% (as of July 30) | $6.00 → cut ~80% | ~1.05M tokens |
Just today, OpenAI announced a price cut on the two lower tiers: Luna dropped by roughly 80% and Terra by roughly 20%, per OpenAI's own release notes. That's a significant move — it suggests OpenAI is leaning harder into the budget/volume end of the market rather than just competing on raw intelligence.
Consumer Subscriptions
| Plan | Price | What You Get |
|---|---|---|
| ChatGPT Free / Go | — | Terra by default |
| ChatGPT Plus | $20/mo | Terra as default, model selection expanding |
| ChatGPT Pro | $200/mo | Full access including Sol |
Note: exact model access by subscription tier has shifted slightly since launch as OpenAI rolls features out in stages, so it's worth checking OpenAI's own plan page before assuming what your subscription includes.
4. How GPT-5.6 Actually Stacks Up Against Claude and Gemini
This is where a lot of the online coverage gets sloppy, so let's separate OpenAI's own launch numbers from independent third-party testing.
On OpenAI's launch charts (Terminal-Bench 2.1, a coding-agent benchmark):
- Sol Ultra: 91.9%
- Sol (standard): 88.8%
- Claude Mythos 5: 88.0% (close second)
- Claude Opus 4.8: 78.9%
Independent testing site Artificial Analysis, which runs the same evaluation across vendors rather than relying on each company's own charts, found a tighter race. On its Intelligence Index, GPT-5.6 Sol came in a close second to Claude Fable 5 — but at roughly a third of the cost, giving Sol a strong performance-per-dollar edge. Sol also led Artificial Analysis's Coding Agent Index specifically within OpenAI's own Codex harness, though that's a narrower claim than "best coding model" overall.
For real-repository software engineering work (as opposed to terminal/agent benchmarks), the picture flips: on SWE-bench Pro, Claude models — including Anthropic's newer Claude Opus 5, released later in July — have maintained a clear lead over Sol, with one comparison putting Opus 5 roughly 14–15 percentage points ahead. In short: Sol tends to win at fast, terminal-driven agent tasks; Claude tends to win at deep, multi-file repository engineering. Neither model dominates across the board, and the ranking has already shifted at least twice in the three weeks since GPT-5.6 launched, as Anthropic and OpenAI have traded blows with new releases.
For multimodal work — image, video, and deep Google Workspace integration — Gemini remains a common recommendation, particularly for teams already living inside Google's ecosystem, though we didn't find a recent independent benchmark head-to-head specific enough to cite a hard number here.
5. ChatGPT Work: What It Does and Who Gets It
The other major piece of the July 9 launch was ChatGPT Work, a new agent mode that OpenAI positions as fundamentally different from a normal chat session.
Instead of answering one question at a time, ChatGPT Work is designed to take a goal, gather context from your connected apps and files, break the job into steps, and stay on a task for minutes or hours before handing back a finished deliverable — a spreadsheet, slide deck, document, report, or even a simple web app — rather than a chat reply. It's built on GPT-5.6 and connects to external tools (Slack, Microsoft Teams, Gmail, Google Drive, Salesforce, and SharePoint among them) through a plugin system invoked with an "@" command. Sensitive actions require your approval before Work executes them.
The launch also came bundled with two related changes:
- A rebuilt ChatGPT desktop app for Mac and Windows that merges standard Chat, Work, and Codex into a single application — with the older interface renamed "ChatGPT Classic" for people who aren't ready to switch.
- ChatGPT Sites, a lightweight tool for building simple web pages and apps directly from a prompt.
Who Gets Access
Rollout has been staged by plan:
- Pro, Enterprise, and Edu users got Work first, on both web and mobile.
- Plus and Business plans followed within days.
- Free and Go users get Work inside the new desktop app only, not on web or mobile.
- Usage is metered rather than flat-rate, similar to how Codex billing already works — so heavy users should watch consumption.
6. What Real Users Are Saying
OpenAI held a Reddit AMA on July 10 — the day after launch — where its Codex team shared that Codex itself had crossed 5 million weekly users, double the figure from three months earlier, with roughly 150 feature updates shipped in that window. That's a genuinely strong adoption signal for the coding side of the business.
Reaction to the broader release has been more mixed than that headline number suggests. Independent developer and frequent OpenAI commentator Simon Willison noted that Sol at "medium" reasoning could become his default for coding work, while also describing the expanded lineup of models and reasoning levels as confusing to navigate. Wharton professor Ethan Mollick raised a fair question that a lot of users have echoed: it's not immediately obvious how ChatGPT Work is meant to be different from Codex, since both are agent-style tools built on the same underlying model.
Separately, the redesigned desktop app itself drew sharper criticism. Some longtime users reported that the update disrupted features they relied on daily, including custom GPTs and saved projects, and reactions to the new interface on Reddit were blunt and unfavorable in places. OpenAI has said it's actively patching issues with browser control, stuck threads, and connection reliability in the new app, and that ChatGPT Classic will keep running in parallel while those fixes land.
The honest takeaway: the underlying model (Sol especially) is getting real praise from technical users. The new app and the Work/Codex overlap are where the friction is.
7. Comparison Table: Sol vs. Terra vs. Luna vs. Competitors
| Model | Best For | Terminal-Bench 2.1 (vendor) | Input / Output ($ per 1M) | Recommended For |
|---|---|---|---|---|
| GPT-5.6 Sol | Terminal/agent coding, speed | 88.8% (91.9% Ultra) | $5 / $30 | Developers needing fast agentic coding |
| GPT-5.6 Terra | Everyday tasks, balance | ~82–84% | $2.50→~$2 / $15→~$12 | General users, cost-conscious teams |
| GPT-5.6 Luna | High-volume, low-cost | ~82–84% | $1→~$0.20 / $6→~$1.20 | Batch processing, simple/high-frequency tasks |
| Claude Opus 5 / Fable 5 | Deep repo engineering, long-form reasoning | 88.0% (Mythos 5) | Higher, but competitive on SWE-bench Pro | Complex, multi-file software work |
| Gemini 3.1 Pro | Multimodal, Google Workspace | Not directly comparable | Mid-range | Teams inside Google's ecosystem |
Pricing shown reflects the July 30, 2026 cuts to Terra and Luna where confirmed; verify current rates on OpenAI's pricing page before budgeting, since both companies have adjusted prices multiple times this summer.
8. Who Should Actually Use What
- Solo developers and small teams doing agent-driven coding: Sol is a reasonable default, especially if your workflow is terminal- and CLI-heavy. If your work is deep, multi-file repository engineering, it's worth testing Claude Opus 5 or Fable 5 side by side before committing.
- General knowledge workers on a budget: Terra plus the July price cut makes it a genuinely cheap way to get most of the capability at a fraction of Sol's cost.
- High-volume, simple tasks (support ticket triage, log summarization, bulk classification): Luna, especially after its ~80% price cut, is hard to beat on cost per task.
- Teams wanting an agent that ships finished documents, not just chat replies: ChatGPT Work is worth a pilot — but go in expecting some overlap and confusion with Codex, and expect the desktop app to still have rough edges.
- Heavy Google Workspace users or multimodal-first teams: Gemini remains a strong default, particularly for anything video- or Drive-integration-heavy.
9. Frequently Asked Questions
- Q1: Is GPT-5.6 Sol actually better than Claude?
- It depends what "better" means. Sol tends to lead on terminal-driven agent benchmarks like Terminal-Bench 2.1. Claude's models (Fable 5, and later Opus 5) have kept a clear lead on real-repository engineering benchmarks like SWE-bench Pro. Neither wins everywhere.
- Q2: Why was GPT-5.6 delayed?
- It wasn't delayed in the traditional sense — OpenAI previewed it in late June but limited access to a small partner group while the model went through a U.S. government cybersecurity review tied to a Trump administration executive order, before opening it up publicly on July 9.
- Q3: What's the actual difference between ChatGPT Work and Codex?
- Both run on GPT-5.6 and both act as agents rather than simple chatbots. Work is aimed at general knowledge-work deliverables (docs, slides, sheets, reports, simple web apps) across connected business apps. Codex is specifically the coding-focused agent. Even industry commentators have noted the line between them isn't perfectly clean yet.
- Q4: Did prices change recently?
- Yes — as of July 30, 2026, OpenAI cut Luna's price by roughly 80% and Terra's by roughly 20%, according to OpenAI's own release notes. Sol's price was not part of that cut as of this writing.
- Q5: Can I use ChatGPT Work on mobile?
- Yes, but only on paid plans (Pro, Enterprise, Edu at launch, with Plus and Business following). Free and Go users only get Work through the new desktop app.
- Q6: Is the new merged desktop app worth switching to right away?
- Not necessarily on day one. Several users reported disruptions to existing workflows (custom GPTs, saved projects) after the update. ChatGPT Classic is still available in parallel while OpenAI works through fixes, so cautious users may want to wait.
- Q7: What happened to Claude Fable 5 around the same time?
- Claude Fable 5 and Claude Mythos 5 went through their own brief suspension in mid-to-late June tied to a separate U.S. export-control review, with access restored around July 1 — a parallel storyline that's worth knowing since several GPT-5.6 comparisons reference Fable 5 as "suspended" at points during this period.
- Q8: Which tier should a small business actually pick?
- For most day-to-day writing, research, and light coding, Terra is the practical default given the recent price cut. Reach for Sol only when a task specifically benefits from stronger agentic coding performance, and consider Luna for high-volume, low-complexity workloads like ticket triage.
10. Conclusion
The headline version of this story — "OpenAI's newest model crushes the competition" — doesn't hold up cleanly once you look past the vendor's own launch charts. What's actually true is more interesting: OpenAI restructured its pricing and product strategy around choice and cost-efficiency rather than chasing a single "smartest model" crown, cut prices on its budget tiers within three weeks of launch, and shipped a genuinely new kind of agent in ChatGPT Work — even if that agent's relationship to Codex is still a little muddy.
If you're deciding what to use this week: pick based on your actual workload rather than a benchmark chart. Terminal-heavy agentic coding leans toward Sol. Deep, multi-file engineering still leans toward Claude. High-volume simple tasks lean toward Luna's new pricing. And if you try ChatGPT Work, go in expecting a capable but still-maturing tool, not a finished product.Related Reading on Mustrend