Claude vs ChatGPT in August 2026
An even-handed comparison of Claude and ChatGPT as of August 2026 - current models, plan pricing, verified benchmark results, and the categories where each one clearly loses.
Claude and ChatGPT are close on raw intelligence and far apart on scope. As of August 2026, Anthropic's Opus 5 tops the independent Artificial Analysis Intelligence Index and Fable 5 tops arena.ai's human-preference board, while OpenAI's GPT-5.6 family wins most agentic-breadth benchmarks, generates images and video, and gives away unlimited text chat for free. Which one you should pay for depends almost entirely on whether you need generative media.
This page reports the losses as carefully as the wins. Every benchmark below names the model version and who ran the test, because in August 2026 almost every published table is run by one of the two vendors against the other's previous generation.
#Head to head, August 2026
- Anthropic lineup
- Fable 5, Opus 5, Sonnet 5, Haiku 4.5
- OpenAI lineup
- GPT-5.6 Sol, Terra, Luna (released 9 July 2026)
- Largest context
- Claude 1,000,000 tokens · GPT-5.6 1,050,000 tokens
- Cheapest paid plan
- Claude Pro $20/mo · ChatGPT Go $8/mo
- Image / video generation
- Claude: none · ChatGPT: GPT-Image-2, Sora 2
- Biggest single benchmark gap
- ARC-AGI-3 - Opus 5 30.16% vs GPT-5.6 Sol 7.78%
#What models does each company sell right now?
Anthropic ships four callable models; OpenAI ships three general-purpose ones plus a wide media and voice range. The tiers do not line up neatly, which is why vendor comparison tables so often talk past each other.
| Tier | Anthropic | OpenAI | Notes |
|---|---|---|---|
| Top-end | Fable 5 - 1M context, $10/$50 per MTok | GPT-5.6 Sol - 1.05M context | Sol's output price is disputed between OpenAI's own pages (see below) |
| Flagship default | Opus 5 - 1M, $5/$25 | GPT-5.6 Terra - 1.05M, $2/$12 | Opus 5 has the newer knowledge cutoff of the Claude line (May 2026) |
| Workhorse | Sonnet 5 - 1M, $2/$10 | GPT-5.6 Luna - 1.05M, $0.20/$1.20 | Luna is roughly a tenth of Sonnet 5's input price |
| Cheap and fast | Haiku 4.5 - 200k, $1/$5 | (Luna serves this role) | Haiku 4.5 is the only current Claude model without a 1M window |
| Media | None | GPT-Image-2, Sora 2, GPT-Realtime-2.1, GPT-Transcribe | Categorical gap, not a scoring gap |
OpenAI's two official price lists do not agree on GPT-5.6 Sol's output rate. The openai.com API pricing page shows one figure; the developer docs pricing page shows a different, context-tiered rate card. The input rates reconcile if the marketing page is quoting long-context pricing; the output rates do not reconcile at all. Until OpenAI resolves it, this page states no Sol output price. Price your own workload against the docs card, and expect a further uplift for regional processing endpoints.
One structural note that matters more than the tier names: all three GPT-5.6 models carry the same 1.05M-token window, so OpenAI has flattened context across its range. Anthropic has done nearly the same - the 1M window is the default on Fable 5, Opus 5 and Sonnet 5 with no beta header and no long-context surcharge - but Haiku 4.5 is still capped at 200k. The Claude model index has the full specification table.
#Which model wins on benchmarks?
Neither, cleanly. Claude wins novel-problem reasoning and human preference; OpenAI wins agentic breadth and computer use. Here is the honest split, with the source for every row.
| Benchmark | Claude | GPT-5.6 | Winner | Who ran it |
|---|---|---|---|---|
| ARC-AGI-3 | 30.16% (Opus 5, high effort) | 7.78% (Sol, max) | Claude, by roughly 4× | ARC Prize, independent |
| ARC-AGI-2 | 90.4% (Opus 5, max) | 92.5% (Sol, max) | OpenAI, by 2.1 points | ARC Prize, independent |
| ARC-AGI-1 | 97.5% (Opus 5) | 97.5% (Sol, xhigh) | Tied | ARC Prize, independent |
| Artificial Analysis Intelligence Index | 63 (Opus 5, max) - #1 overall | 61 (Sol, max) - #5 | Claude | Artificial Analysis, independent |
| arena.ai Elo (human preference) | 1507 (Fable 5) - #1 | No GPT model in the top 15 | Claude | arena.ai, independent |
| Agents' Last Exam | 40.5% (Fable 5) | 52.7% (Sol) | OpenAI, by 12 points | OpenAI, self-reported |
| OSWorld 2.0 (computer use) | 54.8% (Opus 4.8) | 62.6% (Sol) | OpenAI | OpenAI, self-reported |
| Terminal-Bench 2.1 | 83.8% (Claude Code + Fable 5) - top of the official board | 91.9% (Ultra multi-agent mode) | OpenAI holds the highest published number | Official board vs OpenAI self-report |
| SWE-bench Pro | 80.3% (Mythos 5, invitation-only) | 64.6% (Sol) | Claude, but on a model with no self-serve access - OpenAI's table publishes no Opus 5 or Fable 5 figure | OpenAI's own launch table |
| GDPval-AA v2 (Elo) | 1,759.6 (Fable 5) | 1,747.8 (Sol) | Claude, narrowly | OpenAI, self-reported |
| BrowseComp | 84.3% (Opus 4.8) | 92.2% (Ultra) | OpenAI | OpenAI, self-reported |
#Read those numbers with three caveats
Scaffolding is part of the score. Terminal-Bench measures an agent plus a model. "Claude Code + Fable 5" and "Codex + GPT-5.6 Sol" are different harnesses, so a two-point difference tells you as much about the tooling as the weights. A third-party roundup published on 2 August 2026 put Sol at xhigh effort on 89.5% and Opus 5 at max on 89.1% - a gap inside the error bars, but not a Claude lead.
Vendors benchmark against last generation. OpenAI's GPT-5.6 launch table compares Sol against Claude Opus 4.7-era models and Opus 4.8 on most rows, using Fable 5 only selectively. The 54.8% OSWorld figure above is Opus 4.8, not Opus 5, because Anthropic has never published a numeric Opus 5 OSWorld result - only a relative claim that it beats Fable 5 at a third of the cost.
Anthropic's own transparency is the weaker of the two. OpenAI and Google publish full numeric comparison tables as text. Anthropic published the Fable 5 benchmark table only as an image, and the Opus 5 announcement gives relative claims ("3× the next-best model on ARC-AGI-3", "within 0.5% of Fable 5 at half the cost") instead of absolute figures. That makes Anthropic's claims harder to check independently than its competitors', and it is a fair thing to hold against them. The point is expanded in our Claude review.
Benchmarks that used to settle these arguments - GPQA Diamond, AIME, MMMU and SWE-bench Verified - have effectively been retired at the frontier. None of them appear on OpenAI's GPT-5.6 tables or Google's current model cards. If you see a 2026 comparison leaning on them, it is recycling stale numbers.
#Where does ChatGPT clearly beat Claude?
Four places, and three of them are structural rather than a matter of a few benchmark points.
#1. Generative media - Claude has none at all
Claude models are text and image in, text out. There is no image generation, no video generation, no audio or music generation, and no video or audio input either. OpenAI ships GPT-Image-2 for images, Sora 2 for video (priced per second of output), GPT-Realtime-2.1 for speech-to-speech, and a transcription and translation line. If your work touches any of those, Claude is not a candidate and no amount of reasoning quality changes that. Our feature and plan matrix lists what Claude actually does instead.
#2. The free tier, and the gap above it
Since 6 August 2026 ChatGPT Free offers unlimited text chats with GPT-5.6 Luna, funded by advertising for US users, and ChatGPT Go at $8/month extends that with roughly ten times the file, upload and image allowances. Claude Free is a metered weekly pool with a five-hour rolling refresh, and Anthropic sells nothing between $0 and $20. That leaves an $8–$12 band that OpenAI and Google both occupy and Anthropic does not. What $0 actually gets you on Claude is itemised on the free plan page.
#3. Distribution through Microsoft
Microsoft 365 Copilot puts GPT models in front of every Office seat. Business pricing starts at about $18/user/month on annual billing (promotional through 30 September 2026, regular $21) per Microsoft's pricing page, rising to $32 for Business Premium bundles. Anthropic has no equivalent default-placement channel - it ships Microsoft 365 add-ins including Claude for Excel, but as an install, not as the assistant already in the ribbon for a billion seats.
#4. Agentic breadth and computer use
OpenAI wins OSWorld 2.0, BrowseComp, DeepSWE v1.1, Terminal-Bench 3.0 and Agents' Last Exam on its own published tables, and Codex is a mature competitor to Claude Code rather than a follower. Anthropic's computer use tooling only reached general availability on 19 August 2026, alongside a new browser-use tool. That is late relative to the benchmark record.
#Where does Claude clearly beat ChatGPT?
Also four places, and they are narrower but real.
#Novel-problem reasoning
ARC-AGI-3 is the single largest lead any lab holds over any other right now: 30.16% to 7.78%, verified independently by ARC Prize, and Opus 5 was only tested at high effort rather than max, so the figure is arguably a floor.
#Writing people prefer
Anthropic holds six of the top twelve places on arena.ai's text arena with Fable 5 at 1507. No GPT model appears in the top fifteen at all. If prose quality is the job, see Claude as a writing assistant.
#Repository-scale coding
OpenAI's own launch table concedes SWE-bench Pro at 80.3% (Mythos 5, invitation-only) against Sol's 64.6% - but that is the Glasswing model, not one you can sign up for. Claude Code remains the reference implementation for terminal-based agentic coding.
The fourth is enterprise pricing transparency. Anthropic publishes Enterprise at a flat $20/seat/month plus usage at API rates, with SCIM, audit logs, a Compliance API, custom data retention, HIPAA-ready configuration and IP allowlisting listed openly. OpenAI Enterprise is quote-only. That said, OpenAI advertises customer-managed encryption keys and data residency across ten regions, and Anthropic's public pricing page lists neither - so "better" here depends on which control you actually need.
#How do the plan prices compare?
At $20 and at $100/$200 the two are at parity. The difference is at the bottom of the range.
| Tier | Claude | ChatGPT | Who wins |
|---|---|---|---|
| Free | $0, metered weekly pool with 5-hour rolling refresh | $0, unlimited text with Luna; ads for US users | ChatGPT |
| Entry paid | Nothing offered | Go, $8/mo | ChatGPT |
| Flagship individual | Pro, $20/mo or $200/yr (~$17/mo) | Plus, $20/mo, no annual discount | Claude on annual billing |
| Heavy individual | Max 5× $100 / Max 20× $200 | Pro, $100 (5×) / $200 (20×) | Parity |
| Small team | Team, $25/seat/mo or $20 annual; 2–150 seats | Business, $25/seat/mo or $20 annual; 2-seat minimum | Parity |
| Enterprise | $20/seat/mo + usage at API rates, published | Custom quote, annual | Claude on transparency |
Sources: claude.com/pricing; ChatGPT Plus and Pro tiers per the OpenAI Help Center; ChatGPT Go per OpenAI's launch post; Business and Enterprise per OpenAI Business Pricing. Full Claude plan mechanics are on the pricing page.
Both vendors have moved to compute-weighted pools rather than message counts, and Anthropic publishes relative multipliers only - 5× and 20× Pro for Max, 1.25× for a standard Team seat. Any site quoting an absolute Claude message count is repeating a figure that is not in current documentation. Note too that Claude chat, Claude Code, Cowork, Design and Excel all draw from one shared allocation, so a heavy coding week eats your chat budget.
#What about API cost for developers?
Claude is the more expensive family, and it is not close at the bottom of the range. GPT-5.6 Luna at $0.20 in / $1.20 out per million tokens is roughly a tenth of Sonnet 5's input rate and an eighth of its output rate, and OpenAI's marketing explicitly claims Terra and Luna beat Fable 5 on some evals "at around one-sixteenth the cost". On Agents' Last Exam, Luna's 50.3% does beat Fable 5's 40.5%.
Two levers close some of that gap on the Claude side: the Batch API's flat 50% discount, and prompt caching, where cache reads cost 0.1× base input. Both are worked through with worked examples on the Claude API page. One caveat cuts the other way - Sonnet 5's tokenizer emits roughly 30% more tokens for the same text than Sonnet 4.6 did, so the sticker-price saving from its permanent $2/$10 rate is smaller in practice than it looks.
#Claude or ChatGPT: which should you actually use?
By job, not by brand. This table is the short version of everything above.
| If your job is… | Pick | Because |
|---|---|---|
| Anything involving images, video, audio or music | ChatGPT | Claude cannot generate any of them, and cannot accept video or audio input |
| Repository-scale refactoring and long agentic coding runs | Claude | Top of the official Terminal-Bench 2.1 board at 83.8%, plus Claude Code's tooling depth |
| Desktop and browser automation | ChatGPT | OSWorld 2.0 62.6% vs 54.8%, plus BrowseComp; Anthropic's computer use only reached GA in August 2026 |
| Long-form drafting and editing where tone matters | Claude | Six of the top twelve arena.ai slots; no GPT model in the top fifteen |
| Novel puzzle-like reasoning with no template to copy | Claude | ARC-AGI-3, 30.16% vs 7.78%, independently verified |
| Spending nothing at all | ChatGPT | Unlimited text on Free; Claude Free is metered |
| Spending under $20/month | ChatGPT | Go at $8; Anthropic has nothing between $0 and $20 |
| Working inside Word, Excel, Outlook and Teams all day | ChatGPT via Copilot | Microsoft 365 Copilot from about $18/user/month puts GPT models in the ribbon |
| Regulated enterprise buying against a published control list | Claude | Flat $20/seat + usage with compliance features listed openly, not quoted |
| Highest volume at lowest unit cost | ChatGPT | Luna at $0.20/$1.20; Claude's cheapest is Haiku 4.5 at $1/$5 |
Plenty of people run both, and at $20 each that is a defensible answer. If you want a wider field than these two - including the cheap open-weight models that are now within a few index points of the frontier - see Claude alternatives. The Google side of the argument, where the gaps are sharper still, is on Claude vs Gemini. And if you have already decided on Claude, the model comparison guide works out which of the four to actually call.
#Frequently asked questions
Is Claude better than ChatGPT in 2026?
Neither is better overall. Claude leads the independent Artificial Analysis Intelligence Index and arena.ai Elo, and wins ARC-AGI-3 by roughly four times. ChatGPT wins ARC-AGI-2, OSWorld 2.0, Agents' Last Exam and the top Terminal-Bench score, and generates images and video, which Claude cannot do at all.
Can Claude generate images like ChatGPT?
No. As of August 2026 Claude produces text and code only. It cannot generate images, video, audio or music, and it cannot accept video or audio as input, though it does read images and documents. OpenAI ships GPT-Image-2 for images and Sora 2 for video.
Which has the better free plan, Claude or ChatGPT?
ChatGPT, clearly. Since 6 August 2026 ChatGPT Free offers unlimited text chats with GPT-5.6 Luna, funded by advertising for US users. Claude Free is a metered weekly pool with a five-hour rolling refresh, and Anthropic sells no plan between $0 and $20.
Is Claude more expensive than ChatGPT for API use?
Yes. Anthropic's cheapest current model, Haiku 4.5, costs $1 in and $5 out per million tokens. GPT-5.6 Luna costs $0.20 and $1.20. Claude's Batch API discount of 50% and prompt caching at 0.1x base input narrow the gap but do not close it.
Which is better for coding?
Claude for repository-scale work, OpenAI for computer and browser control. Claude Code leads the official Terminal-Bench 2.1 board; the 80.3% (Mythos 5, invitation-only) SWE-bench Pro row in OpenAI's table belongs to a model with no self-serve access. OpenAI wins OSWorld 2.0, DeepSWE v1.1 and Terminal-Bench 3.0.
Why does this page not list a price for GPT-5.6 Sol?
Because OpenAI's own pages disagree. Its marketing pricing page and its developer documentation publish different output rates for Sol, and the two cannot be reconciled as a context-tier difference. Rather than pick one, this page omits the figure and links both sources.
Claude prices and plan mechanics verified 21 August 2026 against claude.com/pricing and platform.claude.com/docs. OpenAI figures checked against openai.com/api/pricing and openai.com/business/pricing. ARC-AGI results come from the ARC Prize leaderboard, index scores from Artificial Analysis and Elo from arena.ai. Both vendors change models and prices without notice - re-check before committing.