Independent & unofficial. Not affiliated with Anthropic. Facts verified 21 August 2026. Always confirm pricing at claude.com/pricing.
Review

Claude review: an honest assessment

A critical review of Claude in August 2026 - what it is genuinely best at, the five places it clearly loses to rivals, and whether Pro, Max or the API is worth the money.

Claude is the best available model for long agentic coding runs and for prose people actually prefer, and it is the wrong purchase for anyone who needs media generation, reliable retrieval across very large documents, low latency, or a bill under $20 a month. That is the whole verdict. As of August 2026 Anthropic sells the most expensive and the slowest frontier family in the market, and it wins enough of the hard benchmarks to make that defensible for some buyers and indefensible for others.

This is claudiai.com's own assessment, written from public documentation and independently-run benchmarks. It is not a roundup of other people's reviews, it contains no testimonials, and this site has no relationship with Anthropic.

#The verdict in one box

Best at
Agentic coding, novel reasoning (ARC-AGI-3), long-form writing
Worst at
Generative media (nothing at all), long-context retrieval, price, speed
Buy Pro ($20/mo) if
You code, write or run agents daily and value output quality over cost
Do not buy if
You need images, video or audio, or your budget is under $20/mo
Biggest unforced error
Publishing flagship benchmarks as an image rather than a table
Assessment date
21 August 2026

#Five things Claude is genuinely worse at

Stated first, because a review that saves its criticism for the last paragraph is not a review.

#1. It generates no media of any kind

No images, no video, no audio, no music - and it will not accept video or audio as input either. Claude models are documented as text and image in, text out. OpenAI ships GPT-Image-2, Sora 2 and a realtime speech line; Google ships Veo 3.1, Imagen and Lyria 3. That is a categorical absence, not a quality gap: it disqualifies Claude from any media workflow regardless of how well it reasons.

#2. It is the most expensive frontier family, and the slowest

Fable 5 at $10 in / $50 out per million tokens is the highest headline price in this market. Against it, Gemini 3.7 Flash costs $0.75/$3.75, DeepSeek V4-Pro $0.435/$0.87 and GPT-5.6 Luna $0.20/$1.20. Speed is worse still: Artificial Analysis measures Fable 5 at roughly 75 output tokens per second against Gemini 3.7 Flash's roughly 3,901. Opus 5 at around 626 tokens/sec is far more usable, and at $5/$25 it is half Fable 5's price for a score one point higher on the independent Intelligence Index - which raises a fair question about who Fable 5 is actually for.

#3. Its long-context retrieval is the weakest of the big three

Anthropic ships a 1,000,000-token window as the default on Fable 5, Opus 5 and Sonnet 5, with no beta header and no long-context surcharge. That is genuinely good engineering. But on GDM-MRCR v2, Google's model card reports Gemini 3.7 Flash at 97.0% against Sonnet 5's 81.5%, with GPT-5.6 Terra at 93.5% in between. A window measures what you can load; retrieval measures what the model can find again. Anthropic has published very little long-context evidence of its own and loses the benchmarks that exist. The exception is narrow: on GraphWalks BFS at 1M tokens, OpenAI's own table gives Anthropic's Mythos 5 79.4% against Sol's 77.1%.

#4. There is nothing between $0 and $20, and the free tier is metered

ChatGPT Free and ChatGPT Go both offer unlimited text chat, ad-funded for US users. Google AI Plus sits at roughly $5–$8, and ChatGPT Go at $8. Anthropic's free plan is a metered weekly pool with a five-hour rolling refresh, and its next step up is $20. That missing tier is the band where students and people still evaluating the product live, and Anthropic has ceded it. Distribution compounds the problem - Microsoft 365 Copilot puts GPT models in front of every Office seat from about $18/user/month, and Gemini is the default in Search, Gmail and Android. Claude has to be sought out.

#5. Its benchmark transparency is the worst of the three big labs

This is the one that should embarrass Anthropic most, because it is entirely self-inflicted. OpenAI and Google both publish full numeric comparison tables as machine-readable text on their launch posts and model cards. Anthropic published the Fable 5 benchmark table only as an image, and the Opus 5 announcement gives relative claims - "3× the next-best model on ARC-AGI-3", "within 0.5% of Fable 5 at half the cost" - instead of absolute figures.

The consequence shows up all over this site: there is no published numeric OSWorld figure for Opus 5, no Humanity's Last Exam figure for Opus 5 or Fable 5, and no Anthropic-run table anyone can cite. Almost every Fable 5 number in circulation comes from OpenAI's or Google's tables - a strange position for the leader on two independent leaderboards. A company that asks to be trusted on safety should be the easiest to verify, not the hardest.

#Category by category: what is actually good?

Assessments below are this site's own, based on the evidence cited. No external review scores are aggregated, because comparable ones do not exist.

CategoryAssessmentReasoning
Coding and agentsBest in classClaude Code plus Fable 5 tops the official Terminal-Bench 2.1 board at 83.8%. OpenAI's own launch table also concedes SWE-bench Pro at 80.3% (Mythos 5, invitation-only) against GPT-5.6 Sol's 64.6%, but Mythos 5 is not generally available and no figure is published for Opus 5 or Fable 5, so the verdict rests on Terminal-Bench. Caveat: OpenAI's Ultra multi-agent mode holds the highest published number at 91.9%, and OpenAI wins OSWorld 2.0 and DeepSWE v1.1
WritingBest in classSix of the top twelve arena.ai human-preference slots, Fable 5 at #1 with 1507 Elo, and no GPT model in the top fifteen. Caveat: Meta's Muse Spark 1.2 at 1498 outranks Opus 5 at 1493, and Opus 5 sits below Anthropic's own Opus 4.6 and 4.7 - the newest model is not the best-liked
Long-context workAdequate, over-marketed1M window by default with no surcharge, but the weakest published retrieval of the big three. Fine for structured material you control; risky for needle-finding in a large unfamiliar corpus
MultimodalityPoorText and image in, text out. No generation of any media type, no video or audio input. There is no partial credit available here
Ecosystem and toolingStrong but narrowClaude Code, Cowork, Claude in Chrome, Claude for Excel, Design and Science, plus MCP - now a Linux Foundation standard adopted by OpenAI too. But no search engine, no operating system, no office suite: distribution is the structural weakness
Price and valueWeak on paper, defensible in narrow casesThe most expensive family in the market. Justified only where a failed task costs more than the tokens. For classification, extraction and routing, cheaper models win outright
Rate limitsOpaqueTwo overlapping windows, a separate Opus counter, one shared pool across every surface, and no published absolute figures - only relative multipliers
Enterprise readinessStrong, with carve-outs that matterPublished flat pricing and a deep control list, undercut by exclusions buried in the BAA and retention documentation

#Coding and agents - the reason to buy it

This is where the premium is earned. On long-horizon tasks - a multi-hour refactor, a migration across dozens of files, an agent loop that must recover from its own mistakes - Claude's failure rate is what you are paying to reduce, and the benchmark record supports that. The tooling is also the most developed: skills, subagents (four built in, up to twenty concurrent, three levels of nesting), agent teams, roughly twenty hook events and a plugin marketplace, all covered on Claude for developers.

Where it is not best: desktop and browser control. OpenAI leads OSWorld 2.0 by roughly eight points, and Anthropic's computer-use tooling only reached general availability on 19 August 2026, alongside a browser-use tool. That is late.

#Writing - better than the benchmarks suggest, and honestly assessed

Human preference is the only measure that matters for prose, and Anthropic wins it. The interesting detail is internal: Opus 5 ranks below Opus 4.6 and Opus 4.7 on arena.ai despite topping the Intelligence Index - evidence that it was optimised for agentic capability rather than conversational feel. If drafting is your main use, test the older Opus models against the newest rather than assuming the latest is best - see Claude as a writing assistant.

#Rate limits - the most common complaint, and a fair one

Anthropic enforces two overlapping windows: a rolling five-hour session window and a weekly limit that resets on a fixed day assigned to your account. Opus models are tracked on a separate weekly counter. Chat, Claude Code, Cowork, Design and Claude for Excel all draw from one shared allocation, so a heavy day of agentic coding consumes the budget you were going to use for chat. Anthropic publishes relative multipliers only - 5× and 20× Pro for Max, 1.25× for a standard Team seat - and no absolute counts, so you cannot forecast whether a plan fits your workload before buying. Extra usage credits exist on Pro and Max at API rates, capped at $2,000/day redemption. The mechanics are laid out on the pricing page.

#Enterprise readiness - strong list, read the footnotes

Enterprise is published at a flat $20/seat/month plus usage at API rates, covering SCIM, audit logs, a Compliance API, custom data retention, customer-managed encryption keys, US-only inference, IP allowlisting and custom roles. Against OpenAI's quote-only Enterprise tier that transparency is a real advantage - though OpenAI advertises data residency across ten regions and Anthropic's public page does not.

Three carve-outs buyers miss

Cowork is not covered by Anthropic's BAA, and Claude Code is covered only under zero data retention in specific modes; Design and the Office betas are excluded entirely. Claude for Excel bypasses custom data retention, audit logs and the Compliance API. And default retention is indefinite unless an owner configures it - the 30-day figure is a minimum you must set, not a default you inherit. If you are buying for a regulated environment, confirm each surface individually rather than trusting the plan-level list.

#What Claude is genuinely best at

#Novel reasoning

ARC-AGI-3: 30.16% for Opus 5 against GPT-5.6 Sol's 7.78%, verified by ARC Prize. That is the largest single-benchmark lead any lab holds over any other, and Opus 5 was tested only at high effort, so it is arguably a floor.

#Independent rankings

#1 on the Artificial Analysis Intelligence Index (Opus 5 at 63) and #1 on arena.ai human preference (Fable 5 at 1507). Both third-party, both standardised, neither run by Anthropic.

#Terms you can live with

Commercial use is permitted on every plan including Free; the Commercial Terms state Anthropic may not train on customer content; consumer training is opt-in. Details on commercial use rights.

#Should you buy it? Verdicts by who you are

You are…VerdictReasoning, and what to buy instead
An individual writer Buy Pro, at $200/year Human preference is the only prose measure that counts and Anthropic wins it. $200/year works out cheaper than ChatGPT Plus, which has no annual discount. If you also need images for the same work, buy ChatGPT instead and accept slightly weaker prose
A solo developer Buy Pro; upgrade to Max only when you hit limits Claude Code is included on Pro and is the strongest agentic coding tool available. Do not start on Max - the shared usage pool means you cannot predict your consumption until you have run a few real weeks. If your work is browser and desktop automation rather than repository work, Codex is the better tool
An engineering team of 5–50 Buy Team, standard seats $20/seat/month annually, Claude Code on every seat, no training on your content by default, and enterprise search across Slack and Microsoft 365. Be aware Team cannot access the Compliance API or audit logs - if you need either, you are buying Enterprise, not Team
An enterprise buyer Shortlist it, but verify per surface The published control list is the best of the three at the base price. The carve-outs above are real and are not on the pricing page. If customer-managed keys plus multi-region data residency is the hard requirement, OpenAI currently advertises more; if EU sovereignty or on-premise is the requirement, Mistral is the only credible answer
Budget-conscious or occasional Do not buy The metered free plan will frustrate you and there is nothing at $8. Use ChatGPT Free for unlimited text, or ChatGPT Go at $8, or Google AI Plus at roughly $5–$8. For API work at volume, DeepSeek V4-Pro at $0.435/$0.87 is a fraction of Claude's cheapest rate. The field is surveyed on Claude alternatives

#How to test it before committing

Run your three hardest real tasks - not benchmark prompts - through Sonnet 5 first, the model most people should default to, and escalate to Opus 5 only where Sonnet fails. Run the same tasks on a competitor. Then check the bill, not the benchmark. Two API-side details change the arithmetic: the Batch API applies a flat 50% discount and prompt-cache reads cost 0.1× base input, both explained on the Claude API page. One cuts the other way - Sonnet 5's tokenizer emits roughly 30% more tokens for the same text than Sonnet 4.6, so its permanent $2/$10 rate saves less than the sticker price implies.

If the comparison you actually need is against a specific rival, the detailed head-to-heads are Claude vs ChatGPT and Claude vs Gemini. If you have already decided on Claude and only need to pick a model, start at the model index.

#Frequently asked questions

Is Claude worth paying for in 2026?

For daily coding, agent work or professional writing, yes - Pro at $200 a year is the cheapest of the flagship $20 tiers because ChatGPT Plus has no annual discount. For occasional use, media work or high-volume API traffic, no: cheaper and more capable options exist for each of those jobs.

What is Claude worst at?

Generative media, which it cannot do at all. It produces no images, video, audio or music and accepts no video or audio input. After that: long-context retrieval, where Gemini scores 97.0% to Sonnet 5's 81.5%, and price, where Claude is the most expensive frontier family available.

Is Claude Pro or Max better value?

Start on Pro. Max costs $100 or $200 a month for 5x and 20x Pro usage, but because chat, Claude Code, Cowork, Design and Excel all draw from one shared pool, you cannot predict your consumption until you have run several real weeks on Pro first.

Why is Anthropic criticised over benchmark transparency?

Because it publishes less verifiable data than its rivals. The Fable 5 benchmark table exists only as an image, and the Opus 5 announcement gives relative claims rather than absolute numbers, while OpenAI and Google publish full numeric tables as text. That makes independent verification harder for Claude than for either competitor.

Is Opus 5 better than Fable 5?

For most work, yes. Opus 5 costs half as much, runs roughly eight times faster, has a newer knowledge cutoff of May 2026, and scores one point higher on the independent Artificial Analysis Intelligence Index. Fable 5 leads human preference and is aimed specifically at long-running agent workloads.

Does this review include scores from other publications?

No. Every assessment here is this site's own, drawn from public documentation and independently-run benchmarks that are named and linked. There are no aggregated review scores, no user testimonials and no author credentials being claimed. claudiai.com is an independent reference site with no relationship to Anthropic.

Verify it yourself

Plans, limits and API rates verified 21 August 2026 against claude.com/pricing, support.claude.com and platform.claude.com/docs. Independent benchmark figures come from ARC Prize, Artificial Analysis and arena.ai. Anthropic changes models, prices and limits without notice - re-check before you buy.