Back to Home

Claude Tops the Benchmarks. So Why Is Everyone Still Using GPT-4o Mini and DeepSeek?

Softcore Future Editorial
August 24, 20268 min readAI & Automation
Claude Tops the Benchmarks. So Why Is Everyone Still Using GPT-4o Mini and DeepSeek?

Anthropic's Claude models rank at or near the top of nearly every major AI benchmark in 2025, and the company is still reportedly struggling to convert that technical lead into market share, according to the Financial Times. The story, which generated 668 upvotes on Hacker News within a day, centers on a straightforward tension that anyone shopping for an AI model comparison will recognize immediately: the highest-scoring model is not automatically the one businesses choose to pay for.

This matters to anyone building products, running a startup, or managing an engineering budget in 2025, because the Claude situation is a live case study in what actually drives AI purchasing decisions. It is not raw capability. It is capability per dollar, and increasingly, capability that is "good enough" at a tenth of the price.

What the FT Report Actually Says Happened

The Financial Times reporting describes Anthropic's flagship models — the Claude line, including Opus and Sonnet variants — as commercially underperforming relative to their benchmark standing. Anthropic has built a reputation, backed by independent evaluations like LMSYS Chatbot Arena and various coding-specific benchmarks, for producing some of the most capable models available. Multiple developer surveys throughout 2025 have shown Claude Sonnet and Opus scoring at or near the top for coding tasks specifically, a category Anthropic has explicitly targeted as its wedge into enterprise accounts.

Despite that, the report frames Anthropic as losing ground on adoption to lower-cost alternatives. That includes OpenAI's smaller, cheaper models like GPT-4o mini, and open-weight models from labs like DeepSeek and Alibaba's Qwen family, which can be run at a fraction of the per-token cost of Claude's top-tier offerings. The FT's framing is explicit: technical superiority has not translated into the usage numbers or revenue trajectory that would justify Claude's pricing in the eyes of a large share of the market.

Anthropic has not disputed that cheaper competitors are pulling volume. The company's public strategy has instead focused on enterprise contracts, coding-specific tooling like Claude Code, and API partnerships — a bet that high-value, high-complexity use cases will pay a premium even if the mass market does not.

Who Gains When the Cheapest Model Wins

The clearest winners in this shift are the labs offering "good enough" models at commodity prices. DeepSeek's V3 and R1 models, released at a fraction of the training and inference cost of frontier labs, forced a public repricing conversation across the entire industry in early 2025. OpenAI's GPT-4o mini, priced at $0.15 per million input tokens against Claude Opus's roughly $15 per million, gives developers a 100x cost gap to justify before they even open a benchmark chart.

Startups building on top of these APIs are the second group of winners. A company running millions of customer support queries, content classifications, or summarization tasks a day does not need frontier-level reasoning for most of that volume. They need consistent, cheap, fast output — and multiple cheaper models now clear that bar. This is the practical center of most ai model comparison decisions being made inside procurement teams right now: not "which model is smartest," but "which model is smart enough."

Google also benefits indirectly. Gemini's tiered pricing structure, with Gemini 1.5 Flash and 2.0 Flash positioned explicitly as low-cost, high-volume options, has captured budget-conscious developers who might have defaulted to Claude a year ago on reputation alone.

pricing charts AI models compared pricing charts AI models compared.

Who Loses If Benchmarks Stop Mattering

Anthropic loses the most in this framing, and the stakes are direct: the company has raised funding at valuations built substantially on the premise that model quality is a durable moat. If the market increasingly treats frontier models as interchangeable commodities once they clear a "good enough" threshold, that moat narrows fast. The FT's reporting suggests this is already happening in segments where Claude's technical edge exists but doesn't matter enough to offset the price gap.

Enterprise sales teams pitching premium AI tooling lose a talking point too. "Best in benchmarks" has been a standard slide in vendor pitches across the industry for two years. If buyers are demonstrably choosing cheaper alternatives anyway, that argument weakens for every lab selling on quality rather than price, not just Anthropic.

There's a subtler loser here: model quality itself as a research priority. If the market rewards cost efficiency over marginal capability gains, the economic incentive to fund extremely expensive frontier training runs weakens. That's a real tension for a company whose research roadmap depends on continued heavy investment.

The Strongest Case Against This Framing

The obvious counterargument: benchmarks and revenue operate on different timelines, and judging Claude's commercial position on a single FT report risks mistaking a snapshot for a trend. Anthropic's revenue has grown substantially through 2024 and 2025 by multiple public reports, and enterprise contracts — the segment Anthropic is explicitly targeting — often take 12-18 months to close and don't show up in consumer-facing usage statistics at all.

There's also a real distinction between "losing users" and "losing the users who determine long-term profitability." A hobbyist switching from Claude to a free-tier competitor for casual chat represents a very different loss than an enterprise customer choosing a cheaper model for a high-volume pipeline. Anthropic's own public positioning has leaned into being the "developer's model" and the coding-specialist choice rather than chasing consumer volume, which makes raw adoption numbers a flawed proxy for the metric that actually matters to the business.

This counterargument holds up structurally, but it doesn't fully answer the FT's core point. Even granting that enterprise contracts lag and that Anthropic isn't chasing casual users, the reporting describes a company failing to convert acknowledged technical leadership into proportional commercial results in the categories it is actually targeting, including coding tools where Claude's benchmark advantage is largest and most publicly marketed. If the premium doesn't hold even in Anthropic's chosen battleground, the counterargument about timelines buys explanation, not exoneration.

developer choosing AI model options developer choosing AI model options.

What This Means for Your Own AI Model Comparison

The practical lesson for anyone running their own ai model comparison right now is that benchmark leaderboards answer the wrong question for most use cases. The right question is task-specific: does the cheaper model fail often enough, on your actual workload, to justify paying 10-100x more per token for the top-ranked alternative.

For high-stakes coding, complex multi-step reasoning, or tasks where a wrong answer is expensive, Claude's benchmark performance likely still justifies its price for many teams — this is precisely the segment Anthropic is defending. For high-volume, lower-complexity tasks — classification, extraction, basic summarization, customer support triage — the FT's underlying data point stands: cheaper models are closing the gap fast enough that the premium is hard to justify by usage volume alone.

The market signal from this story is not that Claude is a worse product. It's that "best" is no longer a single axis, and pricing power in AI now depends on defending a use case, not a leaderboard position.

benchmark leaderboard versus revenue graph benchmark leaderboard versus revenue graph.

Steps to Run Your Own Model Cost-Benefit Test

  1. Identify your three highest-volume AI tasks and run them through both a frontier model (Claude Opus, GPT-4o) and a budget model (GPT-4o mini, DeepSeek V3) using identical prompts.
  2. Score outputs on a simple pass/fail basis tied to your actual acceptance criteria, not subjective quality — does the cheap model's answer work for your use case or not.
  3. Calculate the real cost delta at your expected monthly volume, not per-token cost alone; a 50x price gap matters differently at 10,000 calls versus 10 million.
  4. Reserve the frontier model for the subset of tasks where the cheap model's failure rate creates measurable downstream cost, and route everything else to the cheaper option.
  5. Revisit this comparison quarterly — the price and capability gap between frontier and budget models has moved substantially every few months through 2024 and 2025.

Frequently Asked Questions

Is Claude actually better than cheaper AI models like GPT-4o mini or DeepSeek?

On most published benchmarks, yes — Claude's top-tier models consistently score higher on coding and complex reasoning tasks. Whether that advantage matters depends entirely on your specific task; for simpler workloads, the performance gap often doesn't justify the price gap.

Why would a company use a "worse" AI model on purpose?

Cost per token can differ by 10-100x between frontier and budget models, and for high-volume, low-complexity tasks, a "good enough" cheaper model produces acceptable results at a fraction of the cost. The FT's reporting suggests this calculation is now driving real market share shifts away from premium models.

Does this mean Anthropic is failing as a company?

No — the reporting describes a gap between benchmark leadership and broader market adoption, not company failure. Anthropic's enterprise and coding-focused strategy operates on a longer sales cycle than consumer usage statistics capture, and the company has shown substantial revenue growth through other reported figures.

Related Articles