How Good Is Claude Opus 4.7 vs Other AI Models?

Claude Opus 4.7 is out, and the short answer to how good it is: it depends entirely on what you're asking it to do. In most industry-standard benchmarks, Opus 4.7 outperforms its predecessor Opus 4.6 — but it consistently falls short of Claude Mythos Preview, a model most people still can't access. On real-world professional tasks, Anthropic claims it's ahead of every generally available model. But dig even a few pages into their system card and you'll find genuine regressions, deliberate capability caps, and a compute bottleneck that's already frustrating paying customers. This is not a clean win. It's a complicated upgrade.

What Do Claude Opus 4.7 Benchmarks Actually Show?

Across standard benchmarks — coding, obscure knowledge, computer navigation — Opus 4.7 beats Opus 4.6 but trails Claude Mythos Preview. That pattern holds consistently. On generalized knowledge work benchmarks, Opus 4.7 appears to edge out competitors like Gemini 2.5 Pro, which is presumably what prompted Anthropic to state on page three of the system card that the model is ahead of all generally available models on real-world professional tasks.

Benchmark chart showing Opus 4.7 vs Opus 4.6 vs Mythos Preview across multiple task categories 02:15 Benchmark chart showing Opus 4.7 vs Opus 4.6 vs Mythos Preview across multiple task categories Watch at 02:15 →

But that framing conveniently sidesteps some unflattering comparisons. On BrowseComp — a benchmark testing agentic web search, the ability to hunt down hard-to-find snippets across the internet — Opus 4.7 actually underperforms Opus 4.6. Strikingly, even Mythos Preview underperforms GPT-4.5 on that same test. On vibe coding (building a web app from scratch), Val's AI ranked Opus 4.7 as the best model available, beating GPT-4.5 on performance and speed. But on OCR — visually parsing dense documents — Opus 4.7 loses to Gemini 2.0 Flash, a model that costs more than ten times less. That's not a rounding error. That's a meaningful gap in a genuinely common workflow.

On cybersecurity vulnerability reproduction, Opus 4.7 deliberately underperforms — and Anthropic admits it. Page 48 of the system card states plainly that during training, they experimented with reducing these capabilities on purpose. They don't want the model to be too capable at finding certain security vulnerabilities. Whether you find that reassuring or paternalistic probably depends on your threat model.

Why Does Opus 4.7 Underperform Opus 4.6 in Some Areas?

Some of the regressions are intentional; others appear to be genuine surprises. On certain long-context reasoning tasks — like locating the fourth poem buried across 1 million tokens — Opus 4.7 regresses compared to 4.6, even at max settings. The lead creator of Claude Code acknowledged this benchmark publicly but argued it's being phased out anyway because it relies on stacking distractors to trick the model rather than testing genuine reasoning.

Real-world demo: Opus 4.7 skipping the tooltip step that every prior Claude model handled automatically 06:40 Real-world demo: Opus 4.7 skipping the tooltip step that every prior Claude model handled automatically Watch at 06:40 →

A more telling real-world example: every previous Claude model, when adding itself to a benchmarks leaderboard on one developer's web app, would automatically attach an Open Roster tooltip to new entries. Opus 4.7 was the first model to skip that step entirely — it just didn't bother. The developer had to instruct it explicitly. That's not a catastrophic failure. But it illustrates something real about how the model calibrates effort.

What Is Adaptive Thinking and Why Does It Matter?

Adaptive thinking is Opus 4.7's approach to inference compute: if the model decides your task is easy, it spends less time reasoning through it. That sounds efficient. The problem is that the model sometimes misjudges difficulty. Simple Bench — a benchmark built around trick questions requiring common sense to see through — actually scores worse for Opus 4.7 than Opus 4.6, apparently because the model treats those deceptively simple questions as genuinely simple and under-thinks them.

More controversially, adaptive thinking is now mandatory if you want extended thinking from Claude at all. You cannot force the model to always think longer or always use more compute. You can encourage it toward deeper reasoning, but you can't lock it in. Before Opus 4.7 even launched, one AMD senior AI director flagged that thinking character counts in Claude 4.6 had already dropped by three-quarters — far less reasoning, far more bailing out. The lead creator of Claude Code confirmed that medium effort is now the default. To get high or max effort, you have to set it explicitly. This is a meaningful change for developers who relied on deep reasoning as a baseline.

Is Anthropic's Compute Problem Hurting Claude's Quality?

This is arguably the most important context for understanding all of the above. According to an internal OpenAI memo leaked to The Verge, OpenAI believes Anthropic has made a strategic error in not securing enough compute — and that this will show up in the product. Customers may already be experiencing it through throttling, weaker availability, and a less reliable experience.

Market share graph showing Claude and Gemini both 4x-ing traffic while OpenAI approaches sub-50% 12:10 Market share graph showing Claude and Gemini both 4x-ing traffic while OpenAI approaches sub-50% Watch at 12:10 →

There's some irony here. Claude and Gemini have both roughly quadrupled their share of generative AI web traffic compared to this time last year. OpenAI's share may drop below 50% this month. But the runaway success of Claude seems to have created a genuine supply-side problem. Sam Altman has implicitly joked about Claude's rate limits. One of the Codex leads pointed out that Codex is compute-efficient, always available, never down. OpenAI's chief revenue officer went further in the memo, claiming Anthropic's story is built on fear and restriction, and that their revenue run rate is overstated by around $8 billion — putting it closer to $22 billion, still behind OpenAI.

What Is Claude Mythos Preview and Is the Hype Real?

Claude Mythos Preview remains inaccessible to the public — available only to select insiders like US government agencies and major tech companies. Anthropic's most headline-grabbing claim about Mythos was that it was accelerating their own engineers' work output by 4x, based on an internal survey. That number spread quickly and prompted serious people to ask whether recursive self-improvement was now imminent.

Page 29 of the Opus 4.7 system card gives more detail about that survey — and it falls apart under scrutiny. The actual question asked was how much more output engineers produced over the past week compared to having no model access. Not time saved, not quality improved — just volume of output. And it was opt-in, meaning the people most likely to respond were those who had used Mythos most and benefited most. It was, in short, an extraordinarily unscientific survey dressed up as a capability milestone.

On the cybersecurity front, an external lab called Vidok tried to replicate Mythos's most publicized vulnerability discoveries using other models like Opus 4.6 and GPT-4.5. In nearly every case, those models — with the right scaffolding — were able to reach the same core vulnerability. The better framing, as one researcher put it, isn't that one lab has a magical model. It's that finding vulnerability signals is getting cheaper across the board. That still matters enormously for banks and security teams. It's just a different kind of milestone than the one being implied.

The 9-Year Dario vs Brockman Rivalry Behind the AI Wars

Understanding the Anthropic vs OpenAI dynamic requires going back to around 2016, when Dario Amodei joined OpenAI. According to reporting likely drawn from Amodei's own notes, he would work late into the night with Greg Brockman — now the lead on Codex, OpenAI's direct competitor to Claude Code. The tension between them escalated through a series of incidents: a mass layoff Amodei found needlessly cruel, a proposal to sell AGI access to nuclear powers including Russia and China that Amodei considered tantamount to treason, and ultimately a 2020 declaration that he simply couldn't work with Brockman anymore.

That rivalry is now playing out in the most commercially significant arena in AI: coding assistants. Brockman recently revealed in an interview why OpenAI fell behind Anthropic in coding — they optimized for abstract programming competitions while Anthropic grounded their training data in messy, real-world codebases. Brockman believes OpenAI has now caught up, pointing to a strong pipeline of upcoming models. Amodei, meanwhile, is pushing Mythos toward an enterprise tier that remains largely out of public reach.

What New Claude Code Features Actually Ship With Opus 4.7?

Amid all the controversies, Anthropic did ship some genuinely useful Claude Code upgrades. Routines (currently in research preview) lets you trigger prompts on a schedule — your laptop doesn't even need to be open. The new ultra review command offers a comprehensive code audit, and in testing it caught a bug that GPT-4.5 had missed (though GPT-4.5 also caught one that Claude missed). Dispatch lets you assign tasks to Claude from your phone, which then runs them on your local machine via the desktop app. These are incremental but meaningful quality-of-life improvements for developers living inside Claude Code full-time. The underlying model may be imperfect. The tooling around it is getting more capable by the week.