Claude Opus 4.5 is better than GPT-5.5 on both coding and writing benchmarks, at least according to a week of internal testing at Every. On the senior engineer benchmark, Opus 4.5 scores a 63 compared to GPT-5.5's 62. On the writing benchmark, it pulls further ahead with a 79.6 versus GPT-5.5's 73. Anthropic could have called this Opus 5 and nobody would have argued. So if you have been sleeping on Claude lately, it is time to wake up.

Is Claude Opus 4.5 Better Than GPT-5.5?

The short answer is yes, in most categories that matter for day-to-day knowledge work. After spending a week putting Opus 4.5 through its paces, the verdict is that this is the top-of-pack model right now. It edges out GPT-5.5 on coding by a single point, beats it more decisively on writing, and just feels better to use in a way that is hard to quantify but impossible to ignore.

The internal reach test rating system explained — gold means paradigm shift, green means solid daily driver 01:45 The internal reach test rating system explained — gold means paradigm shift, green means solid daily driver Watch at 01:45 →

To put that in context: for the past month or two, even the most devoted Claude users on the Every team had quietly started defaulting to Codex or GPT-5.5 for most of their work. Opus 4.7 was technically an improvement on paper, but it was slow, hard to love, and not particularly usable. The vibes were bad. Opus 4.5 changes that completely. It is not a marginal upgrade. It feels like a different category of model.

Kieran Classen, GM of Kora and one of the internal testers, called it the most human model he has ever worked with. That is high praise from someone who has been using these tools since the beginning. On the internal reach test — a simple but revealing measure of whether you actually grab for a model in real situations — Kieran gave Opus 4.5 a rare gold paradigm-shift rating. That does not happen often.

How Does Claude Opus 4.5 Score on Coding Benchmarks?

The senior engineer benchmark is one of the most demanding tests in Every's internal suite. It takes a vibe-coded codebase and asks the model to rewrite it from first principles — the same task given to two human senior engineers whose work serves as the baseline. Human engineers typically score in the 80s or 90s. Here is how the models compare:

  • Claude Opus 4.5: 63 out of 100
  • GPT-5.5: 62 out of 100
  • Claude Opus 4.7: approximately 33 out of 100

That 30-point jump from Opus 4.7 to Opus 4.5 is not a rounding error. That is a generational leap. And while both top models are still meaningfully below what a human senior engineer produces, the gap is closing fast.

Side-by-side comparison of the cozy island benchmark output from Opus 4.5 vs GPT-5.5 04:10 Side-by-side comparison of the cozy island benchmark output from Opus 4.5 vs GPT-5.5 Watch at 04:10 →

Beyond raw scores, Kieran's LFG Bench tested the models on real-world style tasks like building a SaaS product, an e-commerce site, and a 3D game landscape. Opus 4.5 consistently wrote more readable code and showed something a little unexpected: creativity. The outputs bridged the gap between technical correctness and genuine artistry. In the cozy island benchmark — where models are asked to generate a 3D island scene — Opus 4.5 produced something lush, rich, and detailed. GPT-5.5's version had more diversity of elements, but it felt flatter. Less alive. Opus 4.5 has depth and character. GPT-5.5 feels like it is optimized to ship.

Is Claude Opus 4.5 the Best AI Writing Model Right Now?

Based on internal benchmarks that tested introductions, promo emails, and mid-article paragraphs, yes. Claude Opus 4.5 is currently the best writing model that the Every team has tested. A score of 79.6 out of 100 versus GPT-5.5's 73 is a meaningful gap, not a statistical blip.

What makes it stand out is not just quality but feel. It is expressive. It does not read like a robot hedging every sentence. And critically, it picks up on a writer's voice from context in a way that is genuinely impressive. Give it a paragraph in your own voice and ask it to continue, and it will — accurately, naturally, in a way that does not immediately betray the seams.

The auto-generated slide deck on compound engineering produced by Opus 4.5 06:30 The auto-generated slide deck on compound engineering produced by Opus 4.5 Watch at 06:30 →

One caveat: performance degrades noticeably on medium reasoning settings. If you are using Opus 4.5 for writing and wondering why it feels flat, try bumping up to high or extra high. The difference is real and significant.

Claude App vs Codex: Which Should You Actually Use?

Here is the honest answer: Codex is still the better daily driver, and that has nothing to do with the model. It is about the harness. The Codex desktop app is fast, clean, and simple. The in-app browser works well and changes the game for knowledge work. It feels like the future.

The Claude desktop app, by contrast, carries the scars of how Anthropic got here. There is a chat tab, a code tab, a co-work tab — and each one feels like it was built by a different team because it probably was. Every time you open it, you spend a few seconds wondering which tab you should be in. That friction adds up.

That said, Opus 4.5 is so good that it is now worth switching between apps depending on the task. The model is pulling people back to the Claude app who had fully committed to Codex. Think of it as using the best tool for each job rather than picking one and never leaving.

Which Reasoning Settings Get the Best Results in Opus 4.5?

This is one of the most important practical tips for anyone starting to use the model. Claude Opus 4.5 is very sensitive to reasoning settings, more so than previous models. Performance on extra high reasoning is noticeably stronger than on high or medium — for both writing and coding.

For your hardest programming challenges, use extra high. For important writing, use at least high. For casual tasks, medium is fine, but do not be surprised if the output feels less sharp. Think of it like a sports car that really needs premium fuel to hit its top speed. The engine is excellent; just do not cheap out on the settings.

How Good Is Claude Opus 4.5 for Knowledge Work and Presentations?

One of Every's standard knowledge work tests is generating a slide deck explaining a complex topic from scratch. Most models produce something that works but feels thin — technically correct, visually passable, intellectually hollow. Opus 4.5 did not do that.

When asked to create a beginner's deck on compound engineering, the output had depth, real styling, and a coherent structure that held up as a first draft. For anyone who has tried to auto-generate a slide deck before and been disappointed, this felt meaningfully different. It is not a finished product, but it is a strong foundation — exactly what you want from an AI collaborator.

More broadly, Opus 4.5 handles the kind of context-switching that real knowledge work demands. Move from a coding task to a writing task to a strategic question in the same thread, and it tracks. It does not lose the thread. That versatility is rarer than it sounds.

Can Claude Opus 4.5 Help With Personal and Interpersonal Thinking?

Surprisingly, yes. One of the more unexpected findings from testing was how well Opus 4.5 handles emotionally intelligent, interpersonal work — management situations, personal decisions, friendship dynamics, internal psychology. It is not just empathetic in a surface-level way. When you look at the model's thinking traces, it is genuinely working through the permutations of a situation before responding.

What makes it particularly useful here is that it pushes back on your frame without being disagreeable. It is not a yes-machine. It expands how you might think about something rather than validating whatever you already believe. For anyone using AI as a thought partner rather than just a task executor, this is a meaningful capability.

Bottom line: Claude Opus 4.5 is a legitimately great model, probably the best available right now. If you have been living in Codex, add this to your rotation. If you are a longtime Claude user who got frustrated with 4.7, come back. Anthropic underpromised and overdelivered — and that is worth paying attention to.