If you're asking whether GPT 5.6 Soul is the best AI model for knowledge work in 2025, the short answer is: yes, for most people, most of the time. After a month of internal testing at Every — using it across coding, writing, design, and real-world automation — GPT 5.6 Soul earns the title of gold standard for everyday AI-assisted work. It's fast, ergonomic, relatively cheap, and smart enough to power the kind of agentic workflows that were previously only available to developers. Put on your sunglasses. Here's the full vibe check.
Is GPT 5.6 Soul the Best AI for Knowledge Work?
Think of GPT 5.6 Soul like a Porsche. It's a low-slung, precision-engineered machine built for daily use — fast, tightly controlled, and elegant rather than flashy. It's not trying to be the biggest model in the room. Instead, it's optimized to be the one you actually reach for every day, whether you're drafting emails, building internal tools, doing research, or managing a fleet of AI agents running in the background.
02:14
Dan Shipper introduces the Porsche analogy for GPT 5.6 Soul — powerful, precise, and built for daily use
Watch at 02:14 →
At Every, we spent a month putting this model through its paces before its public release, and the takeaway is clear: GPT 5.6 Soul inside the new ChatGPT Codex or ChatGPT Work desktop app is the current gold standard for knowledge work. It's not the best coding model. It's not the best design model. But it threads the needle across all of these categories better than anything else we've tested, and it does it at a speed and price point that makes it genuinely usable at scale.
GPT 5.6 vs Claude Fable: Which Model Wins?
This is the matchup everyone wants to see, and the honest answer is: it depends on what you need.
Claude Fable is still the S-tier model for hard coding tasks and design. It's chunky — you can almost smell how large it is. When you prompt Fable, it does things you didn't ask for (in a good way), like adding an "About" section to a video game it built unprompted, or producing design outputs that look like a senior creative director made them. It's world-changing genius, but it's slow, expensive, and your credits evaporate fast. One task can spin up 100 agents before you've had your coffee.
GPT 5.6 Soul, by contrast, is the everyday driver. OpenAI appears to have built a smaller, exceptionally well post-trained model that prioritizes ergonomics, speed, and reliability over raw power. For most knowledge work — writing, summarizing, automating workflows, processing emails — 5.6 doesn't just keep up, it often wins on the things that matter: clarity, speed, and the ability to run reliably in a loop.
05:30
Side-by-side comparison of the Library of Babel game built by GPT 5.6 vs Claude Fable — same prompt, noticeably different results
Watch at 05:30 →
The real pro move? Use both. Feed Fable a difficult coding task and explicitly instruct it to use GPT 5.6 as a sub-agent. You get Fable's architectural intelligence with 5.6's speed and cost efficiency doing the heavy lifting. It's a match made in heaven.
How Good Is GPT 5.6 Soul at Coding?
On our internal Senior Engineer Benchmark — which asks models to rewrite a vibe-coded mess of a codebase from scratch, from first principles — GPT 5.6 scored 56 out of 100. Fable scored 91. That gap is real, but the raw number undersells how 5.6 actually performs in practice. GPT 5.5 hit 62.5 on its best run, so there's meaningful variance in this benchmark.
The practical difference comes down to architectural judgment. Fable rewrites codebases with fewer abstractions and more elegance — it thinks like a senior engineer who's been burned by complexity before. GPT 5.6 does the rewrite, but the result tends to be more complicated than it needs to be.
For a more intuitive sense of the gap, we ran both models on what we call the Babel Bench: a one-shot prompt asking each model to build the Library of Babel from the Borges story as a playable video game. GPT 5.6 built something functional and reasonably well-designed. Fable built something with noticeably better graphics, more detail, and that unprompted "About" section — it just felt more considered.
The verdict: GPT 5.6 is A-tier for coding. It's what most of us use by default. When tasks get big, complicated, or architecturally ambitious, that's when you flip to Fable.
10:45
The Tend app in action — GPT 5.6 inside ChatGPT Codex processing emails into actionable decision cards
Watch at 10:45 →
Is GPT 5.6 a Better Writer Than Claude Opus?
Yes — and this might be the most underrated thing about this model.
Both Claude Opus 4.8 and Fable have a tendency to overexplain. They get a little literary. Fable in particular will run so long and lean so hard into metaphor that it starts speaking its own private language. GPT 5.6 just gets to the point. It's clean. It's clear. It doesn't have AI-isms.
The best example: our head of growth Austin reports that 5.6 is the first model he can genuinely one-shot marketing emails with. No cleanup, no heavy editing, no tone-deaf corporate phrasing. It just sounds like a human wrote it. Need to schedule a meeting? "Tucker — 4:30 ET on Tuesday the 14th works for me. Hope that still works on your end. Looking forward." Done. Fast. Right.
For metaphors, taglines, and pressure-testing a piece of writing you're trying to refine, 5.6 is now the go-to. The fact that it's also dramatically faster than Opus or Fable makes it even more useful — you're not sitting there watching a cursor blink for 40 seconds waiting for a subject line.
What Is the New ChatGPT Codex Desktop App?
The new ChatGPT desktop app is a merger of the old ChatGPT desktop app and the Codex desktop app — they're now one unified product. You'll see two modes: ChatGPT Work and ChatGPT Codex. Under the hood, they're essentially the same thing. The Work mode hides the code from you and presents outputs in a cleaner, more document-oriented way. The Codex mode looks more developer-facing and shows you what's happening at a technical level.
Together, these two modes plus the 5.6 Soul model represent OpenAI's vision for AI-assisted work: a single environment where you can go from writing a document to building an automation to managing a background agent — all in one place.
How to Use GPT 5.6 to Automate Your Knowledge Work
This is where GPT 5.6 Soul changes the game for non-developers. The big shift it enables isn't just "better answers" — it's the ability to move from doing knowledge work to managing a system that does knowledge work.
Developers have been living this reality for a while: instead of manually fixing a bug, you set up a system where a reported issue triggers an agent, the agent does the work, reviews it, and pushes it to production. GPT 5.6 is now smart enough, fast enough, and reliable enough to bring that same model to everyone else.
Here's what that looks like in practice at Every:
- Email triage: An internal app called Tend converts incoming emails into action cards. GPT 5.6 inside Codex processes each email, decides what action to take, and presents the decision for a quick approve or reject. Over time, it gets better at predicting the right call.
- Meeting summaries: When a meeting runs long or you have to leave early, Codex reads the transcript and surfaces what happened after you left — decisions made, follow-ups needed, people who need to be informed.
- Personal life automation: Voice notes and photos from your phone feed into a macro-tracking app. A photo of your lunch, a quick voice note about dinner — 5.6 pulls the macros and logs it automatically, running in a loop in the background.
The mental model here is important: you're not prompting the model to do tasks. You're tuning the system that runs the tasks, and then you're making decisions when decisions need to be made. Everything else runs itself. This is what "working on the company, not in the company" looks like when applied to individual knowledge work — and 5.6 is the first model reliable and affordable enough to make it practical.
Is GPT 5.6 Soul Actually Good at Design?
Better than 5.5 — which set a low bar — but not at the level of Fable or Opus 4.8. The model now shows genuine aesthetic awareness. When you ask it to design a website, it'll pause to think through a concept first: "I want a warm, cream-colored paper feel." That deliberate design thinking is new and real.
But the outputs still fall short of what Fable produces with the same prompt and the same underlying image model. The difference isn't the image model — both use GPT's image generation. The difference is in how each model prompts it. Fable's prompting produces something that looks considered and intentional. GPT 5.6's version looks busier, less refined.
If design quality is critical to your work, you'll still want Fable or Opus in the mix. For quick mockups, visual concepts, or designs where "good enough" is genuinely good enough, 5.6 will get you there faster and cheaper than anything else.
Final Verdict: Should You Use GPT 5.6 Soul?
Yes. For most people, most of the time, GPT 5.6 Soul is the model to have open all day. It's fast, affordable, a genuinely excellent writer, a capable coder, and — most importantly — it's the engine that makes real knowledge work automation finally accessible to non-developers. The ChatGPT Codex and Work apps give it the right harness to do that job well.
Use Fable when you need world-class coding or design. Use 5.6 for everything else, and consider pairing them together for the best of both worlds. The era of managing AI systems rather than just prompting them is here — and GPT 5.6 Soul is the model that makes it real.








