If you're wondering what GPT 5.6 can do, the short answer is: a lot more than you'd expect from a dot release. This isn't a minor patch. OpenAI has effectively squeezed every last drop of performance out of the GPT 5 training run, and the result is one of the most capable, practical AI models available right now. From cloning Excel with pivot tables to building a fully playable Minecraft world — all from a single prompt — GPT 5.6 paired with Codex is genuinely shocking in its output.

Let's break down everything this model can do, how it compares to the new Fable model, and how you should be thinking about routing between Luna, Terra, and Sol for your own workflows.

The Excel clone running live — sorting, formulas, and pivot tables all functional 0:45 The Excel clone running live — sorting, formulas, and pivot tables all functional Watch at 0:45 →

What Can GPT 5.6 Actually Do?

The capabilities of GPT 5.6 become immediately clear when you see what it produces under long autonomous loops. The model shines brightest when paired with Codex and given an open-ended goal with a simple directive: continue until feature parity. That's it. Eight words. And what follows is days of continuous, structured, increasingly sophisticated output.

Here's what makes GPT 5.6 stand apart from previous versions:

  • Deep browser and computer use: It can open applications on your desktop, interact with them, and use what it sees to inform its own output. It literally opened Excel, studied it, and used that as a reference while building its own clone.
  • Long-horizon autonomous task completion: It ran for 5 to 7 days on single prompts without losing context or coherence.
  • Clean, functional code output: The apps it builds aren't just demos — they have real interactivity, formula support, pivot tables, and 3D rendering.
  • More efficient token usage: It reaches the same result with fewer tokens compared to GPT 5.5, which translates directly into cost savings.

How Did AI Build a Full Excel Clone in 5 Days?

The Excel clone demo is one of the most convincing showcases of what GPT 5.6 with Codex can produce. The prompt was almost insultingly simple: /goal make an Excel clone continue until feature parity. That's all it took to kick off a 5-day autonomous build session.

GPT 5.6 using computer use to open real Excel and reference it while building the clone 3:20 GPT 5.6 using computer use to open real Excel and reference it while building the clone Watch at 3:20 →

What came out the other side is a single-page HTML app that includes:

  • Sortable columns (ascending and descending)
  • Formula support — including chained cell references like =A1+B1
  • An inline formula bar that populates when you double-click a cell
  • Data validation and conditional formatting
  • Tables, find and replace, and toggling
  • A full pivot table builder — select data with headers, configure rows and filters, generate the table

One of the most remarkable parts of this build was how it worked. GPT 5.6 used computer use to open the actual Microsoft Excel on the desktop, study specific behaviors and UI patterns, and then replicate them in the clone. It was essentially using the real product as a live reference document. That kind of feedback loop — observe, replicate, refine — is what separates this model from anything before it.

The AI-built Minecraft clone with 3D shadows, mobs, farming, and world generation 5:10 The AI-built Minecraft clone with 3D shadows, mobs, farming, and world generation Watch at 5:10 →

Yes, there are rough edges. It was manually stopped after 5 days and was nowhere near done. But the fact that a week of compute time produced a functional, demo-ready spreadsheet application from an 8-word prompt is genuinely hard to wrap your head around.

Can AI Actually Clone Minecraft? Here's What Happened

The Minecraft clone is arguably even more impressive. Using the same /goal structure — create a clone of Minecraft, feature parity — Codex ran for approximately 7 days. Within just the first day, it had produced something that visually and functionally resembled actual Minecraft.

Box AI benchmark results comparing GPT 5.6 Sol, Terra, Luna, and Fable across enterprise tasks 7:45 Box AI benchmark results comparing GPT 5.6 Sol, Terra, Luna, and Fable across enterprise tasks Watch at 7:45 →

But it didn't stop there. It kept going deeper:

  • Building out biomes that didn't previously exist in the clone
  • Adding mobs directly sourced from the real Minecraft game
  • Implementing farming mechanics (farmland, carrots, harvestable crops)
  • Rendering glass blocks with breakable behavior
  • Generating unique world seeds
  • Building out a full inventory system

The shadowing and 3D animation quality are noticeably better than any previous AI-generated Minecraft clone. It's not a perfect replica — Minecraft took years of development by a full team — but as a benchmark for what a single autonomous AI loop can produce, it's remarkable. This is the best AI-built Minecraft clone produced to date.

GPT 5.6 vs Fable: Which Model Wins?

This is the comparison everyone is asking about, and the honest answer is: it depends on what you need right now.

GPT 5.6 is the absolute peak of an existing training run. Think of it as the most souped-up Honda Civic you've ever seen — every horsepower extracted, tires optimized, spoiler on. It's refined, it's fast, it's efficient, and it gets the job done with a direct line of sight to the goal.

Fable feels like a Ferrari fresh off the manufacturing line — unoptimized, but with vastly higher potential. It sees around corners better. It feels like a genuinely new model, not an iteration. If GPT 5.6 is the pinnacle of what's possible on one training run, Fable is the beginning of something much larger.

Box AI ran their own enterprise benchmark across real knowledge work tasks — document reading, number reconciliation, due diligence, expert error review — and the results showed GPT 5.6 Sol outperforming GPT 5.5 across public sector, life sciences, and healthcare categories. Terra and Luna perform slightly lower on accuracy but win significantly on speed and cost.

For most practical use cases today, GPT 5.6 is the better choice. Fable's ceiling is higher, but it hasn't been optimized yet.

Is GPT 5.6 Cheaper Than GPT 5.5?

Yes — significantly. Here's the pricing breakdown:

  • Input tokens: $5 per million for GPT 5.6 vs $10 per million for Fable
  • Output tokens: $30 per million for GPT 5.6 vs $50 per million for Fable
  • Cache hits: Much cheaper on GPT 5.6 as well

But the price difference isn't the only savings. GPT 5.6 also uses fewer tokens to reach the same result. That means you're paying less per query and consuming less quota to get equivalent output quality. For high-volume use cases or long agentic loops, this adds up fast.

Why Is Codex Browser Use a Game Changer?

Browser use inside Codex has quietly become one of the most powerful features in this entire ecosystem. It's increasingly replacing a standard browser for day-to-day tasks. Real examples of what it handles well:

  • Sorting and triaging Gmail inboxes
  • Making complex DNS record changes from a single prompt
  • Using real applications (like Excel) as live references during code generation

The combination of browser use plus computer use means GPT 5.6 isn't just generating code in isolation — it's interacting with the real world, gathering live context, and feeding that back into its output. That feedback loop is what made both the Excel and Minecraft clones as accurate as they are.

Luna, Terra, and Sol: How to Route GPT 5.6 Models

GPT 5.6 comes in three sizes — Luna (small), Terra (medium), and Sol (large) — and each supports multiple levels of reasoning effort, up to Ultra on Sol. This creates a flexible routing system for complex workflows.

A practical model routing strategy looks like this:

  • Sol on high reasoning — planning, architecture decisions, complex analysis
  • Terra on standard reasoning — most implementation work, code generation, iteration
  • Luna — deployment steps, low-complexity tasks, anything that doesn't need heavy compute

You can now handle entire agentic pipelines entirely within the GPT model family, without needing to call out to Claude or other external models. A skill has been written specifically for this kind of delegation inside Codex — routing tasks automatically between Luna, Terra, and Sol based on complexity — which is available on GitHub for anyone who wants to save quota while maintaining output quality.

GPT 5.6 isn't the future of AI — Fable probably is. But right now, for real work, GPT 5.6 is the most capable and cost-effective model in the lineup. And if a single 8-word prompt can produce a working Excel clone in 5 days, it's worth paying attention to what comes next.