GPT 5.4 outperforms human workers on their first attempt 70.8% of the time across 44 white-collar occupations — and if you include ties, that figure rises to 83%. That is the headline number from OpenAI's new GDP-Val benchmark, and it landed just 48 hours after GPT 5.3 Instant was released. Whether that pace reflects genuine acceleration or a very deliberate attempt to dominate headlines is, frankly, both. The answer is probably both. But the number itself deserves serious attention from anyone who earns a living at a desk.

How Does GPT 5.4 Compare to Human Workers?

The GDP-Val benchmark was designed specifically to measure AI performance against human output across occupations selected for their contribution to GDP — hence the name. GPT 5.4 was blind-graded by domain experts who evaluated its outputs alongside human outputs without knowing which was which. The result: the model wins more often than it loses, by a significant margin.

The GDP-Val benchmark results showing GPT 5.4 beating human first attempts 70.8% of the time across 44 occupations 01:45 The GDP-Val benchmark results showing GPT 5.4 beating human first attempts 70.8% of the time across 44 occupations Watch at 01:45 →

The self-driving analogy is apt here. We may not have reached the threshold of complete safety or reliability, but we may have crossed the milestone where, mile for mile — or spreadsheet for spreadsheet — an autonomous agent like GPT 5.4 is statistically more accurate than a human first attempt. As Waymo has demonstrated in transportation, even 10x better performance does not guarantee immediate mass adoption. So white-collar work is not vanishing overnight. But the trajectory is unmistakable.

There are important caveats, though. The tasks in the benchmark are self-contained and digital. They are not representative of the full range of responsibilities, relationships, and judgment calls that define most professional roles. And the benchmark does not adequately penalise catastrophic failures — mistakes a human would simply never make — which still occur with meaningful frequency.

Is AI Actually Replacing White-Collar Workers in 2026?

The blunter question underneath the benchmark numbers is whether professionals should be worried about their jobs. The honest answer is: it depends on how you use the tools. Not using the best AI available in 2026 is increasingly a risky move — not because a model will replace you directly, but because a competitor who uses these tools effectively will outproduce you significantly.

One of the more striking observations from GPT 5.4's capabilities is what it means for non-developers. If AI can handle 98% of the coding required for world-class software, skeptics might argue the remaining 2% still justifies full-time developer employment. But the other consequence is that non-developers can now perform at a level almost indistinguishable from the very best. The lines between professions are blurring rapidly. A marketer can build dashboards. A strategist can ship functional prototypes. The professional advantage increasingly goes to whoever adapts fastest.

Live demo: GPT 5.4 Codex building an animated league table for Stockport County FC from scratch 04:10 Live demo: GPT 5.4 Codex building an animated league table for Stockport County FC from scratch Watch at 04:10 →

What Is the GDP-Val Benchmark and Why Does It Matter?

GDP-Val is OpenAI's attempt to move beyond the usual abstract benchmarks and measure AI against actual professional work. The 44 occupations were chosen specifically because of their economic weight — the jobs that collectively drive the most GDP output. Tasks were drawn from real work performed in those roles, evaluated blind by subject-matter experts.

The benchmark is internally designed, which is worth noting. OpenAI grading OpenAI is a familiar limitation. But the methodology — blind grading, domain experts, real occupational tasks — is meaningfully more grounded than most model-generated leaderboards. The result that deserves more scrutiny, however, is the one that barely got discussed: GPT 5.4 Pro, available only to the highest-paying subscribers, actually scored worse than standard GPT 5.4 on this benchmark. We will come back to that.

Does GPT 5.4 Hallucinate More Than Other AI Models?

This is where the hype deserves a genuine check. According to benchmarking from Artificial Analysis, GPT 5.4 performs well on overall accuracy — close to state-of-the-art, though not quite matching GPT 5.3 Codex. But when it gets things wrong, it is significantly more likely to confabulate a confident-sounding wrong answer rather than admit uncertainty.

The hallucination benchmark showing GPT 5.4 at 89% confabulation rate when it gets answers wrong 07:55 The hallucination benchmark showing GPT 5.4 at 89% confabulation rate when it gets answers wrong Watch at 07:55 →

On the hallucination chart, GPT 5.4 sits at 89% — meaning nearly nine times out of ten when it is wrong, it will manufacture an answer rather than flag its own ignorance. You want to be low on that chart. GPT 5.4 is not. This matters enormously for professional use cases where a plausible-sounding wrong answer is often worse than no answer at all.

It is also worth noting — with some irony — that this comes almost three years after OpenAI's CEO publicly predicted that hallucinations would no longer be a meaningful topic of conversation by this point in time.

What Can OpenAI Codex Actually Do Now?

Setting aside the caveats for a moment, the progress in near-autonomous software development is genuinely breathtaking. OpenAI's Codex — now available on both Mac and Windows — can one-shot surprisingly complex tasks. In a live demonstration, it was asked to produce an animated league table tracking a football club's position across an entire season. It did so, incorporating live web searches, generating accurate data, and producing an interactive visualisation that would have taken a developer hours to build manually.

GPT 5.4 incorporates the coding capabilities of GPT 5.3 Codex while improving how the model operates across tools and software environments. The practical implication: the loop between writing code, testing it, seeing the output, and correcting mistakes is almost closed. The model can now see its own outputs with increasing accuracy and iterate on them. A Viking raid timeline built in the same session went from missing graphics and mislabelled locations on first pass to a visually polished, geographically accurate interactive map after one round of self-correction.

The Viking incursion interactive map — before and after GPT 5.4 self-corrected its own output 14:30 The Viking incursion interactive map — before and after GPT 5.4 self-corrected its own output Watch at 14:30 →

That loop closing is the development that matters most for the medium term. When a model can write, deploy, test, and fix its own software reliably and repeatedly, the economics of software development change permanently.

Why Does GPT 5.4 Pro Score Worse Than GPT 5.4?

This was the quiet narrative violation in the GPT 5.4 launch. The Pro tier — the most expensive, highest-access version of the model — underperformed standard GPT 5.4 on the GDP-Val benchmark. This is not unprecedented. Certain Pro models from OpenAI have historically underperformed significantly cheaper alternatives from other providers on specific benchmarks.

What makes this interesting is that the same Pro model scored the best of any OpenAI model ever on an independent private benchmark designed to probe for common-sense reasoning and trick-question resilience. So the performance profile is genuinely uneven — stronger in some domains, weaker in others, and not simply correlated with price tier. The broader lesson is that benchmark performance is spiky and domain-dependent. No single score tells you everything about how a model will perform on your specific work.

Why Did Anthropic Lose the Pentagon Contract to OpenAI?

In arguably the most consequential story running alongside the GPT 5.4 launch, Anthropic was officially designated a supply chain risk by the US Department of Defense and effectively removed from a major government AI contract. OpenAI stepped in and secured the deal.

A leaked internal memo from Anthropic CEO Dario Amodei revealed the company's account of what happened. Anthropic declined the contract because the Pentagon wanted to retain the ability to use Claude models for applications that Anthropic said could include domestic surveillance and fully autonomous warfare. Those were described as hard red lines. OpenAI, according to Amodei, accepted the contract with a safety layer provided by Palantir — a classifier system that flags and denies certain applications.

Amodei described this arrangement as 80% safety theater — a system designed more to placate employees uncomfortable with military applications than to meaningfully prevent misuse. He alleged that Palantir explicitly framed the service as one that makes uncomfortable realities invisible to staff. OpenAI disputed this characterisation. Sam Altman stated that the military's operational decisions with its models are ultimately a matter for the government, not for OpenAI to adjudicate.

The story has a significant postscript. The Washington Post reported that Claude — deployed inside a Palantir system — had already been used to suggest hundreds of targets in Iran, provide precise location coordinates, and prioritise those targets by importance. This may not technically violate Anthropic's usage terms given the Department of Defense's six-month transition window. But it complicates any clean narrative about who the responsible actors in this space are.

Amodei later apologised for the tone of the leaked memo, describing it as not reflecting his careful or considered views. Both companies are navigating genuinely difficult territory. There are no obvious heroes here — only organisations making consequential tradeoffs in real time, while simultaneously accelerating toward revenue figures that would have seemed like science fiction three years ago.

What Should Professionals Actually Do With All of This?

The practical answer is uncomfortable in its simplicity: use the best tools available, use them regularly, and test them critically. Platforms like Gemini 2.0 Pro, GPT 5.4, Claude 4.6 Opus, and a range of other models including Chinese alternatives offer meaningfully different performance profiles across different task types. The professionals who will navigate this period best are those who develop fluency across multiple models rather than loyalty to one.

Test outputs. Verify facts. Do not trust a confident-sounding answer on a topic where the stakes are high. And pay attention — not just to the benchmarks, but to the decisions being made about how these models are deployed, by whom, and for what purpose. Because increasingly, that is not an abstract concern. It is a professional and civic one.