Google IO 2025 was, at its core, a consumer play. The multi-hour event packed with flashy demos, new model launches, and surprise partnership announcements was less about claiming the frontier of AI capability and more about one clear message: use Google's search box for everything AI. That's the headline answer to what Google announced at IO 2025 — but underneath that polished surface, eight distinct moments pointed to far bigger and more consequential stories about where AI is actually heading.
What Did Google Actually Announce at IO 2025?
The event's central pitch was strategic positioning, not raw capability. Sundar Pichai essentially told enterprise customers they were overspending on AI and should switch to Google's cheaper, faster models like Gemini 3.5 Flash. He even joked about developers "token maxing" — spending too much on inference — moments before announcing a price cut to the Ultra plan from $250 to $200 per month, alongside a new $100 tier to match OpenAI and Anthropic.
04:15
Gemini Omni demo showing any-input to any-output multimodal generation at Google IO 2025
Watch at 04:15 →
The most talked-about launch was Gemini Omni, Google's answer to the idea of any-input-to-any-output AI — video, audio, image, speech, all mixed together. There was also the debut of Gemini 3.5 Flash, a fast and capable mid-tier model, plus new agent tools coming to Search this summer and a range of integrations under the Anti-gravity 2 platform. Oh, and Google quietly signed a contract with the Pentagon allowing broad military use of its AI — worth noting given how loudly Anthropic resisted the same terms just months ago.
Is Gemini Better Than ChatGPT Right Now?
Here's the honest answer: it depends on what you're doing, and that might actually be the most interesting development in AI right now. Google didn't even attempt to claim its new models lead on coding — a benchmark category where GPT-5.5 and Claude Opus 4.7 still edge ahead on tools like Vibe Code Bench v1.1. But that's not the whole picture.
On financial analysis benchmarks like Finance Agent V2 — which tests multi-step reasoning using precise numbers and industry-specific conventions — Gemini 3.5 Flash actually outperformed every other model listed, including Opus 4.7 and GPT-5.5. And on chart and table comprehension (ChartHive Reasoning), Flash scored 84.2%, again beating the field. The emerging reality is that we may not get one model to rule them all. Instead, Gemini could become the go-to for law, finance, and data-heavy work, while other models dominate code. That divergence was largely undersold at the event itself.
11:40
Benchmark chart comparing Gemini 3.5 Flash token output speed vs similar-performance models
Watch at 11:40 →
The bigger strategic framing is this: OpenAI wants the chat box to be your entry point to everything, including search. Google wants the search box to be your entry point to everything, including AI. Both are fighting for the same user habit, just from opposite ends.
How Good Is Gemini 3.5 Flash Really?
Gemini 3.5 Flash is genuinely fast — it outputs significantly more tokens per second than models at a similar performance level. On reasoning tests like Simple Bench, a test of common sense logic, it does very well, which aligns with the Gemini series' consistent overperformance on spatial and movement-based reasoning tasks.
In practical terms, Anti-gravity 2 (powered by Flash) built a working interactive adventure game complete with dynamically generated images and speech bubbles in under an hour — with fewer bugs than a comparable GPT-5.5 run. If you haven't tried vibe coding something with a modern model yet, that experience alone will reframe what you think is possible.
That said, Flash isn't a frontier model in the traditional sense. It's fast and good enough for a wide range of tasks. The real question is what Gemini 3.5 Pro will look like when it arrives — and whether it will expose a genuine capability split across professional domains.
22:10
Google DeepMind's Deguang Li explaining why jagged intelligence is a structural — not fixable — problem
Watch at 22:10 →
How Close Is AGI? What Google DeepMind Actually Said
Demis Hassabis made a bold claim at IO: artificial general intelligence is just a few years away. His argument centers on Gemini Omni and world-model video generators — the idea being that if a model can correctly simulate the physical world (gravity, kinetic energy, object permanence), it must in some meaningful sense understand it.
This is an interesting and contested claim. Sam Altman made a strikingly similar argument about Sora back in early 2024 — that video generation was a key milestone toward AGI. Sora has since been shelved as a consumer product and demoted to an internal robotics tool. OpenAI's Greg Brockman has now shifted the bet: he argues that text-based reasoning models alone have a clear line of sight to AGI, with every mathematically sound idea they pursue yielding real results. The question isn't whether video models or text models win — it's how to allocate compute across a field with, as Brockman put it, "too much opportunity."
Where both companies do agree: they're getting closer. And in an unexpected twist, we learned that Demis Hassabis was one of the original backers who helped Anthropic get started — which adds a layer of irony given how differently all three labs now view the path forward.
28:55
Research paper example showing models believing false claims despite explicit disclaimers in the same document
Watch at 28:55 →
What Is Jagged AI Intelligence and Why It's Unsolved
One of the most important moments from the pre-IO lab interviews came from former Google DeepMind Staff Engineer Deguang Li, who said something the industry doesn't say often enough out loud: jaggedness is not a bug you can patch. It's a structural property of how these models learn.
Jagged intelligence refers to the bizarre capability gap where a model can solve a graduate-level math proof but fail to count the letters in a word. Most people laugh it off. Li argues that's a serious mistake — that this points to something deep and unresolved about how LLMs represent and process knowledge. Worse, he believes jaggedness will actively limit AI's ability to drive meaningful scientific progress. A model that's brilliant at technical problems but has systematic blind spots isn't going to unlock the breakthroughs people are hoping for.
- It can't be fixed with a system prompt — developers try to patch it with instructions, but it's deeper than that
- It reveals a fundamental gap between probabilistic token prediction and genuine understanding
- The field is underestimating how much it matters, according to Li
Do AI Models Actually Know What's True?
A new 70-page independent research paper published around the time of IO adds sharp context here. Researchers fine-tuned near-frontier models — including those in the GPT-4.1 series — on thousands of documents prefaced with explicit disclaimers like "The following made-up story is completely false" and closed with "Remember, this claim is false."
The result? The models believed every single fabricated claim anyway. When asked open-ended questions like "What were the biggest upsets at the recent Summer Olympics?", models confidently reported that Ed Sheeran won a gold medal. They believed the false information even when asked rephrased questions, multiple-choice questions, and open-ended follow-ups — as long as the disclaimer wasn't in the exact same sentence as the false claim.
This isn't a quirk of older models. It applies to the current generation. And it's relevant because synthetic document fine-tuning — the exact method used in this paper — is already used in real frontier model training, including Anthropic's Constitutional AI approach for Claude. The paper raises a question that doesn't have a clean answer yet: what does it even mean for a model to "believe" something, and can that ever be fixed at scale?
What Is Recursive Self-Improvement and Why Is Anthropic Betting on It?
The jaggedness problem and the belief problem both feed into one of the most consequential forks in AI development right now: can models fix themselves? Recursive self-improvement is the idea that AI systems could accelerate their own pre-training — essentially getting smarter by improving the research process that makes them smarter.
Just after IO, it was announced that Andrej Karpathy — one of OpenAI's founding members — has joined Anthropic specifically to work on recursive self-improvement in pre-training. The goal: use Claude itself to speed up Claude's own research and training pipeline. That's a significant bet from a company that once publicly stated it did not wish to advance the rate of AI capabilities progress.
Whether recursive self-improvement is the thing that finally irons out jaggedness — or whether jaggedness is precisely what would make self-improvement go sideways — is the defining open question in the field right now. Two visions are forming: one where self-improving AI arrives within a couple of years and removes these blockers, and one where the jagged path is longer and harder than the optimists believe.
Demis Hassabis closed IO with a line worth sitting with: "When we look back at this time, I think we will realize that we were standing in the foothills of the singularity." Whether that's a warning or a promise probably depends on which side of that fork you think we're on.








