OpenAI:

We’re introducing GPT‑6 Astra, the world’s most intelligent and aligned model.

GPT‑6 Astra brings together years of research and big bets across pre-training, reinforcement learning, and alignment. Astra is state-of-the-art on computer use, browsing, software engineering, cybersecurity, science, and professional work. Astra saturates FrontierMath Tier 4 with a 98% score, having already helped solve long-standing open problems in mathematics. Astra also saturates ARC-AGI-3 with a 99.9% score and ExploitBench with a 100% score. It also sets a new frontier on computer and browser use, handling the most demanding professional work with unmatched speed, accuracy, and judgment.

Anthropic, earlier in the week:

We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They’re the world’s most advanced models for coding and knowledge work—and their research capabilities offer an early glimpse of how AI models will contribute to scientific progress…

Claude Fable 5.1 sets a new standard for coding, knowledge work, and long-running problem-solving tasks. The charts below show that Fable 5.1 is capable of much higher performance than its predecessor, Fable 5. And when set to Low or Medium effort, Fable 5.1 achieves results similar to or better than Fable 5’s at a much lower cost. (Note that Fable 5.1 defaults to High effort in Claude Code, and to Medium in Claude Cowork and on Claude.ai.)

The charts indeed say Claude Fable 5.1 is the most powerful model for most tasks, though GPT-6 Astra is certainly close. That hasn’t stopped Greg Brockman, OpenAI’s president, from declaring GPT-6 Astra as artificial general intelligence, an arbitrary benchmark that much of the artificial intelligence industry hinges on. GPT-6 Astra and Claude Fable 5.1 are slightly better than their predecessors. They work uninterrupted for longer periods, excel at coding and knowledge work, and hallucinate less. (OpenAI models still tend to hallucinate more than Claude models.) GPT-6 Astra, in particular, is excellent at computer use and using Blender to render scenes. The model’s latter capability has been all over social media, boosted in part by OpenAI employees who want us to believe GPT-6 Astra truly is AGI.

All of this is a pantomime. It’s tough to argue that model capability hasn’t been anemic in recent months, and it’s even tougher to argue that GPT-6 Astra is AGI. I say so because I again think AGI is largely a marketing term. It was intended to describe an AI that performs better than humans at all or most tasks, but what exactly is “better?” Is GPT-6 Astra better than Shakespeare at writing dramas? Can GPT-6 Astra invent Google Search again? I’m not saying that AI is useless because it can’t write like Shakespeare or code like Larry Page and Sergey Brin; I’m trying to prove that a term like “AGI” is an unattainable goal that’s largely irrelevant unless you’re trying to build a computer that will actively put humans out of work. The term “AGI” is only relevant to prove to enterprises that their money is better spent on AI tokens than a new worker. “This computer is better and cheaper than that human worker.

GPT-6 Astra is still not better than human workers in many, many fields. I would say that it’s better at coding than most human programmers, and that coding is one of the few things large language models consistently perform better than humans. But it’s not a better journalist, writer, musician, or artist. Silicon Valley has largely forgotten about these jobs in its “AGI” calculations, but they’re vastly consequential. The world is filled with creative people working creative jobs. Why are AI labs loath to consider their jobs in this definition of AGI? The answer is evident: because enterprises in those fields won’t fire their creative workers for AI — not now and not ever. Creative work doesn’t matter to the people for whom OpenAI and Anthropic constructed this “AGI” term. AGI has never been about model capability.

I’d reckon that we’ll see a lot of this over the next year: AI companies doing everything in their power to make everyone collectively forget about creative people. They’ll show how their models are better at making horrible Blender scenes after 19 hours and thousands of dollars in tokens. They’ll use AI’s immense potential in mathematics and theoretical computer science to downplay concerns that their technology isn’t nearly contributing to society as much as it is taking away from it. They’ll try to impress us with what the models are good at rather than exposing what the models are terrible at. The AI “labs” are becoming less akin to research organizations and more like for-profit oligopolists.