5 min read

Benchmarks Are Dead

After GPT-6 Astra, We’re Entering the Age of Unscorable AI

Hello, AI Enthusiasts!

OpenAI made it official this week: welcome to the Age of AGI.

GPT-6 Astra is out, and OpenAI is selling it as a generational jump — big leaps in computer use, coding, research, and the kind of long, messy work that used to need a human babysitting it the whole way.

I'll take the marketing at face value for a second, because it surfaces a question the industry has been quietly dodging: how do we actually measure this thing?

For a few years now, AI has run on benchmarks. Every launch ships with a table proving the new model is 3% better here, 7% better there, and mysteriously 12% smarter according to a benchmark that didn't exist last spring. We rank them, color-code them, stack them into leaderboards, and crown a winner. It worked well enough when models were basically answering questions. It stops working the moment models start doing things.

Here's the core issue: a benchmark is just a question bank. You hand the model a problem, it produces an answer, you check the answer against a key. Even the fancy computer-use benchmarks are, underneath, a pile of pre-written tasks running inside pre-built environments. That's genuinely useful. It's also nothing like what happens when you drop an agent onto a real machine and tell it to get something done.

Real life doesn’t ship with a benchmark file. Nothing on your desktop tells you exactly what “done” looks like.

A real computer is a browser that crashes mid-task, a site that quietly redesigned its UI yesterday, a spreadsheet with three contradictory versions, a permission nobody mentioned, a PDF buried six folders deep, an API that times out on the worst possible request, and a human who says "can you just take care of this?" and then walks away without defining "this."

The agent has to figure it out. And the part nobody benchmarks well: it has to recover when the first attempt fails. That is a completely different skill from getting the right answer.

Which is why the reports that OpenAI bought tens of thousands of consumer Macs for reinforcement learning and computer-use training are more telling than they look. The exact numbers have been reported, not audited — but the direction is unmistakable. The frontier labs want their models on real machines, not trapped in a chat box. And that quietly breaks the whole benchmark game.

Picture two agents, same computer, same vague brief:

Research this market, pick three promising companies, build a financial comparison, and prep a deck for the investment committee.

There isn't one right way to do that. There are a hundred. One agent browses for two hours. Another writes a script. A third stumbles onto a better data source halfway through and throws out its original plan. A fourth makes a mistake, catches it, and fixes it. A fifth hands you a gorgeous deck built on completely wrong numbers.

So which one won? The benchmark author can't pre-write the answer, because there isn't one.

And this is where it gets genuinely uncomfortable. The hardest part of grading an agent may no longer be was the answer correct. It's was any of this actually useful. That sounds like hair-splitting. It isn't.

Take an agent that clears 95% of a computer-use benchmark, then drop it into a real company for a week. It builds the wrong folder structure, overwrites a live spreadsheet, burns six hours chasing a dead lead, confidently cites a stat from 2023, and gets stuck on a permissions wall it never thinks to ask about. Every benchmark task: perfect. Nobody would keep it employed for a second week.

Now take an agent that only scores 85% — but knows when it's confused, asks the one question that unblocks it, changes tack when something breaks, and quietly ships work you can actually use. Which one is smarter? The leaderboard has no idea. Honestly, neither do we.

This is why I think independent benchmarks are running into a wall. Not because the researchers are slacking — they're building harder, more realistic tests than ever. The problem is that reality keeps getting bigger than the test. You can make the benchmark longer, add more tools, more sites, more apps, more steps, more randomness. It doesn't matter. You're still building a game with rules, and the moment a model learns to play the game, you're back to square one.

So we'll get the scores. 91.4%. 94.7%. 97.2%. 98.1%. People will argue whether Model A beats Model B. Someone will launch another leaderboard. Someone else will launch a leaderboard of the leaderboards. And the whole time, the model will be sitting on a real computer doing things none of those numbers can describe.

That's the benchmark paradox: the closer AI gets to general-purpose intelligence, the less a bank of pre-written questions tells you about it.

So what replaces it? I don't think the answer is "a better benchmark." I think it's something different in kind. Stop handing the agent a question and hand it an objective. Stop giving it a tidy sandbox and give it a messy one. Stop scoring whether it followed the expected steps and score whether it got the outcome. Stop re-running the same public test and throw it into environments that keep changing, on private tasks it has never seen.

And stop asking only whether it succeeded. Ask how reliably it succeeds, what it cost, how long it took, how often it needed a human, how safely it behaved, and whether the work survived contact with reality.

Put simply: the next generation of AI evaluation should look less like an exam and more like an internship.

Don't ask the model:

Did it pass the test?

Give it a laptop, an account, a budget, a goal, and a week. Then watch what it does. That would tell you far more about its intelligence — and it would be a nightmare to standardize. Which is exactly the point.

Real-world intelligence is messy, contextual, adaptive, and open-ended. The second you flatten it into a clean list of questions with fixed answers, you've stopped measuring intelligence. You're measuring test-taking.

And that's why Astra lands differently for me. OpenAI is telling us, out loud, that we've entered the Age of AGI. Fine. But if that's true, we should probably stop handing AGI the standardized exam we designed for chatbots.

The interesting question is no longer "what did it score?" It's "what happens when you give it a computer and tell it to get something done?" And I suspect we're about to find out we don't have a good answer yet.

So — welcome to the Age of AGI. And, quietly, at the same time: welcome to the Age of Unscorable AI. The benchmark didn't get worse. The thing we're trying to measure just walked out of it.


READ MORE

Let the Future Come to Your Inbox

Stay ahead without drowning in information. We turn the most important signals across AI, tech, marketing, and future products into 5-minute reads you can actually finish.


TOGETHER WITH US

AI Secret Media Group is the world’s #1 AI & Tech Newsletter Group, reaching over 2 million leaders across the global innovation ecosystem, from OpenAI, Anthropic, Google, and Microsoft to top AI labs, VCs, and fast-growing startups.

We've helped promote over 500 Tech Brands. Will yours be the next?

Email our co-founder Mark directly at mark@aisecret.us if the button fails.