Is GPT-6 Astra actually AGI? A few days ago I covered its launch and said that question would need more than a hype video and a launch day to settle. The independent numbers have started coming in, and they tell a messier story than OpenAI’s “welcome to the AGI era” line.

What OpenAI actually claims

OpenAI’s president Greg Brockman used that exact phrase, “welcome to the AGI era,” at the Astra launch briefing. He also said he personally believes OpenAI has reached AGI. Coming from the company’s own president rather than an outside commentator, that combination is hard to just wave off.

On paper, the numbers back up the confidence. Astra scores 97.6 percent on FrontierMath Tier 4, a benchmark of unpublished expert level math problems.

Image credit: OpenAi

It hits 100 percent on ExploitBench, a benchmark for finding and chaining security exploits. It is also the first OpenAI model to cross the company’s own “Critical” threshold for cybersecurity capability, meaning it can find unknown vulnerabilities with very little human guidance. Those gains are measurable and verifiable.

The number that got the most attention is 98.6 percent on ARC-AGI-3, and it needs more scrutiny than the rest.

Image credit: OpenAi – ARC-AGI-3 tests how well agents learn as they solve unfamiliar interactive tasks. GPT‑6 Astra saturates the eval, scoring 99.9%. The average human tester scored 48%. GPT‑6 Astra was measured with our responses API harness, which better reflects real-world performance than the original benchmark harness, which discards past reasoning and past messages. With this harness, we estimate Sol would score in the ballpark of ~30%.

What is ARC-AGI-3?

I wrote a full explainer on ARC-AGI and the AGI goalposts back in June, so I will keep the background short here. ARC-AGI (Abstraction and Reasoning Corpus for Artificial General Intelligence) is a test built by researcher François Chollet specifically to resist memorization. It is a benchmark to measure fluid intelligence. Models cannot study for it in advance. They have to figure out new rules from scratch, the way a person would. There’s more to dig in on their page.

At the time I wrote that piece, ARC-AGI-2 was the current version, a set of harder visual puzzles that models solve in one shot. It got beaten fast once labs threw more compute at it. So in March, ARC Prize replaced it entirely with ARC-AGI-3, which ditched static puzzles for hundreds of interactive game environments with no instructions and no stated goals. When it launched, the best model in the world scored under one percent. Humans solved essentially all of it.

Astra supposedly aced that same test at 98.6 percent, but ARC Prize tested it two different ways, and only one of those numbers is the one that got repeated everywhere. Under the standard harness, the same rules every model gets tested under for the public leaderboard, Astra scored 62.7 percent.

Image credig: ARC-AGI

Still the best score anyone has posted, but nowhere near 98.6. The higher number came from a setup that let Astra keep its own working notes across moves within a game, something the other models on that leaderboard were not given. ARC Prize ran both tests and published both numbers themselves. They have also said plainly that even a clean win here is not proof of AGI, since these are still closed, rule bound game worlds, not the messy world we actually live in.

What real testers are saying

Benchmarks are one thing. What people who actually have their hands on Astra are saying is another, and it is more interesting than the benchmark fight.

Theo, a well known developer who has had early access for a few weeks, described it as a major leap in computer use, 3D work, data analysis, and coordinating multiple agents at once. He said it has real quirks and rough edges, but that at times it feels like a taste of AGI. Enthusiasm paired with an honest admission it is not flawless is a useful combination to hear from someone who uses the thing daily.

I have been logging reports like that on trackai, the release tracker I built to pair every model’s claims against how it actually performs in the hands of real users. Two reports came in on Astra within days of launch. One called it world class at Blender and 3D reasoning. Another said early access left them feeling like everyone now has a 3D designer at their fingertips. Those are specific, credible reactions, not just hype for hype’s sake.

The question everyone actually cares about. Are we going to loose jobs now?

None of that benchmark debate is actually what most people are worried about. The real question is jobs.

Sam Altman told CNBC that Astra represents a new capability level and that he expects it to spark “a boom of entrepreneurship.” OpenAI’s own launch ad backed that framing with something concrete: Astra finishing a full presentation, editing a legal document, and booking a reservation, all without a person at the keyboard. That capability, a model that finishes the task rather than just helping with it, is the exact thing driving both the excitement and the fear.

Because that fear is not new. Anthropic’s own CEO warned back in 2025 that AI could wipe out half of entry level white collar jobs within five years. That warning built into a real “jobs apocalypse” narrative through most of 2025 and early 2026. Then it cooled, because the layoff numbers never matched the prediction. By this past May, the share of CEOs expecting AI to meaningfully cut headcount had dropped from around 46 percent to 20 percent. A few outlets even ran pieces about the apocalypse being called off.

Astra’s launch ad brought the fear straight back. Within a day, people online were reacting to the idea that it could replace entire job categories, and outlets were publishing lists of jobs now considered at risk.

Image credit: LadBible

So is this the moment those fears actually start coming true? Partly. Every previous warning was about a model that could draft or suggest things, still needing a human to execute. Astra is built specifically to close that exact gap, operating a browser and finishing tasks start to finish, which puts it a real step closer than anything that came before. But OpenAI did not publish a score on GDPval, its own benchmark built specifically to measure performance on paid work across real occupations, a notable gap for a model being pitched as job ready. Testing so far, including what is landing on trackai, shows strong but uneven results, task by task, not the broad reliable competence you would need to hand someone’s job over completely. The last two years have also shown a consistent gap between a viral launch reaction and layoff data that shows up, if it shows up at all, months later.

So is this AGI or not

My honest read: Astra is the first model where this fear stopped being hypothetical and became technically plausible. That is not the same as proof it has arrived. The capability gap closed in a meaningful way this week. Whether that turns into job losses or the entrepreneurship boom Altman is betting on is a question real deployment will answer over months, not something a launch ad can settle in a day.

AGI itself was invented specifically to move the goalposts every time a system got too good, too fast. Astra is the clearest example of that pattern I have seen since. It is not a settled milestone. It is the goalposts moving again, and we are watching it happen live.

Frequently asked questions

Is GPT-6 Astra actually AGI?
Not by the clearest available evidence. It shows a real jump in capability, but its most striking benchmark score came from a non standard testing setup, and OpenAI has not shown it beats humans at most economically valuable work, which is its own definition of AGI.

What is ARC-AGI-3 and how is it different from ARC-AGI-2?
ARC-AGI-2 was a static puzzle test that got solved faster than expected once labs added more compute. ARC-AGI-3, launched in March 2026, replaced it with interactive game environments that have no instructions, designed to test genuine on the fly reasoning rather than memorization.

Will GPT-6 Astra actually cause job losses?
It is the first model built to carry out full computer based tasks rather than just suggest them, which makes past warnings more plausible than before. Real world deployment, not a launch ad, will determine whether that turns into job losses.


Discover more from The August Dispatch

Subscribe to get the latest posts sent to your email.

Leave a Reply

Trending

Discover more from The August Dispatch

Subscribe now to keep reading and get access to the full archive.

Continue reading