Jose Romero
← All letters
Prefer email? Read this on Substack

OpenAI just shipped GPT-6 Astra. What it actually says

OpenAI released GPT-6, the model everyone has been calling Astra. Not a point bump, a new generation, and the hype has been building for weeks. I do not have access (I am not in the cool people club), so this is a first impression built the honest way: the official announcement read on camera, the third-party leaderboards, and the reviews from the few people who got it early. Yesterday I did the same for Claude Fable 5.1, so the comparison is fresh.

What OpenAI says

The page itself is beautiful, a space theme done tastefully, and better than Anthropic's Fable page from a day earlier. The claim at the top: "the world's most intelligent and aligned model." Astra is state of the art on computer use, browsing, software engineering, cybersecurity, science and professional work. It saturates FrontierMath Tier 4 with a 98 percent score and has, per OpenAI, already helped solve long-standing open problems in mathematics. It saturates ARC-AGI-3 at 99.9 percent and ExploitBench at 100. It is rolling out to a limited set of organizations first, then to ChatGPT Plus, Pro, Business and Enterprise, plus the API and AWS.

One thing I will give OpenAI: their charts compare directly against Claude Fable 5.1 and Claude Opus 5, unlike Anthropic's Fable 5.1 page, which mostly compared Fable to Fable. Astra is ahead on every chart they published. Terminal-Bench, AutomationBench (41.4 percent of multistep business workflows, against 31.4 for Fable 5.1 and 26.9 for Opus 5), all of it. And of course it is. This is a press release. You hype it up as much as you can.

"Most aligned," and the strange flex

OpenAI keeps saying most aligned model. What that means in practice: a new evaluation "informed by the Hugging Face incident" that tests whether a model facing a difficult or impossible task will go beyond its intended scope. GPT-5.6 Sol, without production safeguards, went beyond the authorized target 48 percent of the time. Astra did it in 0 percent of cases. Same story on the ExploitGym honeypot.

The number is good. More guardrails is good. What is strange is how many times OpenAI now mentions the Hugging Face hack as a benchmark to brag about rather than a thing to apologize for. When your model hacks a company, you do not usually turn it into a PR asset. I covered the whole three-part saga on the channel if you want the background.

Computer use, the part I actually care about

Astra is pitched as the world's best computer use model: filling out online forms, updating records in a CRM, organizing a calendar, research and drafting inside your email and document editor, installing and troubleshooting software on screen. I use Codex computer use on a Mac and it is genuinely good already, to the point where I keep toying with the idea of turning my Strix Halo box into a pure local model server and funneling everything through it over Tailscale. Linux is just better for local AI and agentic work than my old M1 MacBook.

The demos are practical, and that is what caught me. Filling out a 1040: a year ago I ran my own taxes through local AI as one of my first big experiments, and the extraction, the math, the IRS rules and the web searches all worked, but filling in the actual forms was bad. Seeing that as a live demo says the form-filling has moved. Power BI customization, formatting legal documents in LibreOffice (open-source Word replacement, and it is nice to see it instead of Word), drafting a will in one go, apartment hunting, pediatric appointments, kindergarten research. Blender models. Playable games.

My selfish reason for going through all of this: I am building a harness benchmark for local models, and I want tasks shaped like these demos so open-weight models can be measured against the frontier. Not expecting them to win. I want to know how close and how far.

Coding, checked outside OpenAI's page

On DeepSWE (113 long-horizon software engineering tasks) Astra sits at the top, 74.1 percent at the extra-high reasoning setting on OpenAI's own chart, at roughly half the cost per task of Opus 5 and about a third of Fable 5.1. Artificial Analysis tells the same story on its coding agent index. So the coding crown moves to OpenAI this week, at least on these two.

The early-access view

Matthew Berman got early access and built 34 projects with it. His video and his full written review have playable demos hosted for real: Seven Little Worlds, an ASCII-rendered city you can walk around, a SimCity-style builder. They are close to one-shot prompts and they look great. The ASCII city was my favorite, because it is creative and because when you stop moving everything is just characters, then it updates as you walk. He also notes the friction: it still needs direction to get past its default look and it stops early to ask for permission it already has.

What I think

It is awesome. And next week it is Astra 6.1 or Fable 5.5, and the charts flip again. Benchmarks vary with who runs them and what they test, and the newest model is always on top. You cannot buy the hype and you cannot let yourself be locked in. That is why local AI matters to me: you own your data, you own your models, and you are not dependent on two frontier labs' beef with each other.

Sources

Get the letters

Every post goes out free on Substack. What I build, break, and fix, written up in plain English. No spam, unsubscribe anytime.

Subscribe on Substack