Artificial intelligence
Astra Does Less Work. That Is Both the Product and the Complaint
Two weeks in, OpenAI's Astra is tied for the top of the independent intelligence rankings and gets there on roughly a third of the output tokens of its nearest rival. The same instinct to cut a task short is also what users are complaining about.
MAI
Astra has been generally available for about two weeks, which is long enough for the launch benchmarks and the user complaints to arrive in the same week. Both are pointing at the same trait.
Independent testing from Artificial Analysis puts Astra at 53 on its Intelligence Index — tied for first with Claude Fable 5.1, and six points above GPT-5.6 Sol. What is unusual is not the score. It is that Astra reaches it using roughly a third of the output tokens Fable 5.1 spends: about 27,000 per task at maximum effort against 78,000. That arithmetic is the entire product.
What you are buying
Astra (gpt-6-astra) | |
|---|---|
| General availability | 4 September 2026 |
| Input context | 1.1M tokens |
| Max output | 128K tokens |
| Knowledge cutoff | April 2026 |
| Input | $10 / M tokens ($1 cached) |
| Output | $50 / M tokens |
| Modalities | Text and image in, text out |
| Time to first token (p95) | 8.75s |
The headline price is 2.5× GPT-5.6 Sol, and that number is what most people reacted to first. It is also the wrong number to react to. Because Astra finishes tasks in far fewer tokens, Artificial Analysis measures its cost per task on the Intelligence Index at $3.26 against Fable 5.1's $7.63 — about 40 per cent — and roughly 60 per cent of Fable's cost on the Coding Agent Index. Per token it is expensive. Per finished job it is currently the cheapest thing at the frontier.
Where the lead is real
The gains concentrate in agentic work: long-running tasks where the model drives a terminal, a browser, or a toolchain.
| Benchmark | Astra | Best rival | GPT-5.6 Sol |
|---|---|---|---|
| Terminal-Bench v4.0 | 59% | Claude Fable 5.1, 52% | 40% |
| AutomationBench-AA | 69% | Grok 4.6, 67% | 60% |
| Coding Agent Index (Codex) | 62 | Claude Fable 5.1, 62 | — |
A nineteen-point jump over Sol on Terminal-Bench is not an incremental release. Practitioner reports line up with it in places benchmarks do not reach: unusually broad web research — one developer counted three to five times as many sources gathered as competing models — and a marked step up in 3D and spatial reasoning, with several people singling out Blender and game-engine work.
Where the same instinct costs it
Artificial Analysis also found Astra about 45 Elo points below Sol on GDPval-AA v2, and attributed it directly to turn count: Astra used 24 turns per task where competitors used 45 to 60. Sol still leads on presentation quality.
That is the same behaviour as the token efficiency, seen from the other side. A model that stops early is cheap when the task was finishable and thin when it was not. Users describe exactly that shape — halting mid-task at unintuitive moments, needing a manual nudge to continue. One hands-on review of game-engine work found autonomous runs burned through plan resets without producing better output than supervised ones, and concluded that human steering still mattered more than raw capability.
There are other gaps worth knowing before switching. Astra takes no native video input, so timing and rendering bugs pass unnoticed unless a person watches. On reverse engineering and security work, testers report the specialised GPT-5.6 Cyber still beats it. And several describe it as over-cautious, reading requests legalistically.
The degradation complaints
About a week after launch, reports began that Astra had got worse: lower code quality, faster but thinner answers, more bullet points. Two researchers ran identical prompts against launch-day and current versions and reported worse results from the newer one.
"We don't have AGI. We have a regression." — Pranjal Paliwal, developer
Treat this carefully. Perceived post-launch degradation is a recurring genre — the same complaints followed GPT-5.6 Sol in July — and it is very hard to separate a changed model from changed expectations or from routing between serving configurations. As of publication OpenAI has not addressed it. That silence is the part worth holding against the company; the claim itself is not yet established.
The verdict
Buy it for agentic work: long autonomous coding runs, browser and terminal automation, deep research where breadth of sourcing matters. On cost per completed task it is currently the best deal at the frontier, and the Terminal-Bench margin is wide enough to matter in production.
Do not buy it as a drop-in upgrade for drafting, formatting, or anything where a thorough pass beats a quick one. Sol is better at presentation, cheaper per token, and less inclined to declare a job finished. And if your work is security-specific, the specialised model still wins.
One thing this review deliberately does not settle: Astra ships part of its reasoning in latent space rather than readable text, and it is the first model OpenAI has rated Critical for cyber capability under its own framework. That argument is a separate piece, and we have written it.
Sources: Artificial Analysis — Benchmarking GPT-6 Astra · OpenAI — The path to Astra · OpenRouter — gpt-6-astra · LLM-Stats — GPT-6 Astra · Decrypt — Users say OpenAI's newest model got dumber · MindStudio — GPT-6 Astra hands-on · Hacker News — Initial thoughts on GPT-6 Astra · CNBC — OpenAI begins rolling out Astra