OpenAI Astra Hits Plus: Autonomous AI Just Arrived!
GPT-6 Astra’s computer use gains, benchmark shocker, and hidden cost trade-offs explained
7 set 2026 (Aggiornato il 7 set 2026) - Scritto da Christian Tico
This image is part of OpenAI's official brand assets, available from their press kit
Azzera gli attriti: Genera codici QR personalizzati per la tua bio
Difficile portare il pubblico degli eventi dal vivo sui tuoi link digitali? Genera un codice QR personalizzato collegato direttamente al tuo profilo.
GPT-6 Astra Rolls Out to Plus Users: What Its Computer Use, Benchmark Scores, Costs, and Hallucinations Mean
GPT-6 Astra has begun rolling out to Plus users, and it is being positioned as a major step forward in computer use, reasoning, and agentic performance. OpenAI’s launch materials highlight a 99.9% score on ARC-AGI-3 under a provider-adapter harness, along with strong results on computer-use and cybersecurity benchmarks, while also noting that high costs and hallucinations still matter.
What GPT-6 Astra Is
GPT-6 Astra is OpenAI’s latest model release, framed as a stronger generation of intelligence focused on agentic work, tool use, and multi-step tasks. The model is presented as especially capable at operating computers, solving difficult benchmark problems, and handling professional workflows with less friction than earlier systems.
Why the 99.9% Benchmark Claim Matters
One of the biggest headlines around Astra is its 99.9% score on ARC-AGI-3, a benchmark designed to test abstract reasoning and adaptation to unfamiliar problems. That result is especially notable because it was achieved with OpenAI’s provider-adapter harness, while the standard ARC-AGI setup produced a much lower result, showing how much the evaluation method can affect the headline number.
- ARC-AGI-3 with provider-adapter harness: 99.9%
- ARC-AGI-3 with standard harness: 62.7%
- Interpretation: the model is highly capable, but benchmark setup matters a lot
Computer Use Is a Major Strength
GPT-6 Astra’s strongest practical improvements appear to be in computer use and agentic workflows. OpenAI’s reported results show clear gains on tasks that require interacting with software interfaces, navigating screens, and completing longer workflows across applications.
- OSWorld 2.0: 72.6%
- ScreenSpot-Pro: 92.7%
- Agents’ Last Exam: 59.3%
These results suggest Astra is designed less as a chat-only model and more as a system that can actively work through digital tasks on a user’s behalf.
Costs Still Matter
Despite the performance gains, cost remains a concern. Reports around Astra show that the strongest benchmark results can come with expensive runs, especially when using more advanced harnesses and longer evaluation setups. That means the model’s top-end performance may not be equally practical for every use case.
- High-performance benchmark runs can be expensive
- Different harnesses can produce very different cost and score profiles
- Real-world adoption will depend on whether the gains justify the spend
Hallucinations Have Improved, But Not Disappeared
Hallucination reduction is another important part of the Astra story. OpenAI has reported a lower hallucination rate than earlier models, but the issue is not fully solved, which means users still need verification for important outputs.
- Hallucinations are lower than before
- Errors still appear in complex or ambiguous tasks
- Human review remains essential for high-stakes use
What This Means for Plus Users
For Plus users, the rollout signals earlier access to a more capable model for browsing, computer use, and advanced task execution. The biggest upside is better automation and stronger reasoning, while the biggest caveats are cost efficiency, benchmark interpretation, and the fact that hallucinations are still part of the experience.
Conclusion
GPT-6 Astra looks like a meaningful leap in computer use and benchmark performance, especially for users who want an AI that can do more than generate text. Its 99.9% ARC-AGI-3 result, strong automation scores, and reduced hallucinations make it impressive, but the fine print on cost and evaluation method shows that the hype should be read carefully.
The real story is not that GPT-6 Astra is better at using computers; it is that AI is crossing from content generation into operational labor, where benchmark wins matter less than whether the model can reliably reduce human oversight at scale. If the headline scores only hold inside expensive harnesses, the competitive advantage may belong less to the smartest model and more to the one that is cheapest to trust repeatedly.
Does GPT-6 Astra still hallucinate?
