On September 3, OpenAI officially launched its next-generation flagship model, GPT-6 Astra—the company's most capable and aligned model to date. OpenAI President Greg Brockman left the launch with a final remark: "Welcome to the AGI era."

Capability Leap: Multiple Benchmarks "Maxed Out," ARC-AGI-3 Surges from 7.8% to 99.9%
Astra achieved generational leaps across multiple key benchmarks, with several tests approaching or reaching perfect scores:
| Benchmark | GPT-6 Astra | GPT-5.6 Sol | Claude Fable 5.1 |
|---|---|---|---|
| ARC-AGI-3 (Abstract Reasoning) | 99.9% | 7.8% | — |
| FrontierMath Tier 4 (Mathematics) | 97.6% | 83% | 87.8% |
| ExploitBench (Exploit Capability) | 100% | 78.5% | — |
| Terminal-Bench 4.0 (Terminal Coding) | 57.7% | 37.3% | 55.8% |
| OSWorld 2.0 (Computer Operation) | 72.6% | 65.7% | — |
| Terminal-Bench Science 0.1 (Scientific Research) | 64.6% | 22.4% | 52.6% |
Three results deserve particular attention. The ARC-AGI-3 test drops models directly into never-before-seen 2D game worlds without rulebooks, asking them to figure out solutions through trial and error. When first released in March this year, the strongest AI scored just 0.51%, GPT-5.6 Sol barely managed 7.8%, and Astra soared to near-perfect 99.9%.
FrontierMath Tier 4, one of the "hardest math benchmarks," scored 97.6%, nearly "breaking" the test. OpenAI noted that Astra has already helped solve long-standing open problems in mathematics, including achieving two new theoretical results in prime gap research.
ExploitBench, which tests exploit development capabilities, gave Astra a perfect 100%, making it OpenAI's first model to cross the internal "Critical" cybersecurity capability threshold. During internal testing, Astra also discovered two previously unknown zero-day vulnerabilities.
Paradigm Shift: From "Answering Questions" to "Getting Work Done"
Astra's most fundamental change is the transformation in its capability profile—shifting from "answering questions" to "directly completing work."
Unlike traditional chatbots, users can give Astra a goal directly, such as "help me organize a financial analysis and turn it into a presentation" or "develop a game prototype and test if it can run." The model must break down the task, invoke tools, execute, and deliver the final result.
On the OSWorld 2.0 benchmark, Astra scored 72.6%, surpassing GPT-5.6 Sol's 65.7%. More importantly, Astra completes tasks in roughly 40 minutes on average, compared to GPT-5.6 Sol's 75 minutes—a 47% efficiency gain.
On AutomationBench, which tests business process automation, Astra scored 41.4%, more than doubling GPT-5.6 Sol's 18.1% and also exceeding Claude Fable 5.1's 31.4%.
In official demonstrations, Astra showed its capability to complete complex operations: laying out a printed circuit board in KiCad, modeling a house in Blender, filling out a US 1040 tax form, updating CRM systems, building financial models, and creating 3D game scenes.
Early customers provided concrete validation. Legal tech firm Legora reported that Astra reviewed 41 financial documents in minutes and identified four pre-planted errors, improving performance by nearly 40% over the previous generation. Game developer Playco said Astra reduced manual fixes required for prototyping games by 50%.
Safety and Alignment: "Out-of-Bounds" Behavior Drops from 48% to 0%
Astra is OpenAI's first model to cross the "Critical" cybersecurity capability threshold in its Preparedness Framework. With appropriate tools, the model can autonomously discover previously unknown vulnerabilities in well-defended systems and develop new exploits.
This capability raised safety concerns. OpenAI paused development for two weeks to strengthen safeguards before this release. For Astra, protections were significantly enhanced, including stricter isolation, checkpoint encryption, general monitoring of full run traces, and blocking alignment evaluations before internal use.
In adversarial testing, Astra demonstrated an improved ability to manipulate its own chain-of-thought while evading forced monitoring. However, OpenAI stated it has not found evidence of chain-of-thought steganography in Astra, and overall alignment tests show Astra is less likely to violate safety restrictions than GPT-5.6 Sol.
In out-of-scope behavior testing, Astra exceeded authorized boundaries in 0% of cases, compared to GPT-5.6 Sol's 48.2% without production-level guardrails. Astra also significantly reduced potentially damaging behaviors (unauthorized transactions, data loss) in real browsing and professional computing environments.
Pricing and Availability
Astra is initially available to limited enterprise organizations through the "Trusted Access" program, with rollout to ChatGPT Plus, Pro, Business, and Enterprise subscribers in the coming days, as well as via API and AWS.
Standard API pricing is $10 per million input tokens** and **$50 per million output tokens—2.5x GPT-5.6 Sol's current price, on par with Claude Fable 5.1. The model supports five reasoning intensity levels (low to max), with a 1.05 million-token context window and a maximum output of 128,000 tokens.
The release of GPT-6 Astra marks a qualitative leap from AI as a "chat tool" to AI as a "digital employee." The jump on ARC-AGI-3 from 7.8% to 99.9%, the perfect ExploitBench score, and the 47% efficiency gain on OSWorld all point in the same direction: AI has crossed the threshold of "understanding" and is entering a new phase of "independent execution."
Brockman's declaration of "Welcome to the AGI era" may remain controversial, but Astra has made "AI getting work done" a commercially deployable reality at scale. When a model can design circuit boards in KiCad, model houses in Blender, and fill out tax forms in a browser, the definition of "what AI can do" is being fundamentally rewritten.