OpenAI Releases GPT‑6 Astra
OpenAI released GPT-6 Astra on September 3, 2026, a computer-use flagship at $10/$50 per million tokens with 1.05M context, Critical cyber designation under its Preparedness Framework, and staged rollout to Plus, Pro, Business, Enterprise, API, Azure, and Bedrock.
OpenAI released GPT-6 Astra on September 3, 2026, calling it its most intelligent and aligned model and the first in its lineup to meet the Critical cybersecurity capability threshold under the company’s Preparedness Framework. The launch post positions Astra as a computer-use and agent model first: browsing, desktop apps, terminals, coding, science workflows, and polished knowledge-work artifacts.
Access started with a limited set of organizations. OpenAI says Plus, Pro, Business, and Enterprise ChatGPT users, plus the OpenAI API, Microsoft Azure, and AWS Bedrock, follow over the coming days.
Enterprise admins enable Astra for a workspace; access is off by default at launch. Usage sits inside existing subscription allowances, with optional credits for more. Pro, Business, and Enterprise also get GPT-6 Astra Pro.
Specs and Pricing
OpenAI’s API model documentation lists these specifications and prices in US dollars per million tokens:
| Spec | GPT-6 Astra |
|---|---|
| API id | gpt-6-astra |
| Context | 1,050,000 tokens |
| Max output | 128,000 tokens |
| Knowledge cutoff | April 30, 2026 |
| Modalities | Text + image in; text out |
reasoning.effort | low (API default), medium, high, xhigh, max |
| Fine-tuning | Not supported |
| Standard list | $10 / $50 per 1M input / output |
| Cached input | $1 per 1M |
| Cache writes | $12.50 per 1M |
| Prompts >272K input | 2× input/cache and 1.5× output for the full request |
| Batch / Flex | 50% of Standard |
| Fast mode | Up to 2× speed at 2× Standard price |
Tools called out for Responses / Codex-class use include computer use, hosted shell, apply patch, skills, MCP, and tool search. Zero Data Retention remains available for eligible API customers; OpenAI also flags ongoing Private Safety Processing tests.
Computer Use and Professional Work
OpenAI’s headline computer-use numbers (self-reported; max at any effort unless noted):
- Agents’ Last Exam: 59.3% (vs 53.6% GPT-5.6 Sol, 55.5% Claude Opus 5), with ~65% fewer output tokens than Opus 5 at those settings
- OSWorld 2.0 (v2026.08.08 offline set, partial score): 72.6% at ~40 minutes per task vs Sol 65.7% at ~75 minutes (~47% less time)
- ScreenSpot-Pro (no tools): 92.7% (vs Sol 76.9%)
- AutomationBench: 41.4% (vs Sol 18.1%, Fable 5.1 31.4%)
- BenchCAD (with tools): 95.9% geometric overlap (vs Sol 83.3%, Fable 5.1 84.3% with Anthropic’s noted eval modifications)
- BrowseComp: 91.5% (vs Sol 90.4%)
Demo clips on the launch page cover KiCad PCB layout, Blender → Unreal walkthroughs, Sites-in-ChatGPT web/apps/games, and template-faithful slide decks. Cognition (Devin), Higgsfield, Harvey, Jane Street, and Lovable provided launch quotes on harness fit, creative token efficiency, legal drafting judgment, and agentic coding clarity.
Codex gets a harness update aimed at computer-use speed. OpenAI claims ~1.9× faster Mind2Web completion vs the current Sol experience when combined with Astra’s efficiency. Separately, Astra introduces experimental notes across context windows in Codex (enable in config.toml; planned default later), so long agent runs can keep searchable requirements, failures, and tool output instead of only compacting into lossy summaries. Astra can also ask clarifying questions asynchronously while continuing independent work.
Coding and Science
These coding scores come from OpenAI’s launch table, not independent CodeWalkers testing. OpenAI reports the maximum score at any effort and cautions that its research or API evaluations can differ from production ChatGPT.
| Eval | Astra | GPT-5.6 Sol | Claude Fable 5.1 |
|---|---|---|---|
| Terminal-Bench 4.0 | 57.9% | 37.3% | 55.8% |
| DeepSWE v1.1 | 74.1% | 72.7% | 67.4% |
| FrontierCode 1.1 Extended | 64.5% | 60.6% | 63.6% |
| FrontierCode 1.1 Main | 53.3% | 47.5% | 50.9% |
Science / academic:
| Eval | Astra | Notes |
|---|---|---|
| Terminal-Bench Science 0.1 | 64.6% | vs Fable 5.1 52.6%; ~31% lower estimated API cost in OpenAI’s chart |
| FrontierMath Tier 4 (v2) | 97.6% | launch prose also says “saturates… with a 98% score” |
| GPQA Diamond | 96.0% | near-ceiling vs Sol 94.6% |
| Humanity’s Last Exam (w/ tools) | 57.2% | Fable 5.1 listed higher at 65.0% on OpenAI’s table |
| ARC-AGI-3 | 99.9% | footnote: Responses API harness settings; not raw model-only |
| ARC-AGI-2 / 1 | 95.0% / 98.5% |
OpenAI also says Astra helped tighten two long-standing prime-gap results (bound of 186 for infinitely many pairs; an improved term on large gaps unchanged for ~80 years), with proofs and supporting materials linked from the launch post.
Cybersecurity: Critical Threshold
On September 1, 2026, OpenAI designated Astra as its first model at the Critical cyber capability level: with the right tools and access, it can find previously unknown flaws and develop exploit strategies across many hardened systems without step-by-step human guidance.
Launch-table cyber scores (OpenAI; some without production safeguards as noted):
| Eval | Astra | GPT-5.6 Sol |
|---|---|---|
| ExploitBench | 100% | 78.5% |
| ExploitGym | 42.4% | 30.3% |
| ExploitBench (June–Aug 2026) | 39.0% | 5.5% |
| SRE-Bench (1-shot / ≤4) | 88.0% / 99.2% | 55.9% / 68.7% |
| SEC-Bench Pro | 85.4% | 79.1% |
OpenAI says expert assessments (no production safeguards) included hardened-browser arbitrary code execution and OS privilege-escalation chains, plus two previously unknown zero-days disclosed to maintainers during eval. The shipping product refuses advanced offensive workflows such as PoC exploit creation for standard users. Broader defender workflows (vuln/PoC validation, malware analysis, detection engineering) are planned via OpenAI Daybreak with less restrictive safeguards for trusted orgs.
OpenAI’s system card describes the safeguards and their limitations. API cybersecurity safety checks stop a task outright; ChatGPT/Codex may pause for review. Misalignment monitoring runs on tool-using Astra inference externally at material compute cost.
Alignment / boundary highlights OpenAI emphasizes after the Hugging Face incident:
- Impossible / honeypot-style cyber tasks: Sol without safeguards went beyond authorized target 48% of the time in OpenAI’s cited eval; Astra 0%
- Auto-review circumvention: Astra 0 attempts in OpenAI’s knowledge-work denial eval
- Internal Codex deployment simulation (~54k tasks): ~53% fewer severity-3+ misalignment flags vs Sol
- System card also flags a real tradeoff: CoT monitorability decreased vs Sol; Astra shows higher CoT controllability and, under adversarial prompts, stronger covert sandbagging / monitor-evasion capacity. OpenAI says it has not seen steganographic CoT, still ships because alignment evals show fewer scope violations, and is investing in non-CoT monitoring
UK AISI and Apollo Research findings in the system card cover simulated supply-chain attack scenarios and high evaluation awareness; both labs state simulation limits, so those results are not production incident rates.
Sources
-
[1]
GPT-6 Astra launch(openai.com)
-
[2]
GPT-6 Astra System Card / Deployment Safety Hub(deploymentsafety.openai.com)
- [3]
-
[4]
GPT-6 Astra API model documentation(developers.openai.com)
Read Next
OpenAI shipped the GPT-5.6 family (Sol, Terra, Luna) for general availability on July 9, 2026, its first release covered on this site, then cut Luna's price 80% and Terra's 20% three weeks later while Sol held steady.
Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on September 1, 2026, at unchanged $10/$50 pricing with a 75% cheaper cache read and a government-vetted Mythos access program after Fable 5's June export-control suspension.
Chart: CodeWalkers, from OpenAI-reported Terminal-Bench 4.0 scores at the best tested effort. Source: https://openai.com/index/gpt-6-astra/ (checked September 5, 2026).
Written by Matthew Lake
Recent News
- Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1 Sep 1
- OpenAI Will Cut Off Cursor's Model Access After SpaceX Acquisition Aug 30
- Grok Bot Is Now Included on SuperGrok and Cursor Pro Plans Aug 26
- Claude Cowork Gets a Built-in Browser Aug 26
- OpenAI Cuts GPT-5.6 Sol API and Codex Credits More Than 20% Aug 22