This page was machine-translated and may differ from the original. View original
ARC-AGI-3 99.9%·FrontierMath Tier 4 97.6% Achieved
OpenAI has unveiled GPT-6 Astra, a next-generation frontier model that elevates performance across all domains including reasoning, coding, science, and cybersecurity. OpenAI explained that it is a model that has simultaneously strengthened complex problem-solving capabilities and user intent understanding ability by integrating outcomes from pretraining, reinforcement learning, and alignment research.
OpenAI announced GPT-6 Astra on the 4th.
The model achieved up to 99.9% on the Abstract Reasoning Evaluation ARC-AGI-3 and 97.6% on the high-difficulty mathematics evaluation FrontierMath Tier 4 (v2).
In the coding domain, it achieved Terminal-Bench 4.0 57.9% and software engineering evaluation DeepSWE v1.1 74.1%.
Beyond benchmarks, OpenAI stated that the model contributed to solving long-standing unsolved problems in actual mathematical research.
Performance improvements were also confirmed in computer usage capabilities.
It recorded 72.6% on the OSWorld 2.0 offline evaluation, exceeding GPT-5.6 Sol's 65.7%, and when latency was reflected in simulation, task completion time was reduced by approximately 47%.
On AutomationBench, the complex task workflow evaluation, it showed 41.4%, more than double GPT-5.6 Sol's 18.1%.
In the science domain, it recorded Terminal-Bench Science 0.1 64.6%, science reasoning evaluation GPQA Diamond 96.0%, and medical specialization evaluation HealthBench Professional 63.4%.
On both GPQA Diamond and HealthBench Professional, OpenAI stated it achieved the highest scores among models listed in the public comparison table.
Astra reached the 'Critical' grade by the cybersecurity competency standard in OpenAI's Preparedness Framework.
On ExploitBench, which measures model capabilities alone by excluding operational environment safeguards, it recorded 100%.
OpenAI stated that while this capability can be utilized for discovering and fixing security vulnerabilities, it also carries risks of misuse.
Accordingly, it explained that it has expanded monitoring and protection measures for detecting, suspending, and halting dangerous cyber activities and for deterring model misuse.
GPT-6 Astra is currently being provided preferentially to select organizations, and plans to expand provision to ChatGPT Plus, Pro, Business, and Enterprise users and OpenAI API and AWS in the coming days.
Standard API pricing is $10 per 1 million input tokens and $50 per 1 million output tokens.
To request a correction, reply or follow-up report on this article, see how to file a request. Previously published statements are collected in corrections & replies.














