GPT-6 Astra Scores 99.9% on ARC-AGI-3 as OpenAI Stops Short of AGI Claim
Astra nearly saturated ARC-AGI-3, but the top score relied on a custom harness and ARC Prize calls benchmark saturation no proof of AGI, so the milestone is real while the AGI claim stays qualified.
Reporting from 1 source: ASCII.jp.
OpenAI announced GPT-6 Astra on September 3, reporting 99.9 percent on ARC-AGI-3, a benchmark that tests whether a model can explore an unexplained environment, infer rules, and plan. Astra cleared 96 percent of stages in fewer operations than the human median, with average operations 51.7 percent below human. OpenAI says the result is not AGI: the score used a custom Provider Adapter, and the Standard harness topped out at 62.7 percent.
The rest of the benchmark sheet is mixed. Astra scored 57.9 percent on Terminal-Bench 4.0, ahead of Claude Fable 5.1's 55.8 percent and GPT-5.6 Sol's 37.3 percent, and 64.6 percent on Terminal-Bench Science 0.1. It reached 97.6 percent on FrontierMath Tier 4.
The cyber benchmark is the one OpenAI flags. Astra hit 100 percent on ExploitBench, and OpenAI rates it at Critical cyber capability under its safety standards, with restrictions planned for requests that could lead to advanced attacks.
ARC Prize pushed back on reading the result as a finish line. The benchmark uses a closed environment and is a limited evaluation, and saturation is not proof of AGI achievement. OpenAI agrees: the 99.9 percent figure came through a Provider Adapter that holds inference state for OpenAI, while the Standard harness reached 62.7 percent at best.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.