NVIDIA Agent AVO Scores 100% on ARC-AGI-3 Where Base Model Manages 30%
The gap between the base model's standalone 30% and the full agent's 100% shows that the execution harness handling memory, tools, and recovery determines long-horizon performance more than the AI model itself.
Reporting from 1 source: GIGAZINE.
NVIDIA's Agentic Variation Operators (AVO) agent system scored 100% on the public set of ARC-AGI-3, a benchmark measuring AI reasoning in unknown game-like environments. The base model 5 used inside AVO scored about 30% in a separate evaluation by ARC Prize, the benchmark's developer. AVO cleared all 183 levels across 25 environment types.
NVIDIA reports that AVO, its agent system built originally for GPU kernel optimization, cleared all 183 levels in the ARC-AGI-3 public set while keeping action efficiency equal to or better than first-time human players. The benchmark's metric, Relative Human Action Efficiency, measures both clearing the game and how few actions the agent takes compared to a human baseline.
ARC Prize, which developed ARC-AGI-3, evaluated the same base model under different conditions and recorded roughly 30% on the same metric. The same agent architecture transferred from GPU code optimization to unknown games without redesign. AVO uses persistent memory of past trials and a supervision function that redirects strategy when progress stalls.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.