ARC-AGI-3 Leaderboard: Claude Opus 5 Leads at 30.2%, GPT-5.6 Sol Top on ARC-AGI-2
What happened
ARC Prize published results on the new ARC-AGI-3 benchmark, which tests AI agents' ability to adapt on the fly to novel interactive environments—a qualitatively new type of evaluation.
Context and impact
ARC-AGI is considered one of the most credible measures of progress toward AGI. Low scores (30%) even for the best model indicate that adaptive intelligence remains a fundamental challenge for frontier models.
Details
- ARC-AGI-3: Claude Opus 5 = 30.2% (1st of 4 evaluated models)
- ARC-AGI-2: GPT-5.5 = 85% (leader); Gemini 3.1 Pro = 77.1% (cheapest within 10% of leader)
- ARC-AGI-1: GPT-5.5 = 95% (leader)
- GPT-5.6 Sol at max reasoning: 13.33% Public, 7.78% Semi-Private on ARC-AGI-2
- ARC-AGI-3 tests 'on-the-fly adaptation in novel interactive environments', not static tasks
- 4 models evaluated on ARC-AGI-3 as of publication
Open original source
ARC Prize