Announcing ARC-AGI-3
The only unsaturated agentic intelligence benchmark in the world
Humans score 100%, AI <1%
This human-AI gap demonstrates we do not yet have AGI
Most benchmarks test what models already know, ARC-AGI-3 tests how they learn
GPT-6.1 Sol from @OpenAI on ARC-AGI (Verified):
- ARC-AGI-3: 52.7%, $7.6K (standard harness), 96.4%, $4.4K (provider adapter harness)
- ARC-AGI-2: 94.2%, $0.25/task
- ARC-AGI-1: 98.5%, $0.06/task
Its 96.4% on v3 was comparable to GPT-6 Astra's 99.9% but at a 77% lower cost.
Over the course of 19 months, the cost to reach 75% on ARC-AGI-1 fell 99.95%: from o3-preview at $26/task to DeepSeek V4 Flash at $0.01/task.
On ARC-AGI-2, it fell 99.44% in just 6 months: from Gemini 3 Deep Think at $13.62/task to Dots3-Note Preview at $0.08/task.
GLM 5.3 Flash from @Zai_org on ARC-AGI (Verified):
- ARC-AGI-2: 65.8%, $0.09/task
- ARC-AGI-1: 91.0%, $0.04/task
GLM 5.3 Flash performs strongly on both benchmarks for under $0.10 per task.