The ARC Prize Foundation has established an internal game studio to develop the ARC-AGI-3 benchmark, a collection of 135 unique puzzle games designed to test the limits of current artificial intelligence. The primary thesis of this initiative is that existing AI models, despite their proficiency in narrow domains like coding, lack true general intelligence. By requiring models to solve novel, interactive puzzles on their first attempt without prior exposure, the foundation aims to provide a rigorous, falsifiable metric for Artificial General Intelligence (AGI) that cannot be bypassed through memorization or training on existing datasets.
Key findings indicate a significant performance gap between humans and machines. While humans consistently achieve high scores on the collection, current top-tier AI models have failed to solve even a single level of the ARC-AGI-3 set, scoring below 1%. The benchmark utilizes a tiered system of 25 public, 55 semi-private, and 55 private games to prevent data leakage. The foundation maintains a strict "no training wheels" policy, refusing to allow models to use specialized prompts or human-assisted harnesses, as the goal is to measure inherent model intelligence rather than optimized performance.
The development of this benchmark involved a dedicated team of eleven game engineers who transitioned from Unity to a custom Python-based engine to streamline production. The project emphasizes interactive environments over static testing, arguing that real-world intelligence requires long-horizon planning and world modeling. As of early 2026, the foundation asserts that the continued failure of frontier models to solve these puzzles serves as empirical evidence that AGI has not yet been achieved, suggesting that a fundamental shift in AI architecture—beyond current transformer-based models—may be required to bridge this intelligence gap.