Abstract reasoning

ARC-AGI open model results

Abstract reasoning benchmark built from novel visual transformation tasks. Results depend heavily on the dataset split and solver system, so each row preserves its evaluation scope.

The ARC Prize community page distinguishes organizer-verified semi-private scores from self-reported public-set scores. The rows below retain that status and must not be compared across different splits as if they were one leaderboard.
SystemModelScoreEvaluation scopeEvidence
Tiny Recursive Model (TRM)Result reported by the open-source implementation authors TRM, 7M parametersOpen-source system 45% ARC-AGI-1 evaluation set Community curated
Tiny Recursive Model (TRM)Recursive reasoning model trained from scratch TRM, 7M parametersOpen-source system 7.8% ARC-AGI-2 Public Train Community curated
Hierarchical Reasoning Model (HRM)Iterative hierarchical recurrent reasoning model HRM, 27M parametersOpen-source system 2% ARC-AGI-2 Semi-Private Organizer verified

Primary sources

Official benchmark Original paper Evaluation repository