Open model benchmark leaderboards
Task-specific results for open-weight models and reproducible open-source systems. Every score links to the original paper, official leaderboard, or first-party result file.
ARC-AGI
Abstract reasoning benchmark built from novel visual transformation tasks. Results depend heavily on the dataset split and solver system, so each row preserves its evaluation scope.
- Metric
- Solved tasks
- Rows
- 3 sourced results
CRAG
Comprehensive RAG benchmark for factual question answering across dynamic facts, long-tail entities, web pages, and mock knowledge-graph APIs.
- Metric
- Answer accuracy
- Rows
- 4 sourced results
LongMemEval
Long-term conversational memory benchmark covering information extraction, multi-session reasoning, knowledge updates, temporal reasoning, and abstention.
- Metric
- End-to-end QA accuracy at top-10 retrieval
- Rows
- 2 sourced results
OSWorld-Verified
Real-computer benchmark for multimodal agents completing open-ended tasks across desktop and web applications.
- Metric
- Task success rate
- Rows
- 4 sourced results
How scores enter this database
We preserve the benchmark version, evaluation scope, system setup, openness, and verification status for every row. A result is omitted when its exact score or provenance cannot be confirmed from a primary source.