Open model benchmark leaderboards

Task-specific results for open-weight models and reproducible open-source systems. Every score links to the original paper, official leaderboard, or first-party result file.

Model scores and agent-system scores are labeled separately. Results from different benchmark versions, splits, tools, or step budgets should not be treated as directly comparable.

How scores enter this database

We preserve the benchmark version, evaluation scope, system setup, openness, and verification status for every row. A result is omitted when its exact score or provenance cannot be confirmed from a primary source.