Skip to main content
Mythos

BenchmarkList is a web-based aggregator of AI model benchmark scores and evaluation leaderboards, collecting published results into a single comparable view with source links back to each original result.

Its coverage spans the categories models are commonly measured on — coding, reasoning, agentic tasks, creative work, games, safety, and multimodal performance — allowing scores to be compared across models and across evals rather than read one leaderboard at a time. The stated emphasis on source-linked provenance addresses a recurring problem in the space: benchmark numbers circulate widely detached from the conditions that produced them, and a score without its harness, prompt, and date is difficult to trust or reproduce.

The site occupies the same niche as individual leaderboards such as LiveBench and the evaluation pages maintained by model developers, but positions itself as an index across them rather than another primary evaluator. That makes it most useful as a starting point for orientation, with the underlying source still worth reading before drawing conclusions.

Contexts

Created with 💜 by One Inc | Copyright 2026