The Benchmark Lottery
“Which benchmarks you pick can decide who wins.”
- Last reviewed
- 15 Sep 2026
- Source types
- 3 peer-reviewed or classic, 1 preprint
- Version
- v1.0 5 Oct 2026
In plain terms
Which benchmarks get reported can decide who looks best. Pick the suite before you see results, and read the whole table.
Takeaways #
- Model rankings often change depending on which benchmarks are included.
- A method can look like a breakthrough on some tasks and ordinary on others.
- Choosing benchmarks after seeing results is a quiet form of cherry-picking.
- Pick benchmarks that match your use case, and pick them before you run anything.
What it means #
There are thousands of benchmarks. Any given paper or product announcement reports a handful. If the handful was chosen after seeing results, the table tells you more about the selection than about the system. Even without bad intent, communities settle on benchmarks that happen to favor certain approaches, and newcomers get judged by rules that weren’t written for them.
The evidence #
reviewed 15 Sep 2026Named and measured. Dehghani et al. (2021) showed that the relative ranking of methods can change substantially depending on the benchmark tasks chosen, and called this the benchmark lottery.
Models weren’t even compared on the same tests. Liang et al. (2023) found that before HELM, language models had on average been evaluated on just 17.9% of HELM’s core scenarios, with some prominent models sharing no scenarios at all. HELM raised that to 96%.
Whose utility? Ethayarajh and Jurafsky (2020) show that a leaderboard’s ranking reflects one implicit set of priorities that may not match any particular user’s.
An interdisciplinary warning. Eriksson et al. (2025) reviewed about 100 studies on benchmark shortcomings and found recurring problems, including data contamination, construct validity issues, misaligned incentives, and gaming of results, all shaped by commercial and competitive pressure.
Use it #
- Choose your benchmark suite before seeing any results, and write it down.
- Weight benchmarks by relevance to your use case, not by popularity.
- When reading a claim, notice which common benchmarks are missing.
- Look at the full results table, not just the highlighted wins.
Questions to ask #
For vendor reviews, model cards, and launch reviews.
Origins #
The term comes from Mostafa Dehghani and colleagues’ 2021 paper, “The Benchmark Lottery.”
Sources #
- [1]Dehghani, M., Tay, Y., Gritsenko, A. A., et al. (2021). The benchmark lottery.arXiv:2107.07002 PreprintOpen ↗ (opens in a new tab)
- [2]Eriksson, M., et al. (2025). Can we trust AI benchmarks? An interdisciplinary review of current issues in AI evaluation.Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (AIES 2025)Open ↗ (opens in a new tab)
- [3]Ethayarajh, K., & Jurafsky, D. (2020). Utility is in the eye of the user: A critique of NLP leaderboards.Proceedings of EMNLP 2020Open ↗ (opens in a new tab)
- [4]Liang, P., et al. (2023). Holistic evaluation of language models.Transactions on Machine Learning Research (TMLR)Open ↗ (opens in a new tab)
Cite this law #
Laws of AI Evaluation. (2026, October 5). The Benchmark Lottery (v1.0). https://josephalfonso.com/laws-of-ai-evaluation/laws/the-benchmark-lottery.html@misc{lai-the-benchmark-lottery,
title = {The Benchmark Lottery},
author = {{Laws of AI Evaluation}},
year = {2026},
month = oct,
note = {Version 1.0},
howpublished = {\url{https://josephalfonso.com/laws-of-ai-evaluation/laws/the-benchmark-lottery.html}}
}https://josephalfonso.com/laws-of-ai-evaluation/laws/the-benchmark-lottery.htmlRevision history #
- v1.05 Oct 2026Published.