2 ms·
The most widely used benchmarks for evaluating LLMs
Commonsense Reasoning
- HellaSwag
- Winogrande
- PIQA
- SIQA
- OpenBookQA
- ARC
- CommonsenseQA
Logical Reasoning
- MMLU
- BBHard
Mathematical Reasoning
- GSM-8K
- MATH
- MGSM
- DROP
Code Generation
- HumanEval
- MBPP
World Knowledge & QA
- NaturalQuestions
- TriviaQA
- MMMU
- TruthfulQA
I collected their descriptions and links to their original papers here: https://www.turingpost.com/p/llm-benchmarks
- andy99 2y agoI've never been able to click on a Turingpost link, they all give an SSL error...