3 ms·
I have a set of independent benchmarks and most also show a difference between reasoning and non-reasoning models: LLM Confabulation (Hallucination): https://g
by zone411 2y ago
I have a set of independent benchmarks and most also show a difference between reasoning and non-reasoning models:
LLM Confabulation (Hallucination): https://github.com/lechmazur/confabulations/ https://github.com/lechmazur/confabulations/
LLM Step Game: https://github.com/lechmazur/step_game https://github.com/lechmazur/step_game
LLM Thematic Generalization Benchmark: https://github.com/lechmazur/generalization https://github.com/lechmazur/generalization
LLM Creative Story-Writing Benchmark: https://github.com/lechmazur/writing https://github.com/lechmazur/writing
Extended NYT Connections LLM Benchmark: https://github.com/lechmazur/nyt-connections/ https://github.com/lechmazur/nyt-connections/
and a couple more that I haven't updated very recently.