3 ms·
Another thing that sort of puzzles me about benchmarks is that LLMs are not deterministic and do not always complete a problem. So what are the results actually
by tudelo 2mo ago
Another thing that sort of puzzles me about benchmarks is that LLMs are not deterministic and do not always complete a problem. So what are the results actually representing? The best run? The average? It is all in some ways a falsehood
- deepsquirrelnet 2mo agoThat's a good point and conventionally if benchmarks aren't run as "one shot", it is denoted as "benchmark@K". Inference time scaling has historically shown improvement. Generally though, many of these fairness complaints do go away if there is "3rd party testing". Right now, companies reporting their own benchmarks has all the problems that 3rd party testing resolves in many other industries.