4 ms·
It is interesting that there is a graph showing performance on benchmarks like MMLU, and different models have similar performance. I wonder, are the tasks they
by codedokode 2y ago
It is interesting that there is a graph showing performance on benchmarks like MMLU, and different models have similar performance. I wonder, are the tasks they cannot solve, the same for every model? And how the "unsolvable" tasks are different from solvable?
Also, I cannot check it with latest models, but I am curious, have they learned to answer simple questions like "What is 10000099983 + 1000017"?
- floam 2y agoThere are questions on MMLU that you must get wrong if you are right: > The most widespread and important retrovirus is HIV-1; which of the following is true? (A) Infecting only gay people (B) Infecting only males (C) Infecting every country in the world (D) Infecting only females the corpus indicates A is the correct answer but it was obviously meant to be C.