7 ms·
This is amazing, thank you!! > My hypothesis is that the Alpaca dataset does not contain any related arithmetic tasks, and the model actively unlearns basic ar
by munro 3y ago
This is amazing, thank you!!
> My hypothesis is that the Alpaca dataset does not contain any related arithmetic tasks, and the model actively unlearns basic arithmetic when it focuses more on other tasks.
I'm surprised this wasn't verified, it's a major benchmark stat. My eyes keep getting drawn to it, because it seems to have the most variance. Does anyone know?
Also throwing it out, I would love to see a Neptune/Wnb of the hyperparameter tuning :)
- pbhjpbhj 3y ago>the model actively unlearns Aka "catastrophic forgetting" (CF).
- moffkalast 3y agoThere's also been a recent discovery that the MMLU benchmark contains a significant percentage of answers with completely wrong ground truth, where models answering correctly lots points. The most popular benchmarks and datasets are mostly haphazardly cobbled together with hardly any oversight or verification, sometimes even synthetically generated from gpt 3.5 et al. without checking the output at all. Frankly it's amazing that any of it even works when people blindly train and test with what's essentially self contradicting garbage.
- munro 3y agoOh wow, yea I see. A web search brings up a lot of examples. >> As a result of an accident, Abdul lost sight in his right eye. To judge the distance of vehicles when he is driving, Abdul is able to rely on cues of >> - A. I only >> - B. II only >> - C. III only >> - D. I and II only > You didn’t read that wrong. The question never explains what I, II or III are. This appears to have been improperly copied from crackap.com. Somehow Platypus 2 still gets the right answer with high confidence. Is this a sign it has merely memorized the answers? I checked the second best ranked model upstage/LLama-2–70b-instruct-v2 and it also somehow got the answer right (the third best Open LLM also gets this question right so I don’t know what is happening). https://derenrich.medium.com/errors-in-the-mmlu-the-deep-learning-benchmark-is-wrong-surprisingly-often-7258bb045859 https://derenrich.medium.com/errors-in-the-mmlu-the-deep-lea...
- kristianp 3y agoWe might get a hint that a model has beem trained on these datasets if the model gets these questionable questions "correct".
- moconnor 3y agoThat Alpaca has no arithmetic tasks? Just look at the data, it’s text...