3 ms·
I think it should be expected that a smaller model is worse at reciting memorized data than a big one. I also don't think it's a good use case of small local mo
by sorenjan 8d ago
I think it should be expected that a smaller model is worse at reciting memorized data than a big one. I also don't think it's a good use case of small local models. Can it find and recite Jabberwocky if given access to a web search tool?
- 0xbadcafebee 8d agoForgetting things isn't lossless though is it? Makes the benchmark and the finding quite suspect
- sorenjan 7d agoIt depends on what you mean by lossless. If both models can perform the same tasks it can be considered lossless for those tasks. That task might be more related to language understanding rather than memorizing, they have several benchmarks in the article.