3 ms·
I agree for the most part, but there is also the other side which is how useful it can be. If you fine tune your model on anime girls, it becomes a very useful
by 22c 3y ago
I agree for the most part, but there is also the other side which is how useful it can be. If you fine tune your model on anime girls, it becomes a very useful model for generating anime girls and perhaps it's now worse at drawing stick figures or cave paintings, but a lot of people don't need the model to draw stick figures or cave paintings and if they did, they could get a model fine tuned to do that.
On the LLM side, there's a similar phenomenon where a huge amount of training could be on SEO spam tags, or /r/counting or producing nonsense or generating obfuscated C or reciting excerpts of Shakespeare word for word, etc.
If what you're measuring is generally more useful then you can end up with a better model using these methods, likely at the cost worse performance on things that you aren't measuring.
I would gladly take a 10b parameter coding assistant that doesn't know how to write in iambic pentameter or recite digits of pi, or translate words from Swahili to Turkish etc. but is much better at code completion.
- thot_experiment 3y agoTotally valid points! However I think a lot of the real "magic" comes from the sort of cross-domain "thinking" that the larger models are capable of and I think that axis hard to benchmark. (though because it seems so hard to quantify, there's a high chance i'm making shit up)
- 22c 3y agoValid, but how much can you distill a model while retaining (or refining) usefulness? You don't need to know how to paint Rembrandt to draw an anime girl, and you don't need to know about the biological taxa of North American Ducks to write a for loop. The cross-over might come to a point where you might need to explain what "duck typing" is, but even then, you only need to know that a duck is something which quacks (What does "quack" mean? Who knows..) If a model forgets how to speak French but gets much better at generating unit tests, that might be perfectly fine for the type of work we want the model to perform. The problem is we can't easily know what the model "forgets" when it gets better at doing something else. The best thing we can do is benchmark/measure their output and hope that those benchmarks cover what users care about. I suspect high quality benchmarks will quickly become almost as important as the tuning process itself.