6 ms·
I think maybe its because he didn't have experimental results that show that it worked. Not a knock against the author, there are just so many things that seem
by chessgecko 3y ago
I think maybe its because he didn't have experimental results that show that it worked. Not a knock against the author, there are just so many things that seem like good ideas that don't end up working well in practice, a paper like this without results is hard to value.
- mlsu 3y agoYes, definitely. If he tried to have it published, the lack of experimental results would definitely be a glaring error. But this is still scientific communication. It's really nice that it's legible! > Even though softmax1 is facially quite boring, I’m 99.44% sure that it will resolve the outlier feedback loop that’s making quantization the subject of cascades of research. If you want to run some experiments and prove me right, DM me on Twitter and we’ll get a paper going. I'm guessing that in the stodgy world of science, a communication like this might happen over lunch at a conference, limited to a small clique of researchers who are zealously guarding their next paper. Who could blame them, publish or perish! But someone will probably test this theory out (after my read, it will probably happen in llama.cpp with preliminary results on GPT-2 by next week) and achieve results, and it will happen quickly and legibly to the outside world, because this was published openly and without all of the pretension that formal science (tm) has. If it works, it works. Stuff like this is the soul of the internet. Sharing knowledge and making it legible for all.
- deleted 3y ago[deleted]
- light_hue_1 3y agoThere's a perfectly good venue for this communication: a workshop. Workshop submissions often don't need evidence. They just need a small kernel to spur discussion. Without experiments, there is no hope of publishing this in anything more than a workshop. Nor should there be.
- WithinReason 3y agoThen again, if you don't have access to giant compute clusters you can't test this, so it's either a blog post or nothing. I believe the outlier problem that this solves only appears for very large models.
- janalsncm 3y agoThat isn’t true at all. Train a smaller model on a smaller dataset. You can even train on your laptop. It’s definitely feasible. This is just a proof of concept, it doesn’t need to beat state of the art.
- WithinReason 3y agoMaybe I edited my comment too late.
- janalsncm 3y ago> I believe the outlier problem that this solves only appears for very large models. Any reason to believe this? The author never mentioned it, and I can’t think of any other a priori reason why it should be true.
- WithinReason 3y agoSee figure 1: https://arxiv.org/pdf/2208.07339.pdf https://arxiv.org/pdf/2208.07339.pdf Outliers appear at model size 6.7B and are not present at 2.7B
- janalsncm 3y agoSure, emergent properties can arise as parameters increase. Everyone knows that. That’s a much less specific claim than to say that the benefit of modifying softmax can only arise as an emergent property after N parameters, and therefore the benefit can only be evaluated on models above a certain size. To my understanding the author of TFA isn’t suggesting the same issue as the one in your linked paper.