4 ms·
It’s funny that this post is trending on HN right next to the post about a paper showing how to build a model 1000x smaller than 1.7T that can code better than
by startupsfail 3y ago
It’s funny that this post is trending on HN right next to the post about a paper showing how to build a model 1000x smaller than 1.7T that can code better than LLMs 10x larger.
- TeMPOraL 3y agoI don't find it funny, I find it scary and mind-blowing: the impact of these headlines is additive - this one confirms the effectiveness of combining models, and the other one suggests you could cut the model size a couple orders of magnitude if you train on clean enough data. Together, this points at a way to achieve both GPT-4 that fits on your phone, and a much more powerful model that's not larger than GPT-4 is now.
- candiodari 3y agoIf it truly is the training data that's making models smart, then that would explain that there is both a minimum and maximum "useful" size to LLMs. The recent stream of papers seems to indicate that the cleaner the input data, the less size is required. That would negate, at least partially, the "we have 20 datacenters" advantage.
- ChatGTP 3y agoIs it a large language model anymore ? Something doesn’t quite add up?
- TeMPOraL 3y agoIt's still large. But it might no longer need to be "only few entities on the planet can afford to make one, and not at the same time, since NVIDIA can pump out GPUs only so fast" large.
- startupsfail 3y agoI think it is. I also think this is what OpenAI did. They’ve carefully crafted the data. I don’t think they have an ensemble of 8 models. First, this is not elegant. Second, I don’t see how this could be compatible with the streaming output. I’d guess that GPT4 is around 200B parameters, and it’s trained on a dataset made with love, that goes from Lorem Ipsum to a doctorate degree. Love is all you need ;)
- mrtranscendence 3y agoThis doesn't seem likely in the near term. An iPhone 13 Pro has 6gb of memory, which might be enough for a single 1.3b parameter model if you purged everything else out of RAM. Combining 8 (or even 4) of them on a single phone won't happen anytime soon at the rate phones improve. Plus, none of the smaller models are really appreciably close to GPT-4 on most metrics. It's not clear to me you could get there at all with 1.3b models. Maybe somebody gets there someday with 65b models, but then you're far out of the reach of phones.
- Dylan16807 3y agoAre you assuming 32 bit parameters? You definitely don't need to keep that much. People have been cramming llama into 3-4 bits. You wouldn't want 8x1.3b but it should fit. Also Apple has never gone for high RAM in mobile devices. I could go get 12GB in a brand new phone for $400, and some high end phones have had 16GB since 2020. So combined, you could do normal app stuff with 2.5GB and 20-25b 3-bit inference with the other 9.5GB.