4 ms·
Mixtral8x22B looking very strong! Finetunes seem comparable to GPT4!
by mcbuilder 2y ago
Mixtral8x22B looking very strong! Finetunes seem comparable to GPT4!
- littlestymaar 2y agoThat's quite crazy to see that a model that's barely bigger than GPT-3 (and only uses a fraction of the compute due to its MoE architecture) can achieve such a thing. It looks like the people who focasted that AI models would need to keep growing to improve their performance where completely misguided. I wonder if in 3 to 4 years we'll end up with models with less than 4B parameters we comparable performance as today's State of the Art.
- int_19h 2y agoIt can also simply mean that benchmarks are not particularly representative of real-world performance on challenging reasoning tasks.
- littlestymaar 2y agoOf course they aren't, but it's still pretty evident that most opensource models are miles ahead of GTP-3, even the ones that are only a fraction of its size so there's still some massive improvement that doesn't depend on the model size itself.
- int_19h 2y agoGPT-3 sure. There's no model that's even close to GPT-4, though. And we don't really know much about the actual size of that, only lots of speculation.
- littlestymaar 2y agoMy entire point since the beginning is that the models out there improved a ton compared to what used to be the state of the art just 4 years ago, with an equivalent when not much smaller number of parameters (GPT-3 was 175B params and on MMLU it performs worse than Mistral 7B [1]). [1]: https://paperswithcode.com/sota/multi-task-language-understanding-on-mmlu https://paperswithcode.com/sota/multi-task-language-understa...
- Vetch 2y agoThey weren't misguided, they just over-focused on scaling parameters instead of appropriately scaling quality data in tandem. A 4B has limited capacity to encode knowledge and algorithmic circuits. It's also too small to learn programs whose execution exceed the sizes of circuits it can encode. There is a hard cap on how much we can squeeze out of small models. What we need is better consumer hardware, so we don't have to hope for miracles. Another hope is that the 1.58 bit/ternary quantization aware training of model scaling pans out. That'd be another axis of inefficiency beyond just parameter count.
- littlestymaar 2y ago> They weren't misguided, they just over-focused on scaling parameters instead of appropriately scaling quality data in tandem. That's exactly what I mean by misguided, they were focusing on the wrong metric. > A 4B has limited capacity to encode knowledge and algorithmic circuits. It's also too small to learn programs whose execution exceed the sizes of circuits it can encode. There is a hard cap on how much we can squeeze out of small models. There must be some kind of cap indeed, but at this point we have no idea where it is, especially now that there are emerging infrastructures that aren't just transformers popping out. > What we need is better consumer hardware, so we don't have to hope for miracles. Hardware improvements are going to be a big factor, but there's no reason to think the software part (+ training data) isn't going to continue improving a lot.