3 ms·
They weren't misguided, they just over-focused on scaling parameters instead of appropriately scaling quality data in tandem. A 4B has limited capacity to enco
by Vetch 2y ago
They weren't misguided, they just over-focused on scaling parameters instead of appropriately scaling quality data in tandem.
A 4B has limited capacity to encode knowledge and algorithmic circuits. It's also too small to learn programs whose execution exceed the sizes of circuits it can encode. There is a hard cap on how much we can squeeze out of small models. What we need is better consumer hardware, so we don't have to hope for miracles.
Another hope is that the 1.58 bit/ternary quantization aware training of model scaling pans out. That'd be another axis of inefficiency beyond just parameter count.
- littlestymaar 2y ago> They weren't misguided, they just over-focused on scaling parameters instead of appropriately scaling quality data in tandem. That's exactly what I mean by misguided, they were focusing on the wrong metric. > A 4B has limited capacity to encode knowledge and algorithmic circuits. It's also too small to learn programs whose execution exceed the sizes of circuits it can encode. There is a hard cap on how much we can squeeze out of small models. There must be some kind of cap indeed, but at this point we have no idea where it is, especially now that there are emerging infrastructures that aren't just transformers popping out. > What we need is better consumer hardware, so we don't have to hope for miracles. Hardware improvements are going to be a big factor, but there's no reason to think the software part (+ training data) isn't going to continue improving a lot.