4 ms·
These techniques are not new. And the reason why they’re usually not used is on page 9 in the paper. They require about 10x as many training iterations.
by fxtentacle 1y ago
These techniques are not new. And the reason why they’re usually not used is on page 9 in the paper. They require about 10x as many training iterations.
- typpilol 1y agoYea I saw that training perplexity and thought hmmm...
- shomp 1y agoTurns out using floats is a feature and not a bug?
- Dylan16807 1y agoNo, I don't think so, in that I don't think anyone has ever called that a bug.
- shomp 1y agoIn the paper summary they did not call it a bug explicitly, but they do say there are 32x improvements in using single bits instead.
- reactordev 1y agoTo memory, sure. At the cost of 32x slower speeds.
- Dylan16807 1y agoThat's an obvious exaggeration. The competition is using smaller weights already, some of which are floating point and some of which aren't. And they use full size floats for training.
- imtringued 1y agoThat means their paper is actually worse than SOTA, which is concerned with training in fp4 natively without full precision [0] for QAT. [0] "full precision" in ML usually means 16 bit floats like bfloat16
- Dylan16807 1y agoI wouldn't say "worse". It's focusing on inference cost and leaving training at a default for now.
- personalityson 1y agoUnless each iteration is 90% faster
- deleted 1y ago[deleted]
- amelius 1y agoThis. In fact, it can be slower because hardware is probably not optimized for the 1-bit case, so there may be a lot of low-hanging fruit for hardware designers and we may see improvements in the next iteration of hardware.
- nlitened 1y agoIsn't digital (binary) hardware literally optimized for 1-bit case by definition?
- reactordev 1y agoPeople are confusing word size… The CPU can handle up to word size bits at once. I believe they mean that a lot of assembly was written for integer math and not bit math. Word size 4+ However, it is unlikely we’ll see improvements in this area because by definition, using 64-bit floats uses max word size. So… that’s the max throughput. Sending 1 bit vs 64 bits would be considerably slower so this entire approach is funny.
- observationist 1y agoNo, because there are algorithmic shortcuts that allow approximations and skipped steps in comparison to a strict binary step-by-step calculation, by using in-memory bit reads and implicit rules, among other structural advantages in how GPUs and CPUs instruction sets are implemented in hardware.
- nickpsecurity 1y agoFPGA's could be highly-competitive for models with unusual, but small, bit lengths. Especially single bits since their optimizers will handle that easily.
- PaulHoule 1y agoWhen I was working for startups trying to develop foundation models circa 2015 we were concerned with training more than inference. Today with models that are actually useful training costs matters much less than inference costs. A 10x increase in training costs is not necessarily prohibitive if you get a 10x decrease in inference costs.
- nickpsecurity 1y agoI still don't have a GPT3-class model that was trained without copyright infringement. I'd have so many uses for it from research to production. What's stopping me is the $30 million training cost for 180B models. Even a 30B like Mosaic cost over a million dollars. So, I strongly disagree unless we're talking about the five or six companies that already spend tens of millions on training and keep repeating that. Outside of them, the medium to large models are done infrequently or one off by a small number of other companies. Then, most of us are stuck with their pretraining efforts because we can't afford it ourselves. On my end, I'd rather see a model that drops pretraining costs to almost nothing but costs 10-32x more to do inference. My uses would produce mere MB of output vs hundreds of GB to TB that pretraining requires. A competitive use that costs 32x current prices would probably be profitable for me. Optimizations, which are plentiful for inference, might bring it down further.
- arthurcolle 1y agoWhy are you making something cheap more expensive than it needs to be?
- nickpsecurity 1y agoIt's not cheap. It costs millions to $100 million depending on the model. I was responding to this tradeoff: "A 10x increase in training costs is not necessarily prohibitive if you get a 10x decrease in inference costs." Given millions and up, I'd like that to be 10x cheaper while inference was 10x more expensive. Then, it could do research or coding for me at $15/hr instead of $1.50/hr. I'd just use it carefully with batching.