3 ms·
Why bring "model parameters ... on demand to DRAM"? Maybe it is better to move LLM processing right unto flash controller chip... (after adding bfloat16 and mat
by pulse7 3y ago
Why bring "model parameters ... on demand to DRAM"? Maybe it is better to move LLM processing right unto flash controller chip... (after adding bfloat16 and matrix multiplication support into controller circuitry)
- jng 3y agoProbably because changing the software to use a different read pattern is doable in a few weeks/months on your existing systems, and changing anything in the flash controller is a wicked project probably only available to hardware manufacturers, and which will take months to years given the immensely slower hardware iteration cycles (even if it's "just" firmware changes).
- YetAnotherNick 3y agoThe bottleneck is anyways going to be flash read speed so it doesn't matter there are 10 extra steps or if output is computed in the flash.
- moffkalast 3y agoThat's what the Google TPU is in a nutshell as I understand it, loading weights into memory cells between fpus.
- keivanalizadeh 3y agoOne of the initial ideas was of course doing computation inside flash, but we didn't try to go that path for two reasons: 1. It's not as easy to change the controller, even if you do it was not obvious for me if we need for software updates at system level. In current way it is a standalone project. 2. I guess for a LLM scale flash controller chip might not be strong enough for computation. Additional hardware inside flash might be required for that.