3 ms·
Could we all get bigger FPGAs and load the model onto it using the same technique?
by punnerud 8mo ago
Could we all get bigger FPGAs and load the model onto it using the same technique?
- fercircularbuf 8mo agoI thought about this exact question yesterday. Curious to know why we couldn't, if it isn't feasible. Would allow one to upgrade to the next model without fabricating all new hardware.
- wmf 8mo agoFPGAs have really low density so that would be ridiculously inefficient, probably requiring ~100 FPGAs to load the model. You'd be better off with Groq.
- menaerus 8mo agoNot sure what you're on but I think what you said is incorrect. You can use hi-density HBM-enabled FPGA with (LP)DDR5 with sufficient number of logic elements to implement the inference. Reason why we don't see it in action is most likely in the fact that such FPGAs are insanely expensive and not so available off-the-shelf as the GPUs are.
- wmf 8mo agoYeah, FPGA+HBM works but it has no advantage over GPU+HBM. If you want to store weights in FPGA LUTs/SRAM for insane speed you're going to need a lot of FPGAs because each one has very little capacity.
- menaerus 8mo agoOk, then I may have misunderstood what you were saying. If the only thing we are interested is to store all the weights into the block RAM or LUTs then, yeah, that wouldn't be possible. I understood the OPs question a bit differently too.
- generuso 8mo agoYou could [1], but it is not very cheap -- the 32GB development board with the FPGA used in the article used to cost about $16K. [1] https://arxiv.org/abs/2401.03868 https://arxiv.org/abs/2401.03868
- sowbug 8mo agoFPGAs aren't very power-efficient. You could do it, but the numbers wouldn't add up for anything but prototyping.