3 ms·
per-chip compute is not the main thing this chip innovates for fast inference, it is the extremely fast memory bandwith. when you do that, you'll loose all of t
by treesciencebot 3y ago
per-chip compute is not the main thing this chip innovates for fast inference, it is the extremely fast memory bandwith. when you do that, you'll loose all of that and will be much worse off than any off the shelf accelerators.
- QuadmasterXLII 3y agoload model, compute a 1k token response (ie, do a thousand forward passes in sequence, one per token), load a different model, compute a response, I would expect the model loading to take basically zero percent of the time in the above workflow