3 ms·
The bottleneck for LLM is fast and large memory, not compute power. Whoever is recommending investing in better chip(ALU) design hasn't done even a basic analy
by frizdny5 2y ago
The bottleneck for LLM is fast and large memory, not compute power.
Whoever is recommending investing in better chip(ALU) design hasn't done even a basic analysis of the problem.
Tokens per second = memory bandwidth divided by model size.