3 ms·
Some kind of compute-in-memory architecture is a good candidate, I think. There are many alternatives here, researched for many years prior to the LLM craze. Ho
by jononor 18d ago
Some kind of compute-in-memory architecture is a good candidate, I think. There are many alternatives here, researched for many years prior to the LLM craze. However economies of scale dominate in chip industries, and this tends to favor more conventional or incremental approaches (to piggyback on existing scale). Alternatively someone needs to have a way of bootstrapping the insane scales needed to be competitive with a better-but-different approach.
So it could be that boring and straightforward stuff like two-chip prefill+decode takes most.