3 ms·
Is there a single reason why we can't just "distribute" the online softmax? Each die-attached PIM accelerator computes online softmax for its own KVs. Then the
by ACCount37 1mo ago
Is there a single reason why we can't just "distribute" the online softmax?
Each die-attached PIM accelerator computes online softmax for its own KVs. Then the central unit gathers the softmax intermediates, one intermediate per die, and uses those to compute the final softmax.
The PIM win is that we crater the memory traffic between the central accelerator and the memory dies for attention ops. Most of the attention bandwidth never leaves the memory.
This isn't "run the entire LLM in PIM", no - this is "offload the parts of LLM that benefit from PIM the most to PIM".