3 ms·
Thank you, that demo was insane! Follow up (noob) question: Are you using a KV cache? That would significantly increase your memory requirements. Or are you fo
by ppsreejith 3y ago
Thank you, that demo was insane!
Follow up (noob) question: Are you using a KV cache? That would significantly increase your memory requirements. Or are you forwarding the whole prompt for each auto-regressive pass?
- tome 3y agoYou're welcome! Yes, we have KV cache. Being able to implement this efficiently in terms of hardware requirements and compute time is one of the benefits of our deterministic chip architecture (and deterministic system architecture).
- ppsreejith 3y agoThanks again! Hope I'm not overwhelming but one more question: Are you decoding with batch size = 1 or is it more?
- tome 3y agoThat's OK, feel free to keep asking! I think currently 1. Unlike with graphics processors, which really need data parallelism to get good throughput, our LPU architecture allows us to deliver good throughput even at batch size 1.