3 ms·
KV caching status? What's the point of 1000tok/s if you have to do prefill on every agentic turn which at 100k depth would make it 1.5 min latency every turn?
by lostmsu 2mo ago
KV caching status?
What's the point of 1000tok/s if you have to do prefill on every agentic turn which at 100k depth would make it 1.5 min latency every turn?
- walrus01 2mo agoInformation about RAM type/size and connection topology of the RAM to be used for context cache seems to be conspicuously absent from the slick looking marketing materials.
- gpm 2mo agoThere's a few more details at the bottom of this page: https://www.cerebras.ai/blog/introducing-cerebras-cs-4 https://www.cerebras.ai/blog/introducing-cerebras-cs-4 44GB on-chip-sram * 3 chips. Per chip: 43.2 PB/s memory access + 53.5 PB/s on-chip fabric bandwidth + 2.4 Tbits/s "IO" bandwidth (I think that means their RoCE v2 RDMA over Ethernet interface). I suspect there might be a certain amount of customization for how much RAM they attach when you order it.
- porridgeraisin 2mo agoThey have managed to make the external link 300GBps/2us. Cs3 was 150/5. This is 1/3rd blackwells nvlink c2c bandwidth already. Not too bad. We can make KV cache offload work with that I suppose. If magically KV cache was not an issue, pipeline parallelism on cerebras can be quite pleasant. As for the KV cache offload, I have hopes their CPO solution they're trying with that canadian company ends up bearing fruit.