3 ms·
One of the amazing things is that when every has one GPU, they will actually have 1k-10k agents at their disposal. LLMs and KV caches have amazing performance
by nbardy 1mo ago
One of the amazing things is that when every has one GPU, they will actually have 1k-10k agents at their disposal.
LLMs and KV caches have amazing performance characteristics with concurrent throughput. It scales very non linearly. So the token throughput within a batch scales WAY faster than the tokens per second of each user.
This is the reason the LLM providers have such crazy margins on their costs.