4 ms·
Even if they could generate tokens at that speed on the chip (which maybe they can in theory?) you need to get user tokens onto the chip and the resulting model
by twothreeone 2y ago
Even if they could generate tokens at that speed on the chip (which maybe they can in theory?) you need to get user tokens onto the chip and the resulting model tokens off again and transport them to the user as well. This means at some point the I/O becomes the bottleneck, not the compute. I also suspect it will get faster still, from the announcement it didn't sound like it's "optimal" yet.
- cma 2y agoUser tokens onto the chip and output tokens out are tiny.
- twothreeone 2y agoNot if you're serving tens of thousands of users at the same time.
- cma 2y agoStill tiny at 100,000.