4 ms·
What an fascinating concept. I guess this won't be useful for any kind of realtime feedback system, though?
by simongray 4y ago
What an fascinating concept. I guess this won't be useful for any kind of realtime feedback system, though?
- colordrops 4y agoWhy not?
- simongray 4y agoHow would one make a reliable realtime system that depends entirely on unknown network conditions? Perhaps inside a closed network it is possible.
- _joel 4y agoThat's orthoganal to a realtime system. You can infer at a fair speed so realtime would be possible.
- simongray 4y agoGuarantees are not orthogonal to realtime feedback, they are essential. If I write a query, it is not irrelevant whether it takes 1 second or 1 minute to return at any given moment. You write that speed can be inferred, but the analogy that was used here is BitTorrent—and my experience with BitTorrent tells me that it certainly cannot be inferred.
- _joel 4y agoIf you read the article text and the response from the dev then yes, inference can happen at 1/s or if parallelised, more. I'm not sure what your parameters are for a realtime system. If you're talking about network reliability, that's a different issue. Yes it can infer quickly, can it do it reliably is another matter.
- borzunov 4y agoA Petals dev here. It is not real-time, but we think the speed of ~1 token/sec may be enough for some interactive apps such as chat bots (especially, if you show tokens to a user once they are generated). You can try one at http://chat.petals.ml http://chat.petals.ml (heads-up: it may be laggy right now due to lots of HN users trying out the system). Of course, you could do better if you have enough high-end GPUs to host the entire model yourself (3x A100 or 8x 3090). But if you don't, 1 token/sec is much faster than what you get with other existing methods.
- dpflan 4y agoI have not read the technical details, apologies for ignorance, but is there an opportunity for caching?
- jerpint 4y agoProbably not, since you need to compute the activations of unknown inputs and there could be infinitely many variations of them
- KaoruAoiShiho 4y agoWhat are the speeds of other existing methods?
- borzunov 4y agoTheoretical best-case for RAM offloading is 5.5 sec/token, for SSD offloading - 22 sec/token. Implementations we've tested are not faster than 10 sec/token though. See details in our paper: https://arxiv.org/pdf/2209.01188.pdf https://arxiv.org/pdf/2209.01188.pdf