3 ms·
It's mentioned in their paper: https://arxiv.org/pdf/2209.01188.pdf https://arxiv.org/pdf/2209.01188.pdf Several recent works aim to democratize LLMs by “o
by taink 4y ago
It's mentioned in their paper: https://arxiv.org/pdf/2209.01188.pdf https://arxiv.org/pdf/2209.01188.pdf
Several recent works aim to democratize LLMs
by “offloading” model parameters to slower but
cheaper memory (RAM or SSD), then running
them on the accelerator layer by layer (Pudipeddi
et al., 2020; Ren et al., 2021). This method allows
running LLMs with a single low-end accelerator
by loading parameters from RAM justin-time for
each forward pass. Offloading can be efficient for
processing many tokens in parallel, but it has inher-
ently high latency: for example, generating one to-
ken with BLOOM-176B takes at least 5.5 seconds
for the fastest RAM offloading setup and 22 sec-
onds for the fastest SSD offloading. In addition,
many computers do not have enough RAM to of-
fload 175B parameters.
- dpflan 4y agoIs a mobile device / edge device a possible participant / source of resources?