5 ms·
7B models are small enough to be usable on a smartphone, so a local handheld assistant sounds like a use case.
by snowram 3y ago
7B models are small enough to be usable on a smartphone, so a local handheld assistant sounds like a use case.
- ComputerGuru 3y agoAre they? Unquantized, Llama 2 7b needs over 14GB of GPU (or shared) memory.
- polygamous_bat 3y ago"Unquantized" is the key word here: with quantization you can get a 4-8x improvement without much performance degradation.
- throeaaysn 3y ago[dead]
- minimaxir 3y ago"usable" is not the same as practical. Even running a quantized and optimized LLM on a smartphone would kill battery life at minimum.
- brucethemoose2 3y agoTry MLC-LLM. Its not as bad as you'd think. In the future(?), they will probably use the AI blocks instead of the GPU, which are very low power.