4 ms·
When you say it can run on consumer gpus, do you mean pretty much just the 4090/3090 or can it run on lesser cards?
by 7speter 3y ago
When you say it can run on consumer gpus, do you mean pretty much just the 4090/3090 or can it run on lesser cards?
- halflings 3y agoI was able to run the 4bit quantized LLAMA2 7B on a 2070 Super, though latency was so-so. I was surprised by how fast it runs on an M2 MBP + llama.cpp; Way way faster than ChatGPT, and that's not even using the Apple neural engine.
- hereonout2 3y agoIt runs fantastically well on M2 Mac + llama.cpp, such a variety of factors in the Apple hardware making it possible. The ARM fp16 vector intrinsics, the Macbook's AMX co-processor, the unified memory architecture, etc. It's more than fast enough for my experiments and the laptop doesn't seem to break a sweat.
- gsuuon 3y agoQuantized 7B's can comfortably run with 8GB vram
- deleted 3y ago[deleted]