4 ms·
I noticed that the app is listed as being ~3Gb in size and the Vicuna 7b model is ~13Gb in size. What did you do to compress it? Same for memory... I think it n
by owyn 3y ago
I noticed that the app is listed as being ~3Gb in size and the Vicuna 7b model is ~13Gb in size. What did you do to compress it? Same for memory... I think it needs 30Gb? And same for CUDA or GPU support... How does that work, or is it just running on the CPU?
- 5cott0 3y agoThe model is compiled for apple silicon/metal with 4-bit quantization using mlc-llm https://mlc.ai/mlc-llm/ https://mlc.ai/mlc-llm/ which uses TVM Unity https://tvm.apache.org/ https://tvm.apache.org/ under the hood.