3 ms·
The model is compiled for apple silicon/metal with 4-bit quantization using mlc-llm https://mlc.ai/mlc-llm/ https://mlc.ai/mlc-llm/ which uses TVM Unity https:/
by 5cott0 3y ago
The model is compiled for apple silicon/metal with 4-bit quantization using mlc-llm https://mlc.ai/mlc-llm/ https://mlc.ai/mlc-llm/ which uses TVM Unity https://tvm.apache.org/ https://tvm.apache.org/ under the hood.