3 ms·
Using the 4bit quant on an M1 64GB I'm getting ~65 tps for prompt processing and ~11 tps token generation using oMLX to serve the models and pi as a harness.
by noman-land 2mo ago
Using the 4bit quant on an M1 64GB I'm getting ~65 tps for prompt processing and ~11 tps token generation using oMLX to serve the models and pi as a harness.