3 ms·
From my experimentation it seems like there's some significant loss in accuracy running the tuned LoRa models through llama.cpp (due to bugs/differences in infe
by antimatter15 4y ago
From my experimentation it seems like there's some significant loss in accuracy running the tuned LoRa models through llama.cpp (due to bugs/differences in inference or tokenization), even aside from losses due to quantization.