3 ms·Depending on how much ram you have, fine-tune then quanitize & run with llama.cpp, which works quite well.by miloignis 3y agoDepending on how much ram you have, fine-tune then quanitize & run with llama.cpp, which works quite well.