4 ms·
Better to run the Q8 model on an epyc pair with 768GB, you'll get the same performance
by chriscappuccio 2y ago
Better to run the Q8 model on an epyc pair with 768GB, you'll get the same performance
- ltbarcly3 2y agoThe Q8 model is totally different?
- manmal 2y agoMy experience with quantizations is that anything below 6 is noticeably worse. Coherence suffers. I’ve rarely gotten anything really useful out of a Q4 model, code wise. For transformations they are great though, eg convert JSON to Markdown and vice versa.