4 ms·
For sure they quickly move to q8, the output quality difference to bf16 is small compared to the speed/capacity gain
by boredatoms 23d ago
For sure they quickly move to q8, the output quality difference to bf16 is small compared to the speed/capacity gain
- NineStarPoint 23d agoYeah q8 made so littler difference back when I was testing such things I'd be surprised if people could quickly notice that as a change. It's got to be either further quantized or some other type of optimization that kicks in when people notice the drop.
- selectodude 23d agoNVFP4 would buy them a huge increase in capacity but I think it would be noticeable.
- Caracas288 23d agoWhy doesn't someone just try to measure this next time!?
- embedding-shape 23d agoCan't really measure without being sure you aren't being messed around with, when it's a remote platform. Stupidly easy to detect when people run such benchmarks/tests against you as well.
- nonethewiser 23d agoCould this explain Opus?
- torginus 23d agoSome people here have remarked previously that while reduced precision doesn't show up in quick prompts, it does severely impact these models' ability to perform long running tasks - to the point that running these big models with severe quantization might be counterproductive as smaller but less quantized ones perform better.