6 ms·
QAT “quantization aware training” means they had it quantized to 4 bits during training rather than after training in full or half precision. It’s supposedly a
by deepsquirrelnet 1y ago
QAT “quantization aware training” means they had it quantized to 4 bits during training rather than after training in full or half precision. It’s supposedly a higher quality, but unfortunately they don’t show any comparisons between QAT and post-training quantization.
- noodletheworld 1y agoI understand that, but the qat models (1) are not new uploads. How is this more significant now than when they were uploaded 2 weeks ago? Are we expecting new models? I don’t understand the timing. This post feels like it’s two weeks late. [1] - https://huggingface.co/collections/google/gemma-3-qat-67ee61ccacbf2be4195c265b https://huggingface.co/collections/google/gemma-3-qat-67ee61...
- llmguy 1y ago8 days is closer to 1 week then 2. And it’s a blog post, nobody owes you realtime updates.
- noodletheworld 1y agohttps://huggingface.co/google/gemma-3-27b-it-qat-q4_0-gguf/tree/main https://huggingface.co/google/gemma-3-27b-it-qat-q4_0-gguf/t... > 17 days ago Anywaaay... I'm literally asking, quite honestly, if this is just an 'after the fact' update literally weeks later, that they uploaded a bunch of models, or if there is something more significant about this I'm missing.
- timcobb 1y agoProbably the former... I see your confusion but it's really only a couple weeks at most. The news cycle is strong in you, grasshopper :)
- osanseviero 1y agoHi! Omar from the Gemma team here. Last time we only released the quantized GGUFs. Only llama.cpp users could use it (+ Ollama, but without vision). Now, we released the unquantized checkpoints, so anyone can quantize themselves and use in their favorite tools, including Ollama with vision, MLX, LM Studio, etc. MLX folks also found that the model worked decently with 3 bits compared to naive 3-bit, so by releasing the unquantized checkpoints we allow further experimentation and research. TL;DR. One was a release in a specific format/tool, we followed-up with a full release of artifacts that enable the community to do much more.
- oezi 1y agoHey Omar, is there any chance that Gemma 3 might get a speech (ASR/AST/TTS) release?
- simonw 1y agoThe official announcement of the QAT models happened on Friday 18th, two days ago. It looks like they uploaded them to HF in advance of that announcement: https://developers.googleblog.com/en/gemma-3-quantized-aware-trained-state-of-the-art-ai-to-consumer-gpus/ https://developers.googleblog.com/en/gemma-3-quantized-aware... The partnership with Ollama and MLX and LM Studio and llama.cpp was revealed in that announcement, which made the models a lot easier for people to use.