4 ms·
Easy. A reasoning model with better performance than QWQ but at 21B (like Reka Flash 3) and good tooling call support. A model as “intelligent “ as Qwen2.5 but
by syntaxing 2y ago
Easy. A reasoning model with better performance than QWQ but at 21B (like Reka Flash 3) and good tooling call support. A model as “intelligent “ as Qwen2.5 but personality and creativity of Gemini (or Gemma at a minimum)
- rcpt 2y agoAnd also you should get prizes for using it
- xiphias2 2y agoSomething even cooler would be a model trained for 4 or less (1.33) bit weights instead of quantized after pretraining. Math units are completely underutilized when I'm inferencing with batch size of 1, and post-training quantization under 8 bits loses too much of the precision to make a real difference compared to smaller models with higher precision.