7 ms·
The release of Gemini Flash 3.7 just 3 weeks after 3.6 confirms my theory, IMO. Only post-training refinement and reinforcement learning (RL) trajectory optimiz
by sinuhe69 2mo ago
The release of Gemini Flash 3.7 just 3 weeks after 3.6 confirms my theory, IMO. Only post-training refinement and reinforcement learning (RL) trajectory optimization could yield such high improvements using the same baseline pre-trained model. Flash, MoE models are basically so efficient that the AI labs can put them in a continues post=training loop.