5 ms·
My theory is that it all boils down to better data and longer post-training period. Cursor got curated data from the trillions reactions of real world developer
by sinuhe69 2mo ago
My theory is that it all boils down to better data and longer post-training period. Cursor got curated data from the trillions reactions of real world developers in real jobs. xAI bought is and used it for its post-training and got Grok 4.5 . Longer post-training on the powerful Colossus cluster helped it get Grok 4.6 , although both versions use the same model with the same number of parameters. Thus, both must use the same pre-trained model as a baseline. See also an article infers the training and release timeline of popular models featured a few days ago here on HN.
Chinese labs must follow similar trajectories plus their specific efficiency improvements. That also explains the jump from DeepSeek 4 performance in April and July releases. They both use the same pre-trained model as well.
- sinuhe69 2mo agoThe release of Gemini Flash 3.7 just 3 weeks after 3.6 confirms my theory, IMO. Only post-training refinement and reinforcement learning (RL) trajectory optimization could yield such high improvements using the same baseline pre-trained model. Flash, MoE models are basically so efficient that the AI labs can put them in a continues post=training loop.
- sinuhe69 2mo agoNow GLM 5.3! And they explicitly confirms my theory: "Scaling post-training is all we did for GLM-5.3. With GLM-5.2 we built the stack: IndexShare for efficient long-context processing, SAO for RL on long-horizon tasks, and slime for large-scale asynchronous training — all running on the long-horizon task environments we have been accumulating. Over the past month we kept scaling on this stack: more environments, more diverse tasks, and more compute spent training on them."