4 ms·
There's no "just" in RL. Fine tuning is very important and could make a lot of difference.
by HeavyStorm 7mo ago
There's no "just" in RL. Fine tuning is very important and could make a lot of difference.
- merlindru 7mo agoapparently GPT-5 uses the same pretrain as 4o did, hah
- lukaslalinsky 7mo agoIndeed, this is quite obvious on Claude models vs Gemini. I fully believe Gemini is more powerful model, but the post training process is nowhere near what Anthropic does, which results in Gemini being horrible at coding sessions, while Claude is excellent.