3 ms·
> I personally think that Gemini 2.5 Pro's superiority comes from having hundreds or thousands RL tasks (without any proof whatsoever, so rather a feeling). Gi
by t55 1y ago
> I personally think that Gemini 2.5 Pro's superiority comes from having hundreds or thousands RL tasks (without any proof whatsoever, so rather a feeling).
Given that GDM pioneered RL, that's a reasonable assumption
- flowerthoughts 1y agoAssuming with GDM, you mean Google-Deep Mind. They pioneered RL with deep nets as policy function estimator. The deep nets being a result of CNNs and massive improvements in hardware parallelization at the time. RL was established, at the latest, with Q-learning in 1989: https://en.wikipedia.org/wiki/Q-learning https://en.wikipedia.org/wiki/Q-learning
- t55 1y agoi didn't say they invented everything; in science you always stand on the shoulders of giants i still think my original statement is fair
- lechatonnoir 1y ago"gdm pioneered rl" is definitely not actually right, but it's correct to assert that they were huge players. people who knew from context that your statement was broadly not actually right would know what you mean and agree on vibes. people who didn't could reasonably be misled, i think.