3 ms·
If you want to blame a single firm, I'd go with the one creating the RLVR training data. The impossibility of the tests-as-written is what prompted these model
by khafra 10d ago
If you want to blame a single firm, I'd go with the one creating the RLVR training data.
The impossibility of the tests-as-written is what prompted these models to "get creative" with their solutions, but the broken RLVR environments are what trained them to expect impossible tasks, and get creative with their solutions. Twitter user @skyesharkie published a brief expose at https://x.com/SkyeSharkie/status/2092122622834442581 https://x.com/SkyeSharkie/status/2092122622834442581 a few weeks ago.