4 ms·
The method of the NS advance involved RLHE (reinforcement learning via human example), and that is only open-ended if users continue to advance the frontier wit
by whatshisface 19d ago
The method of the NS advance involved RLHE (reinforcement learning via human example), and that is only open-ended if users continue to advance the frontier within chats ahead of publications.
- thorum 19d agoSure, but the point is that the labs use more powerful internal models for research work, not public models. Public models tend to lag the internal frontier by a decent margin, and are constrained in other ways by monitoring. It’s just not a useful indicator.