5 ms·
Aside from Meta is there any reason to think the big AI labs are still using LMArena data for training? The weaknesses are well understood and with the shift to
by thorum 9mo ago
Aside from Meta is there any reason to think the big AI labs are still using LMArena data for training? The weaknesses are well understood and with the shift to RL there are so many better ways to design a reward function.
- dk8996 9mo agoSuch as?
- thorum 9mo agoMy favorite is LLM-as-judge with a detailed rubric as discussed here: https://www.dbreunig.com/2025/07/31/how-kimi-rl-ed-qualitative-data-to-write-better.html https://www.dbreunig.com/2025/07/31/how-kimi-rl-ed-qualitati...
- nl 9mo agoI don't think anyone has ever used it as training. But yes labs still do seem to target it as goal (which is a different thing).