3 ms·
RLHBARF - Reinforcement Learning from Human But Actually Robot Feedback
by jeron 3y ago
RLHBARF - Reinforcement Learning from Human But Actually Robot Feedback
- tsunamifury 3y agoUh, I think this should be a thing considering we have LLMs training other LLMs. so we end up with this BARF loop. There is an actual interesting question here if BARF loops can create new synthetic knowledge, or if humans need to be the 'spark of novelty' in that loop to drive it to new knowledge.
- yetanotherloser 3y agoI love your acronym formation there. There's another thread running at the moment about BARF degrading the machine's performance. I don't have enough information to be confident, but it sounded plausible and likely.
- jeron 3y agoif BARF becomes a thing I really hope I'll be credited for it haha