3 ms·
Reinforcement learning w/ human feedback. What u guys are describing is the alignment problem
by meow_mix 4y ago
Reinforcement learning w/ human feedback. What u guys are describing is the alignment problem
- mistymountains 4y agoThat’s just a supervised fine tuning method to skew outputs favorably. I’m working with it on biologics modeling using laboratory feedback, actually. The underlying inference structure is not changed.