5 ms·
Bigger issue with offline RL in the real world (I.e. not Atari video games) has been the assumption of reward labeling. Who’s giving you reward labels at scale?
by polygamous_bat 3y ago
Bigger issue with offline RL in the real world (I.e. not Atari video games) has been the assumption of reward labeling. Who’s giving you reward labels at scale? In my opinion that’s why we haven’t seen any large scale real world success stories using offline RL.
- AndrewKemendo 3y agoI fully agree with you that instrumentation is one of the biggest barriers to state, action, trajectory and reward feedback However, instrumentation assumes that there’s a control regime that could actually control whatever the system is mechanically, and that’s generally not true. So it’s almost a chicken and an egg problem where you can do instrumentation for non-autonomous-control systems in order to get state-action-reward data, but because you don’t actually have an actuated control system that you can specify and build mechanically, your targets for state-action-reward tuple aren’t the same That is to say unless you’re actively collecting data from an autonomous system that’s being used non-autonomously then you’re not gonna be able to transition from a non-autonomous control regime to an autonomous control regime