4 ms·
It's not used a lot in ambiguous tasks like image recognition and audio because you would need to have a state for every single possible wave form or image vari
by CodiePetersen 7y ago
It's not used a lot in ambiguous tasks like image recognition and audio because you would need to have a state for every single possible wave form or image variation. You would have to have a similarity preserving hash or some other related method to reduce related states to a single common state. That's why it's better to use them near the ends of DL networks. If you didn't do that, every image of the same apple with just one pixel different would be a new state.
- guptaneil 7y agoYup, totally agree that Q-learning is not the right tool for classification problems. However, it is great for an agent that needs to act on its world where it can quickly get a reward/punishment for its choice (assuming a relatively limited pool of possible actions). And of course, there's no such thing as the single perfect algorithm. My point is just that I'm surprised Q-learning isn't talked about more.
- CodiePetersen 7y agoWell I don't know about q learning being talked about more but unsupervised reinforcement learning definitely needs more attention in general.
- PeterisP 7y agoCan you name two real world problems with practical applications that fit all the criteria? I.e. where (1) an agent needs to act on its world but it would be okay for it to spend many tries exploring and failing; (2) it will be able to quickly and cheaply (i.e. without 24/7 human supervision) get a reward/punishment for its choice, and (3) the set of world states and possible actions is sufficiently small so that Q-learning is tractable? IMHO Q-learning isn't talked about because it really is not a good fit for the kind of problems people actually want to solve. What behavior does Hiome 'learn' with Q-learning? From your site it's not obvious what actions on the world would be implied where you can actually get some feedback/reward/punish depending on whether these actions were desired by the smart home inhabitants; and the behavior that your page does show - occupancy sensing - is essentially a classification problem.