5 ms·
There might be some truth in what you say for very large image and language models that use supervised learning. It is really hard to see how this 'it's just a
by jah242 4y ago
There might be some truth in what you say for very large image and language models that use supervised learning.
It is really hard to see how this 'it's just a lot of good data' view applies to deep reinforcement learning where the model learns multi step policies from raw input data (e.g a camera on a robot) with only a rough high level reward function to guide it.
If therefore (as seems to be the case) you can abstract the information humans need to provide to the model/learning system to ever high levels of reward function (and thereby vastly reduce the information provided by humans) then it seems very hard to argue that the model (and the training process) isn't doing to some degree what you describe as:
'incredible amounts of experimental work to carve-the-world along its joints, ie., to have the right concepts; and incredible amounts of work to measure along its joints, ie., to have the right units. And then to eliminate all the coincidences and irrelevances.'
For example, imagine a robot learning from scratch to pick objects up based on raw pixel data with only a scalar reward function - where in this process is the human preparing the data so the model only has to average?
- mjburgess 4y ago> For example, imagine a robot learning from scratch to pick objects up based on raw pixel data with only a scalar reward function - where in this process is the human preparing the data so the model only has to average? Great -- so do you have an example of such a system? I'd be inclined, initially, to deny that it exists. If your reward function expresses a reward for the goal of "picking up objects with (pixel-space) properties etc.", you're cheating. In this case, the reward function serves the role of the data: ie., prepared by us to work. Indeed, a function is just a dataset -- and the reward function here is being sampled by the system. You'd need to show me a system whose reward function / dataset didnt "contain the solution", in the manner of animals who respond to the world without already having all the information about it. The relevant capacity a system needs to have, in both cases, is being able to take a profoundly ambiguous environment and produce a dataset/reward-fn which "carves along its joints". Ie., which effectively eliminates that ambiguity. When such ambiguity & coincidence is eliminated, there's basically nothing left to do -- it's that basic nothing which we task machines with doing. Ie. running `mean(sample(unambigious relevant well-carved data))`. You'll note its the *properties* of the data which express intelligence & learning.
- moyix 4y agoI realize it's a much more cleanly defined domain, but how does this interpretation explain MuZero?
- xfs 4y agoPlenty of RL systems learn to play video games just fine without fine-tuned rewards, but I see this line of thought isn't actually what you're getting at. I would assume serious ML people would not be overly ambitious and overstep their claims beyond empirical realms. You were saying ML "uncovers latent representational structure not present in the data", but I would guess the claim, if that is what you're going against, is merely that the latent structures exist, and no Truth is really "uncovered" by ML per se, in the Heideggerian sense. I agree ML hasn't really produced an Understanding of the world. The carving along the joints is in other words a symbolic abstraction of the world that is a radical simplification, for which only Reason is capable of, and ML hasn't shown to be capable of Reason. As an aside, I also would not assume the ambiguity you refer to can be fully eliminated even by human intelligence, just see how languages are fully of ambiguity, or even quantum mechanics. But again, when philosophical critiques are launched against ML, the usual story is ML advocates would retreat to the success of ML in the empirical realms. I'm reminded of the Norvig vs Chomsky debate by this.
- mjburgess 4y agoI think this debate has historically suffered from being conducted purely philosophically. Hediegger, Dreyfus (Ponty et al.) needed a bit more science and mathematics to see through the show. All we need to do to make the Heideggerian point is ask the RL researcher what his reward function is. Have him right it out, and note, that its a disjunction of properties which already carve the environment of the robot. In otherwords, the failure of AI is far less of a mystery than philosophy alone seems to imply. Its a failure in a very very simple sense if one just asks the right technical questions. For RL, all we need ask is, "what will the machine do when it encounters an object outside of your pregiven disjunction?" The answer, of course, is fall over. Hardly what we fear when the wolf learns our movements, or what we love when a person shows us how to play a piano for the first time. The very thing we want, and we are told we have, isnt there... and it's not "not there" philosophically... its not there in the actual journal paper.