5 ms·
This is kind of amazing given that obviously, GPT-4 never contained such tasks and data. I think it puts an end to the claim that "language models are only stoc
by bartwr 3y ago
This is kind of amazing given that obviously, GPT-4 never contained such tasks and data.
I think it puts an end to the claim that "language models are only stochastic parrots and cannot do any reasoning".
No, this is 100% a form of reasoning and furthermore, learning that is more similar to how humans learn (gradient-less).
I still don't understand it and it blows my mind - how such properties emerge just from compressing the task of next word prediction. (Yes, I know this is oversimplification, but not a misleading one).
- chinchilla2020 3y agoI read through the code and tried it out for 15 mins. It's a hard-coded program that can do a text search for it's own hard-coded, human-implemented functions. Apparently it can string those functions together, but doesn't do it correctly. https://github.com/MineDojo/Voyager/tree/main/voyager/control_primitives https://github.com/MineDojo/Voyager/tree/main/voyager/contro... 20 minutes of light reading through the repository pretty much dispels any notions that this is a self-learning system that can reason and think. It's the same minecraft automation we have been seeing for a decade now, with a chatbot text search builtin.
- asperous 3y agoWhile I do believe LLMs can perform some reasoning, I'm not sure this is the best example as all the reasoning you would ever need for Minecraft is well contained in the data set used to train it. A lot has been written about minecraft. To me, it would be more convincing if they developed an enterly new game with somewhat novel and arbitrary rules and saw if the embodied agent could learn this game.
- notamy 3y agoLooking at the paper, as I understand it they're using Mineflayer https://github.com/PrismarineJS/mineflayer https://github.com/PrismarineJS/mineflayer and passing parts of the state of the game as JSON to the LLM that are used for code generation to complete tasks. > I still don't understand it and it blows my mind - how such properties emerge just from compressing the task of next word prediction. The Mineflayer library is very popular, so all the relevant tasks are likely already extant in the training data.
- emptysongglass 3y agoYou declare: > I think it puts an end to the claim that "language models are only stochastic parrots and cannot do any reasoning". But then two sentences later: > I still don't understand it and it blows my mind I've said this before to others and it bears repeating because your line of thinking is dangerous (not sudden AI cataclysm): to feel so totally qualified to make such a statement armed with ignorance, not knowledge, is the cause of mass hysteria around LLMs. What is happening can be understood without resorting to the sort of magical thinking that ascribes agency to these models.
- godelski 3y ago> What is happening can be understood without resorting to the sort of magical thinking that ascribes agency to these models. This is what has (as an ML researcher) made me hate conversations around ML/AI recently. Honestly getting me burned out on an area of research I truly love and am passionate about. A lot of technical people openly and confidently are talking about magic. Talking as if the model didn't have access to relevant information (the "zero-shot myth") and other such nonesense. It is one thing for a layman to say these things, but another to see them on the top comment on a website aimed at people with high tech literacy. And even worse to see it coming from my research peers. These models are impressive, and I don't want to diminish that (I shouldn't have to say this sentence but here we are), but we have to be clear that the models aren't magic either. We know a lot about how they work too. They aren't black boxes, they are opaque, and every day we reduce the opacity. For clarity: here's an alternative explanation to the results that's even weaker than the paper's settings (explains autogpt better). LLM has a good memory. LLM is told (or can infer through relevant information like keywords: "diamond axe") that it is in a minecraft setting. It then looks up a compressed version of a player's guide that was part of its training data. It then uses that data to execute goals. This is still an impressive feat! But it is still in line with the stochastic parrot paradigm. I'm not sure why people don't think stochastic parrots aren't impressive. They are. But right now ML/AI culture feels like Anime or weed culture. The people it attracts makes you feel embarrassed to be associated with it.
- famouswaffles 3y ago
- godelski 3y ago> GPT-4 never contained such tasks and data No task, but we need to be clear that it did have the data. Remember that GPT4 was trained on a significant portion of the internet, which likely includes sites like Reddit and game fact websites. So there's a good chance GPT4 learned the tech tree and was trained on data about how to progress up that tree, including speed runner discussions. (also remember that as of March GPT4 is also trained on images, not just text) What data it was trained on is very important and I'm not sure why we keep coming back to this issue. "GPT4 has no zero-shot data" should be as drilled into everyone's head as sayings like "correlation does not equate to causation" and "garbage in, garbage out". Maybe people do not know this data is on the internet? But I'm surprised if the average HN user thought that way. This doesn't make the paper less valuable or meaningful. But it is more like watching a 10 year old who's read every chess book and played against computers beat (or do really well) against a skilled player vs a 10 year old who's never heard of chess beating a skilled player. Both are still impressive, one just seems like magic though and should raise suspicion.