4 ms·
So the AI is trained from the previous frame and the input and tries to predict the next frame, correct? How could you achieve object permanence this way? Will
by folli 2y ago
So the AI is trained from the previous frame and the input and tries to predict the next frame, correct?
How could you achieve object permanence this way? Will it 'automatically' appear given more training data or more hidden layers? How is this handled in other approaches?
- 383toast 2y agothe same way LLMs are trained to predict next tokens from prev tokens, just longer context = better memory = object permanence
- echoangle 2y agoBut then you can’t just give the previous frame, with the LLM analogy you would have to give the last few thousand frames (that’s the context window, right?). If you only give the previous frame, that’s like having an LLM that only gets the single previous token and has to predict the next one.
- int_19h 2y agoIndeed. Although more recently they figured out a way to feed the hidden state as the new input, which basically allows the model to "continue thinking" in vectors without round-tripping it via words (or pixels). Presumably if you were to take that and build a large enough NN to accommodate all the necessary state it needs to carry and all the rules it needs to be able to execute, then after training it on enough game input you'd have a proper world simulation. Of course, as the article rightly notes, then you have just successfully reimplemented Minecraft in a way that is orders of magnitude more computationally expensive...
- _flux 2y agoPerhaps the trick used by text-based LLMs could be used: when the context window starts filling up, the LLM is asked to summarize the existing data in the context, thus compressing it (lossily..) into smaller space.
- Sharlin 2y agoMore previous tokens in this case would mean more previous frames. But there's really no reason to just stick to rendered pixels as input (except for novelty's sake) because we could train directly on snapshots of full game state.
- philipwhiuk 2y agoYeah but then it's not generalizable
- airstrike 2y agoDoesn't that depend on how such game state is modeled?
- kvdveer 2y agoIf training is indeed done on frame + input, any information that isn't in either of those data sources is simply not there. To achieve object permanence, there needs to be some persistent off-screen data from frame to frame. There's a way to achieve this: train on frame+input+woldstate -> frame+woldstate