4 ms·
There is no text conditioning provided to the SD model because they removed it, but one can imagine a near future where text prompts are enough to create a fun
by refibrillator 2y ago
There is no text conditioning provided to the SD model because they removed it, but one can imagine a near future where text prompts are enough to create a fun new game!
Yes they had to use RL to learn what DOOM looks like and how it works, but this doesn’t necessarily pose a chicken vs egg problem. In the same way that LLMs can write a novel story, despite only being trained on existing text.
IMO one of the biggest challenges with this approach will be open world games with essentially an infinite number of possible states. The paper mentions that they had trouble getting RL agents to completely explore every nook and corner of DOOM. Factorio or Dwarf Fortress probably won’t be simulated anytime soon…I think.
- mlsu 2y agoWith enough computation, your neural net weights would converge to some very compressed latent representation of the source code of DOOM. Maybe smaller even than the source code itself? Someone in the field could probably correct me on that. At which point, you effectively would be interpolating in latent space through the source code to actually "render" the game. You'd have an entire latent space computer, with an engine, assets, textures, a software renderer. With a sufficiently powerful computer, one could imagine what interpolating in this latent space between, say Factorio and TF2 (2 of my favorites). And tweaking this latent space to your liking by conditioning it on any number of gameplay aspects. This future comes very quickly for subsets of the pipeline, like the very end stage of rendering -- DLSS is already in production, for example. Maybe Nvidia's revenue wraps back to gaming once again, as we all become bolted into a neural metaverse. God I love that they chose DOOM.
- energy123 2y agoThe source code lacks information required to render the game. Textures for example.
- TeMPOraL 2y agoObviously assets would get encoded too, in some form. Not necessarily corresponding to the original bitmaps, if the game does some consistent post-processing, the encoded thing would more likely be (equivalent to) the post-processed state.
- hoseja 2y agoFinally, the AI superoptimizing compiler.
- mistercheph 2y agoThat’s just an artifact of the language we use to describe an implementation detail, in the sense GP means it, the data payload bits are not essentially distinct from the executable instruction bits
- electrondood 2y agoThe Holographic Principle is the idea that our universe is a projection of a higher dimensional space, which sounds an awful lot like the total simulation of an interactive environment, encoded in the parameter space of a neural network. The first thing I thought when I saw this was: couldn't my immediate experience be exactly the same thing? Including the illusion of a separate main character to whom events are occurring?
- Jensson 2y ago> With enough computation, your neural net weights would converge to some very compressed latent representation of the source code of DOOM. Maybe smaller even than the source code itself? Someone in the field could probably correct me on that. Neural nets are not guaranteed to converge to anything even remotely optimal, so no that isn't how it works. Also even though neural nets can approximate any function they usually can't do it in a time or space efficient manner, resulting in much larger programs than the human written code.
- mlsu 2y agoCould is certainly a better word, yes. There is no guarantee that it will happen, only that it could. The existence of LLMs is proof of that; imagine how large and inefficient a handwritten computer program to generate the next token would be. On the flipside, human beings very effectively predicting the next token, and much more, on 5 watts is proof that LLM in their current form certainly are not the most efficient method for generating next token. I don't really know why everyone is piling on me here. Sorry for a bit of fun speculating! This model is on the continuum. There is a latent representation of Doom in weights. some weights, not these weights. Therefore some representation of doom in a neural net could become more efficient over time. That's really the point I'm trying to make.
- godelski 2y ago> With enough computation, your neural net weights would converge to some very compressed latent representation of the source code of DOOM. You and I have very different definitions of compression https://news.ycombinator.com/item?id=41377398 https://news.ycombinator.com/item?id=41377398 > Someone in the field could probably correct me on that. ^__^
- _hark 2y agoThe raw capacity of the network doesn't tell you how complex the weights actually are. The capacity is only an upper bound on the complexity. It's easy to see this by noting that you can often prune networks quite a bit without any loss in performance. I.e. the effective dimension of the manifold the weights live on can be much, much smaller than the total capacity allows for. In fact, good regularization is exactly that which encourages the model itself to be compressible.
- godelski 2y agoI think your confusing capacity with the training dynamics. Capacity is autological. The amount of information it can express. Training dynamics are the way the model learns, the optimization process, etc. So this is where things like regularization come into play. There's also architecture which affects the training dynamics as well as model capacity. Which makes no guarantee that you get the most information dense representation. Fwiw, the authors did also try distillation.
- _hark 2y agoSorry I wasn't more clear! I'm referring to the Kolmogorov complexity of the network. The OP said: > With enough computation, your neural net weights would converge to some very compressed latent representation of the source code of DOOM. Maybe smaller even than the source code itself? Someone in the field could probably correct me on that. And they're not wrong! An ideally trained network could, in principle, learn the data-generating program, if that program is within its class of representable functions. I might have a NN that naively looks like it takes up GBs of space, but it might actually be parameterizing a much simpler function (hence our ability to prune/compress the weights without performance loss - most of the capacity wasn't being used for any interesting computation). You're right that there's no guarantee that the model finds the most "dense" representation. The goal of regularization is to encourage that, though! All over the place in ML there are bounds like: test loss <= train loss + model complexity Hence minimizing model complexity improves generalization performance. This is a kind of Occam's Razor: the simplest model generalizes best. So the OP is on the right track - we definitely want networks to learn the "underlying" process that explains the data, which in this case would be a latent representation of the source code (well, except that doesn't really make sense since you'd need the whole rest of the compute stack that code runs on - the neural net has no external resources/embodied complexity it calls, unlike the source code which gets to rely on drivers, hardware, operating systems, etc.)
- basch 2y agoSimilarly, you could run a very very simple game engine, that outputs little more than a low resolution wireframe, and upscale it. Put all of the effort into game mechanics and none into visual quality. I would expect something in this realm to be a little better at not being visually inconsistent when you look away and look back. A red monster turning into a blue friendly etc.
- slashdave 2y ago> where text prompts are enough to create a fun new game! Not really. This is a reproduction of the first level of Doom. Nothing original is being created.
- radarsat1 2y agoMost games are conditioned on text, it's just that we call it "source code" :). (Jk of course I know what you mean, but you can seriously see text prompts as compressed forms of programming that leverage the model's prior knowledge)
- troupo 2y ago> one can imagine a near future where text prompts are enough to create a fun new game Sit down and write down a text prompt for a "fun new game". You can start with something relatively simple like a Mario-like platformer. By page 300, when you're about halfway through describing what you mean, you might understand why this is wishful thinking
- reverius42 2y agoIf it can be trained on (many) existing games, then it might work similarly to how you don't need to describe every possible detail of a generated image in order to get something that looks like what you're asking for (and looks like a plausible image for the underspecified parts).
- troupo 2y agoThings that might work plausible in a static image will not look plausible when things are moving, especially in the game. Also: https://news.ycombinator.com/item?id=41376722 https://news.ycombinator.com/item?id=41376722 Also: define "fun" and "new" in a "simple text prompt". Current image generators suck at properly reflecting what you want exactly, because they regurgitate existing things and styles.
- SomewhatLikely 2y agoVideo games are gonna be wild in the near future. You could have one person talking to a model producing something that's on par with a AAA title from today. Imagine the 2d sidescroller boom on Steam but with immersive photorealistic 3d games with hyper-realistic physics (water flow, fire that spreads, tornados) and full deformability and buildability because the model is pretrained with real world videos. Your game is just a "style" that tweaks some priors on look, settings, and story.
- user432678 2y agoSorry, no offence, but you sound like those EA execs wearing expensive suits and never played a single video game in their entire life. There’s a great documentary on how Half Life was made. Gabe Newell was interviewed by someone asking “why you did that and this, it’s not realistic”, where he answered “because it’s more fun this way, you want realism — just go outside”.
- magicalhippo 2y agoThis got me thinking. Anyone tried using SD or similar to create graphics for the old classic text adventure games?
- deleted 2y ago[deleted]