7 ms·
This is pretty remarkable. So these really do learn humanly interpretable representations and not only doing some magic in the billion dimensional hyperplane th
by dkarras 3y ago
This is pretty remarkable. So these really do learn humanly interpretable representations and not only doing some magic in the billion dimensional hyperplane that we can't hope of deciphering.
- corysama 3y agoAs an old 3D graphics engineer, the fact that albedo is in there is as just striking as it should be expected. The core components of physically based rendering are position (derivable from image XY and depth), surface normal, incoming light, and at least albedo + one of a few variations on surface material properties such as specularity and roughness. That the AI is modeling depth is pretty expected. Modeling surface normal is a nice local convolution of depth. But, modeling albedo separated from incoming light is great. I wonder if specularity is hiding in there too.
- hackerlight 3y agoIt's a good depth map, too. Better than other tools I've seen which require fiddling with lots of knobs to get a good result. Might be useful for textures for parallax mapping.
- viraptor 3y agoI'm mostly ok with the idea that the text-to-image models do this. What blows me away is that the autoregressive model does it too. We're not even asking it to produce a specific object which has a meaning or context assigned. The network learned to produce those representations as just "this is a useful decomposition if I want to recreate the same thing again". The jump from image compression to figuring out how to represent objects in space is insane. And makes me want to feed some optical illusions through that network - what's the depth map of the infinite staircase?
- theygotit 3y ago[dead]
- int_19h 3y agoI find it amazing that, for all the evidence we have of generative models having some fairly complex internal model of the world, people still insist that they are mere "stochastic parrots" who "don't really understand anything".
- mewpmewp2 3y agoFrom the other side however if you really think about it, our understanding of everything must be stochastic as well. So perhaps this sort of thing yields in many complexities that we are not aware of. How would I know I am not a stochaistic parrot of some sort. I am just neural nets traine on current envrionment while the base model that is dependent on DNA, through evolution and natural selection of the fittest. Same as currently competing LLMs where the best one will win out.
- int_19h 3y agoYou're not wrong, but the "stochastic parrot" claim always comes with another, implicit one that we are not like that; that there's some fundamental difference, somehow, even if it is not externally observable. Chinese room etc. In short, it's the good old religious debate about souls, just wrapped in techno-philosophical trappings.
- cthalupa 3y agoYou could leverage the exact same accusation against the other side - we know fundamentally how the math works on these things, yet somehow throw enough parameters at them and eventually there's some undefined emergent behavior that results in something more. What that something more is is even less defined with even fewer theories as to what it is than there are around the woo and mysticism of human intelligence. And as LarsDu88 points out in a separate thread, there are alternative explanations for what we're seeing here besides "We've created some sort of weird internal 3D engine that the diffusion models use for generating stuff," which also meshes closely with the fact that generations routinely have multiple perspectives and other errors that wouldn't exist if they modeled the world some people are suggesting. If there's something more going on here, we're going to need some real explanations instead of things that can be explained multiple other ways before I'm going to take it seriously, at least.