4 ms·
This seems like an unreasonable approach given the quality of their graphics. They might as well train an GAN to replace the entire rendering pipeline if they a
by tehsauce 9y ago
This seems like an unreasonable approach given the quality of their graphics. They might as well train an GAN to replace the entire rendering pipeline if they are using just simple images! Train it to generate images directly from pose information. They really should focus on improving their rendering beyond a 1993 level if they want a robust solution.
- candiodari 9y agoThat is too indirect a problem. Neural networks will find answers to problems where there is a relatively direct hierarchical connection between input and output. Symbolic information would require different and very big architectures. Also doing that is still state-of-the-art and you don't try to solve 2 state-of-the-art problems at the same time. Firstly, using the better rendering approach has been tried and failed. These networks exploit things that perfect rendering has but reality doesn't. For instance, in a rendered simulated environment it is perfectly reasonable to estimate exactly where an edge is and then follow that edge with zero margin with the robot. In reality you need a minimum margin if you expect it to work (or your object will be damaged), and a decent one. Photorealistic rendering tries to emulate optics that are incredibly good, with no distortion, and the robots have cheap webcams with plastic lenses, with both consistent distortion near the edges and random distortions everywhere because they're webcams. Not supercheap ones, but this isn't a RED camera either. You could fix those lenses, but why would you ? It would be more valuable if it works with cheap lenses, and human and animal eyes don't really try to be distortion free. Rather they have the "software" behind it compensate. Second, photorealistic rendering requires modeling accuracy, and material modeling on a scale we just can't do for real datasets. These datasets have millions of objects, and to make the robot more tolerant of the general weirdness of the real world, they are randomly perturbed, to turn millions of examples to billions. But those perturbed objects are not photorealistic anymore. Nobody has a toothbrush with a triangle in the bottom sticking out to one side. Besides, keep in mind that the resolutions used here are quite small. It's not like they're trying to run a renderer at 4k resolution. Think 128x128, or 256x256. Not exactly a big issue.
- nobbis 9y agoPhotorealism (or "better rendering") doesn't mean perfect imaging. You can model the radial distortion, rolling shutter, Bayer filtering of cheap webcams in 5 lines of code. A better simulation is closer to reality, not closer to an ideal.
- gugagore 9y agoA "minimum margin" is needed anyway, even with perfect perception, because in general you don't have perfect actuation.
- zardo 9y ago>They really should focus on improving their rendering beyond a 1993 level if they want a robust solution. Recent work from openai suggests that the 1993 graphics with randomized textures may be good enough, and the push for realistic graphics may be entirely unnecessary to get a robot capable of generalizing it's learning to the real world.
- state_less 9y agoI'm interested to know which simulation feature provides the most information. Sort of a random drop out of features on the scene graph. Include lighting vs. no lighting, color vs monochrome (brightness only), shadowing vs non, texture vs untextured, monocular vs binocular, etc... You can test what simulation features are informative by measuring the resulting real world accuracy with the arm. How simple can the simulation be? How informative is each feature?
- nobbis 9y agoIt's the opposite. Simple rendering makes pose estimation harder, not easier - texture and shadows resolve depth ambiguity. So it's easier to show you've "closed the simulation-to-reality gap."
- darkmighty 9y ago(A note from elsewhere) A reminder that a state of the art GPU renderer can produce photorealistic scenes at 4k resolution (about 8000000 pixels) at 120 fps (i.e. render the scene in less than 8ms). In other words, this performance is completely inaccessible to current DNN architectures (and I suspect will never be as efficient). As another algorithmic illustration, we have already near-optimal algorithms for sorting large lists of random numbers. I guarantee a large neural network even trained for massive lengths of time cannot beat the sheer efficiency of traditional algorithms. It should approach the asymptotics of the optimal algorithm, but there's too much overhead in how DNNs are structured. A better task, I believe, is to use DNNs to write code that improves renderers, shaders, or sorting algorithms while maintaining image quality (in the case of computer graphics) or correctness (in the case of sorting algorithms). --- In this case the problem could be coming up with a variety of good realistic textures. I'm sure Google could hire a few artists to do a fantastic job, but if you want overkill I guess a GAN could generate the textures (to be rendered by a conventional algorithm).