3 ms·
Trust Google to look at their c. 1993 gouraud shaded, untextured renderings without shadows and conclude "let's use deep learning to make it look realistic!"
by nobbis 9y ago
Trust Google to look at their c. 1993 gouraud shaded, untextured renderings without shadows and conclude "let's use deep learning to make it look realistic!"
- alexbeloi 9y agoICCV 2017 this year had dozens of papers using GANs and other methods to make simulated images/videos look more realistic. The hunger for cheap usable data is real.
- kalal 9y agoCan you provide some links? this sounds interesting ...
- alexbeloi 9y agoI spoke in hyperbole, there were dozens of papers using generative networks to create real looking simulated data which is the more common approach. You start with noise + some features you want in your image, then you turn that into a realistic image. The notable exception is if you want to generate segmentation data (where you want labels for each pixel), in this case the set of features you want is essentially an image already. ICCV 2017 * Photographic Image Synthesis with Cascaded Refinement Networks: https://arxiv.org/pdf/1707.09405.pdf https://arxiv.org/pdf/1707.09405.pdf * Synthetic avatars: https://arxiv.org/pdf/1704.05693.pdf https://arxiv.org/pdf/1704.05693.pdf * Synthesizing cell microscopy data: https://arxiv.org/pdf/1708.04692.pdf https://arxiv.org/pdf/1708.04692.pdf * Material Editing Using a Physically Based Rendering Network: https://arxiv.org/pdf/1708.00106.pdf https://arxiv.org/pdf/1708.00106.pdf * 3D-PRNN: Generating Shape Primitives With Recurrent Neural Networks: * for more search 'generative' 'generating' 'simulating' 'synthesis' 'synthetic' on: http://iccv2017.thecvf.com/program/main_conference#schedule http://iccv2017.thecvf.com/program/main_conference#schedule Also worth mentioning: CVPR 2017 best paper, this is the other exception, it's unique because it doesn't require _any_ labeled data * Learning from Simulated and Unsupervised Images through Adversarial Training: https://arxiv.org/abs/1612.07828 https://arxiv.org/abs/1612.07828
- tehsauce 9y agoThis seems like an unreasonable approach given the quality of their graphics. They might as well train an GAN to replace the entire rendering pipeline if they are using just simple images! Train it to generate images directly from pose information. They really should focus on improving their rendering beyond a 1993 level if they want a robust solution.
- candiodari 9y agoThat is too indirect a problem. Neural networks will find answers to problems where there is a relatively direct hierarchical connection between input and output. Symbolic information would require different and very big architectures. Also doing that is still state-of-the-art and you don't try to solve 2 state-of-the-art problems at the same time. Firstly, using the better rendering approach has been tried and failed. These networks exploit things that perfect rendering has but reality doesn't. For instance, in a rendered simulated environment it is perfectly reasonable to estimate exactly where an edge is and then follow that edge with zero margin with the robot. In reality you need a minimum margin if you expect it to work (or your object will be damaged), and a decent one. Photorealistic rendering tries to emulate optics that are incredibly good, with no distortion, and the robots have cheap webcams with plastic lenses, with both consistent distortion near the edges and random distortions everywhere because they're webcams. Not supercheap ones, but this isn't a RED camera either. You could fix those lenses, but why would you ? It would be more valuable if it works with cheap lenses, and human and animal eyes don't really try to be distortion free. Rather they have the "software" behind it compensate. Second, photorealistic rendering requires modeling accuracy, and material modeling on a scale we just can't do for real datasets. These datasets have millions of objects, and to make the robot more tolerant of the general weirdness of the real world, they are randomly perturbed, to turn millions of examples to billions. But those perturbed objects are not photorealistic anymore. Nobody has a toothbrush with a triangle in the bottom sticking out to one side. Besides, keep in mind that the resolutions used here are quite small. It's not like they're trying to run a renderer at 4k resolution. Think 128x128, or 256x256. Not exactly a big issue.
- nobbis 9y ago