32 ms·
Nvidia Research Turns 2D Photos into 3D Scenes
- sennight 5y agoI know that taste in comedy is seasonal (yes, there were a people in a time that thought vaudeville was the cat's pajamas), but has anyone ever greeted a pun with anything other than a pained sigh?
- cogman10 5y agoPuns aren't to make people laugh, the pained sigh is the point. It's schadenfreude for the person making the pun.
- sennight 5y ago> It's schadenfreude for the person making the pun. Nah, if it is a joke at their own expense then it is "self deprecating humor", something which is definitely designed to get a laugh. Humiliation fetish, maybe? Obviously nothing is funny past a certain point of deconstruction... especially if you find yourself defending the distinguishing difference of the "meta". Just stop making puns, easy.
- noduerme 5y agoIt's ones like this that make me shake my head and go "Aiaiai."
- ModernMech 5y agoWatch Bob's Burgers. The whole show is basically puns. I chuckle.
- mkaic 5y agoIdk, personally I find wordplay quite punny -- though I almost always try to greet the person who made the pun, and not the pun itself (they're abstract, inanimate concepts, pretty difficult to say "hello!" to) :P
- ksec 5y agoNvidia is really turning into an AI powerhouse. The moat around CUDA, and how those target customer aren't as stringent about budget, especially when the hardware cost is tiny compare to what they do. I wonder if they could reach a trillion market cap.
- bloodyplonker22 5y agoit's not a matter of "if", but "when" they will reach a trillion dollar market cap.
- ksec 5y agoWell I think that is a little optimistic in the near terms. Considering there has never been a Semi-Conductor player reaching that milestone without a Fab. So Nvidia will be a first. At current P/E of 70, Semiconductor industry average is only ~30. Realistically Nvidia will need to triple their revenue at a fair P/E. I could see that in Data Center, reaching $30B revenue per year within next ~5 years. And this is already larger than Intel's DataCenter record revenue. We are still $30B short coming from Gaming and Professional Visualisation. Intel already has their GPU play ready in 2022. ( Assuming it is competitive. ). i.e Nvidia will need to find another massive market to conquer to reach that Trillion Market Cap status.
- Animats 5y agoWhat's new about this? That it's faster? People have been reconstructing 3D images from multiple photos for over a decade. The experimental work today is constructing a 3D image from a single photo, using a neural net to fill in a reasonable model of the stuff you can't see.
- mkaic 5y agoit's not just faster, it's extremely faster. They're achieving results that are better than SOTA in a fraction of the time. Wildly impressive work.
- csee 5y agoThis type of thing looks like the future of Meta or even Zoom to me.
- idworks1 5y agoFive years ago, I've used common software to do this. I had to take hundreds of pictures of a scene, getting as many angles and details possible. Then when you pass it to the computer. Stitching it all together took well beyond 24 hours. Now that I had a 3d model of the scene, I had to spend countless hours cleaning it up to make sure it was useable. Maybe in the last 5 years, things have improved. But this demo used 4 pictures. And apparently, it rendered the final image in seconds. That's what's new.
- matsemann 5y agoIf I understand it correctly it didn't make a 3d model, though. So you can't extract and reuse the result. Only move around in it and it creates an image for that viewpoint. But no meshes or textures.
- shultays 5y agoDid it really use only 4 pictures? Do you have a source? https://news.ycombinator.com/item?id=30810885 https://news.ycombinator.com/item?id=30810885
- daenz 5y ago>The model requires just seconds to train on a few dozen still photos — plus data on the camera angles they were taken from — and can then render the resulting 3D scene within tens of milliseconds. Generating the novel viewpoints is almost fast enough for VR, assuming you're tethered to a desktop computer with whatever GPUs they're using (probably the best setup possible). The holy grail (from my estimation) is getting both the training and the rendering to fit into a VR frame budget. They'll probably achieve it soon with some very clever techniques that only require differential re-training as the scene changes. The result will be a VR experience with live people and objects that feels photorealistic, because it essentially is based on real photos.
- simsla 5y ago> plus data on the camera angles they were taken from Doesn't seem like much of a stretch to determine the angles as well. E.g. a semi brute forced way with GANs
- c4wrd 5y agoI've spent a lot of time thinking about this (i.e. taking a video and creating a 3D scene) and I don't think that it is feasible in most cases to have good accuracy. If you need to infer the angle, you need make a lot of biased assumptions about things like velocity, position, etc., of the camera and even if you were 99.9% accurate, that 0.1% inaccuracy is compounded over time. Now I'm not saying it's not possible, but I'd believe that if you want an accurate 3D scene, you'd rather be spending your computation budget on things other than determining those angles when it can be simply be provided by hardware.
- krasin 5y agohttps://github.com/NVLabs/instant-ngp https://github.com/NVLabs/instant-ngp has a script that converts a video into frames and then uses COLMAP ([1]) to compute camera poses. You can then train a NeRF model within a few seconds. It all works pretty well. Trying it on your own video is pretty straightforward. 1. https://colmap.github.io/ https://colmap.github.io/
- alanwreath 5y agoI’m probably going to ramp up the number of photos I take in hope that google photos auto applies this tech
- jamiek88 5y agoI’m gonna do the same particularly dog pics. I’d love to bring my old dogs back to virtual life. Have old spanky (died a decade ago) running with lilybean (current version) in iMovie.
- dylan604 5y agoGoogle has done some expirements from people taking very similarly framed images from different times to create timelapse videos, so that would hint to me they definitely have the content to try. There was an article here some time back that showed streets in NYC that used this kind of idea of using older photographs to put one in street view of older NYC. So, yeah, I'm guessing it could be done. Might be weird with the different quality of images (modern digital, polaroids, kodachrome, etc).
- woah 5y agoAre there examples of this being used on large outdoor spaces?
- krasin 5y agoYes, Waymo did the whole San Francisco block: https://waymo.com/research/block-nerf/ https://waymo.com/research/block-nerf/
- gundmc 5y agoWoah, this video is way more interesting than the Nvidia polaroid teaser in the original link.
- krasin 5y agoStill, NVIDIA's achievement (and Thomas Müller in particular) is amazing. Thomas and his collaborators achieved an almost 1000x performance improvement, by a combination of algorithmical and implementation tricks. I highly recommend trying this at home: https://nvlabs.github.io/instant-ngp/ https://nvlabs.github.io/instant-ngp/ https://github.com/NVlabs/instant-ngp https://github.com/NVlabs/instant-ngp Very straightforward and gives better insight into what NeRF is than any shiny marketing demo.
- cinntaile 5y agoWaymo needed 2.8 million images to create that scene, I wonder how many Nvidia would need? Or was the focus only on speed? I skimmed the article and didn't really find info on that.
- krasin 5y agoWaymo essentially trained several NeRF models for Block-NeRF that are rendered together. It's conceivable that NVIDIA's instant-ngp could be used for that.
- noduerme 5y ago
- bogwog 5y agoNvidia is leaving us all behind
- elil17 5y agoMy prediction/hope is that NeRFs will totally revolutionize how the film/TV industry. I can imagine: - Shooting a movie from a few cameras, creating a movie version of a NeRF using those angles, and then dynamically adding in other shots in post - Using lighting and depth information embedded in NeRFs to assist in lighting/integrating CG elements - Using NeRFs to generate virtual sets on LED walls (like those on The Mandalorian) from just a couple of photos of a location or a couple of renders of a scene (currently, the sets have to be built in a game engine and optimized for real time performance).
- andy_ppp 5y agoComputer games, VR and AR could also be pretty amazing uses for this technique too.
- teaearlgraycold 5y agoRIP photo-realistic modelers
- tomatowurst 5y agohmmm well I still think they will be in demand for the same reason software developers will be not automated away. NeRF is really mind boggling good but there are still artifacts, and something that modelers have a good eye for. Having said that, it might be the end for any junior type of roles. Same reason that github copilot really takes a bite of the need to have a junior developer. I'm very curious what will happen because it will become a sort of trend across other industries apart from legal or medical professions (peace of mind from human-in-the-loop).
- teaearlgraycold 5y agoMaybe we'll have people spend their time building IRL sculptures and spaces to get digitized.
- 5y ago
- anyfactor 5y agoTangent I wonder what happens to most people when they see innovation such as this. Over the years I have seen numerous mind-blowing AI achievement, which essentially feel like miracles. Yet literally after an hour I forget what I even saw. I don't find these innovations to have a lasting impression on me or on the internet except for the times when these solutions are released to the public for tinkering and they end up failing catastrophically. I remember having the same feeling about chatbots and TTS technology literally ages ago, but at present time, the practical use of these innovation feel very mediocre.
- fleischhauf 5y agoI have the impression, that now some of them seem to really end up in some practical applications. Funnily enough someone just today showed me a feature of his phone where you can select some undesired objects in youe photo and it would just replace them with a fitting background indistinguishable from the original photo.
- Groxx 5y agoIt's rather entertaining when this happens in the opposite direction automatically too: https://twitter.com/mitchcohen/status/1476351601862483968 https://twitter.com/mitchcohen/status/1476351601862483968
- Kerrick 5y agoThe followup on that is... it didn't happen. There was a leaf in the foreground, and the depth of field in the photo was large enough that it was in focus rather than blurring. https://twitter.com/mitchcohen/status/1476951534160257026 https://twitter.com/mitchcohen/status/1476951534160257026
- efraim 5y agoIn a follow up tweet it looks like there was a leaf on a tree in the foreground that obscured her face, not an AI replacement.
- 5y ago
- xrd 5y agoThis nerf project is cool too. https://github.com/bmild/nerf https://github.com/bmild/nerf I've been trying to get GANs to do this for a while, but NeRFs look like the perfect fit.
- maybelsyrup 5y agoIs anyone else kinda terrified?
- deleted 5y ago[deleted]
- syspec 5y agoIs there a video of this? I'm not sure what's the connection to the top photo/video/matrix-360-effect Was that created from a few photos? I didn't see any additional imagery below --- Update It looks like these are the four source photos: https://blogs.nvidia.com/wp-content/uploads/2022/03/NVIDIA-Research-Instant-NeRF-Image.jpg https://blogs.nvidia.com/wp-content/uploads/2022/03/NVIDIA-R... Then it creates this 360 video from them: https://blogs.nvidia.com/wp-content/uploads/2022/03/2141864_Instant-NeRF_TEASER_GIF.mp4 https://blogs.nvidia.com/wp-content/uploads/2022/03/2141864_...
- shultays 5y agofour source photos Is it just 4 or are there more? I find it hard to believe there is only 4. There are clearly more data in video https://i2.paste.pics/645fe17e418b2cb1f6179e0b6671a170.png https://i2.paste.pics/645fe17e418b2cb1f6179e0b6671a170.png like back side of camera here (it is kinda visible but much poorer compared to video). Or existence of a 2nd white sheet in background. But correct me if it is only 4 and you have a source on that
- jrib 5y agoJust want to say I appreciate the cleverness of the title.
- deleted 5y ago[deleted]
- PaulHoule 5y agoIf you have a graphics card which is unobtainable.
- astrange 5y agoStill plenty of them for OEMs/prebuilts, it's just people don't want to buy the extra PC around it, right?
- baron816 5y agoIt would be really great to recreate loved ones after they have past in some sort of digital space. As I’ve gotten older, and my parents get older as well, I’ve been thinking more about what my life will be like in old age (and beyond too). I’ve also been thinking what I would want “heaven” to be. Eternal life doesn’t appeal to me much. Imagine living a quadrillion years. Even as a god, that would be miserable. That would be (by my rough estimate) the equivalent of 500 times the cumulative lifespans of all humans who have ever lived. What I would really like is to see my parents and my beloved dog again, decades after they have past (along with any living ones at that time). Being able to see them and speak to them one last time at the end of my life before fading into eternal darkness would be how I would want to go. Anyway, there’s a free startup idea for anyone—recreate loved ones in VR so people can see them again.
- heavyset_go 5y agoI think it would be a special kind of hell to have your resurrected loved ones sell you ads in virtual reality.
- rilezg 5y agoThis reminds me a lot of Black Mirror season 2 episode 1. Always good to treasure the time we are given.
- olladecarne 5y agoThere's so many things we invent with good intentions but in the end go terribly wrong and I think this is one of those things. I think it's ok to mourn and remember the past, but moving on and accepting reality is important to a healthy life. Let's be real though, the startup that makes this but appeals to our worst instincts make bank. I can't imagine how much more messed up future generations will be as we keep making more dangerous technology that appeals to our primal instincts.
- XorNot 5y agoSo the part which makes this interesting to me is the speed. My new desire in our video conferencing world these days has been to have my camera on but running a corrected model of myself so I can sustain apparently eye-contact without needing to look directly at the camera.
- jamiek88 5y agohttps://osxdaily.com/2021/05/12/enable-eye-contact-facetime-iphone-ipad/ https://osxdaily.com/2021/05/12/enable-eye-contact-facetime-...
- aaron695 5y ago
- siavosh 5y agoI'm curious for those that work with NeRFs what their results look like for random images as opposed to the 'nice' ones that are selected for publications/demos.
- danamit 5y agoI am kinda skeptical, AI demos are impressive but the real world results are underwhelming. How much it resources it takes to generate images like that? is this the most ideal situation? Can you take images from the web and based on metadata make a better street view? With all this AI where is one accessible translation service? or even an accent-adjusting service? or just good auto-subtitles?
- visarga 5y agoThis is essentially like a 3D JPEG, but instead of modelling the image with Discrete Cosine Transform (DCT) they use a neural net. So the neural net itself will learn just that one image and be able to reproduce it point by point, from various angles. A whole network for just one example, and an innovative way to look at what a neural network can be.
- sorenjan 5y agoI don't really understand why NeRFs would be particularly useful in more than a few niche cases, perhaps because I don't fully understand what they really are. My impression is that you take a bunch of photos in various places and directions, then you use those as samples of a 3D function that describes the full scene, and optimize a neural network to minimize the difference between the true light field and what's described by the network. An approximation of the actual function, that fits the training data. The millions of coefficients are seen as a black box that somehow describes the scene when combined in a certain way, I guess mapping a camera pose to a rendered image? But why would that be better than some other data structure, like a mesh, a point cloud, or signed distance field, where you have the scene as structured data you can reason about? What happens if you want to animate part of a NeRF, or crop it, or change it in any way? Do you have to throw away all trained coefficients and start again from training data? Can you use this method as a part of a more traditional photogrammetry pipeline and extract the result as a regular mesh? Nvidia seems to suggest that NeRFs are in some way better than meshes, but according to my flawed understanding they just seem unwieldy.
- p1esk 5y agoWhat happens if you want to animate part of a NeRF, or crop it, or change it in any way? Do you have to throw away all trained coefficients and start again from training data? You don’t change NeRF (the model). You change the point of view of an observer.
- wokwokwok 5y agoI mean, this is the parent posts point; the use case for a static photo or a static 3d nerf is pretty limited. With other structured data compositing and animating is relatively trivial. It turns out that people have approached this problem before and you can composite nerf too (1) by sampling different functions over the volume. …but, let’s not pretend. The complaint is entirely valid. You’re taking a high resolution voxel grid and encoding it into a model. Working with simple voxel data let’s you do all kinds of normal image manipulation techniques, and it’s not clear how you would do some of those with a nerf. Practically speaking, the applications you can use this for are therefore reasonably limited right now. [1] - https://www.unite.ai/st-nerf-compositing-and-editing-for-video-synthesis/ https://www.unite.ai/st-nerf-compositing-and-editing-for-vid...
- gareth_untether 5y agoAI and 3D content making is becoming so exciting. Soon we'll have an idea and be able to make it with automated tools. Sure having a deeper undertaking of how 3D works will be beneficial, but will no longer be the entry requirement.
- luckydata 5y agoI'm really looking forward to this technology getting applied to home improvement.
- mkaic 5y agoas someone who works in both AI and filmmaking, I remember losing my mind when this paper was first released a few weeks ago. It's absolute insanity what the folks at Nvidia have managed to accomplish in such a short time. The paper itself[0] is quite dense, but I recommend reading it -- they had to pull some fancy tricks to get performance to be as good as it is! [0]https://nvlabs.github.io/instant-ngp/ https://nvlabs.github.io/instant-ngp/
- ranger_danger 5y ago
- visarga 5y agoNo, really, it's a cool application of neural nets. It was unexpected when it first came up and it took a whole day to learn a scene, but a couple of years later it can be done in seconds. I think the cool part is that a neural net can learn to produce (R,G,B) from (X,Y,Z,angle) by using a clever encoding trick with sin() and cos() for the input coordinates. And the fact that a neural net can be a frigging JPEG in 3D.
- nharada 5y agoYou're missing this work in the context of the field as a whole. Labs have been releasing papers boasting 2-4x speedups and getting them published at conferences, and then this group comes in and speeds up the original by 1000x. That's a huge leap in capability.
- jvanderbot 5y ago> Creating a 3D scene with traditional methods takes hours or longer, depending on the complexity and resolution of the visualization. This just isn't true. I can create a 3D scene from 360-degree photos (even 4) in a minute or so using traditional methods, even open-source toolkits. It doesn't look as good as this because it doesn't have a neural net smoothing the gaps, but it's not true that it takes hours to build 3D information from 2D images.
- jayd16 5y agoI think it's comparing to crunching through many more images to get a comparable quality scene.
- wubbert 5y agoIIRC Microsoft had something like this years ago, but the results weren't nearly as smooth or natural looking. I can't remember what it was called, though.
- schemescape 5y agoPhotosynth?
- xsmasher 5y agoI remember a very old video where they rendered 3d scenes for frames 1 and 5 (examples) but interpolated the frames in between, or something like that, instead of re-rendering the whole scene. Maybe they were only redrawing some triangles if the angle changed too much. If it's that tech, I'm pretty sure it got dropped when 3D acceleration made it feasible to just re-render the whole scene every frame. Dedicated hardware won out over software tricks.
- patientplatypus 5y agoWhat's the current state of research on true volumetric displays? That's what I'm excited for, although that takes less AI and more hardware, so quite a bit more difficult.
- Tenoke 5y agoI hope someone can take this, all the images of street view, recent images of places etc. and create a 3d environment covering as much of earth as possible to be used for an advanced Second Life or other purposes.
- dexter89_kp3 5y agoSelf-driving car companies are already working on this. Check out: https://waymo.com/research/block-nerf/ https://waymo.com/research/block-nerf/ NERF is a very active research area, and the progress from 2020 to now has been nothing sort of astonishing. In 5 years, I expect there to be fully generative NERF's in research i.e describe a scene, and a NN produces a full 3d scene that can you interact with.
- krapp 5y agoCan it be done with Google's existing Street View data?
- yonz 5y agoYes!!!! I always get frustrated when ever it does the weird stretch thing as soon as you move around in street view. Just jump to the next frame of you have to.
- asvitkine 5y agoIn the demo video, they mention that they used a lot of footage from self-driving cars to produce that. One thing I noticed is there are no pedestrians and cars in those scenes. So they must do a lot of work to filter them out by combining a lot of footage. Therefore, it likely can't be used (as-is) on the street view dataset...
- dexter89_kp3 5y agoCan be done yes. Is it the best dataset to do what you describe, no.
- pedalpete 5y ago
- socceroos 5y agoENHANCE. ROTATE. I mean, obviously generated images can't be used as proof in the court of law, but this feels like we're slipping into crummy USA show territory.
- ge96 5y agoComment related to top comment Was talking to someone 2 days ago, just died randomly, early 40's. It's trippy, I have data of this person's face eg. videos/base64 strings... it's eerie. Unanswered texts wondering what's wrong. My thinking is I was only exposed to a part of this person, won't be them fully if reproduced.
- deleted 5y ago[deleted]
- m3at 5y agoThis is great, and the paper+codebase they're referring to (but not linking, here [1]) is neat too. The research is moving fast though, so if you want something almost as fast without specialized CUDA kernels (just plain pytorch) you're in luck: https://github.com/apchenstu/TensoRF https://github.com/apchenstu/TensoRF As a bonus you also get a more compact representation of the scene. [1] https://github.com/NVlabs/instant-ngp https://github.com/NVlabs/instant-ngp
- Mandelmus 5y agoThe key difference is that this TensoRF is discrete (like a JPG) whereas NeRFs are continuous (like vector graphics).
- nl 5y agoNext time someone says "why does everyone in AI use NVidia and CUDA"? this is why. They do high quality research and almost inevitably end up releasing the code and models. It's possible to reconstruct all that as a non-CUDA model, but when you want to use it, why would you when it's going to take months of work to get something that isn't as optimised?
- Zenst 5y agoMy first thoughts seeing this is darn, Facebook will with there metaverse, be drinking this up for content. So much so that my thoughts of, would I be shocked if Facebook/Meta made a play to buy Nvidia! Certainly wouldn't shock me as much now as it would before this given how they are banking upon the metaverse/VR being there new growth divergance, what with the leveling of with current services user base after well over a decade and a half. Certainly though, game franchised films would become a lot more imersive, though I do hope that whole avenue dosn't become sameish with this tech overly learned upon. But one thing for sure, I can't wait to bullet-time the film - The Wizzard of OZ with this tech :).
- preommr 5y agoDude, this isn't 2015 anymore - Nvidia has a higher market cap than Meta. The days of Meta being a giant that can continously buyout other companies to keep stayling alive are coming to an end.
- spyder 5y agoThere is an explosion of NeRF papers: https://github.com/yenchenlin/awesome-NeRF https://github.com/yenchenlin/awesome-NeRF It's possible to capture video / movement to into NeRFs, possible to animate, relight, compose multiple NeRF scenes, and a lot of papers are about making faster more efficient and higher quality NeRF. Looks very promising.
- dharma1 5y agoIn terms of practical use - is there a pipeline to use the NeRF 3D scenes in Unreal Engine? How many photos do you need on average vs photogrammetry? 50% less?
- Synaesthesia 5y agoThis is not a textured polygon asset, it's a neural field, so it's like stored the directional data in a neural network, I think using spherical harmonics.
- dharma1 5y agoThat’s ok, just need some level of integration with UE (being able to /integrate with UE camera etc). Specifically interested in using this for LED green screens that sync to a live camera movement (Mandalorian style). Our pipeline uses UE, Nvidia cards and photogrammetry/plates/3D models atm, this could speed things up a lot and require less photos for creating the 3D backgrounds.
- rawoke083600 5y agoI'm guessing if you can "detect/recognize" an object in 2D space, you could guestimate it's "missing-dimension" i.e depth. If you detect an apple in a photo, you could quite reliably guess how the back look Still very cool :)
- shultays 5y agoIs the example the result of just 4 photos? Or more? Are there any other data available, spatial data attached to photos for example? Why they don't explain the scope of achievement properly? edit: I don't think it is just 4 https://news.ycombinator.com/item?id=30810885 https://news.ycombinator.com/item?id=30810885
- shultays 5y agoActually it is explained in the article, I somehow missed it The model requires just seconds to train on a few dozen still photos — plus data on the camera angles they were taken from — and can then render the resulting 3D scene within tens of milliseconds Pretty impressive, but lesser compared to generating it from 4 photos (which imho the movie suggests). Which would be "real magic level of impressiveness" for me
- tirrex 5y agoAre they using four photos or more?
- shultays 5y agoI don't think so but I don't have a source. Please someone correct me if I am wrong https://news.ycombinator.com/item?id=30810885 https://news.ycombinator.com/item?id=30810885
- xsmasher 5y agoI'm guessing at least six photos, judging from the six tripods. (They could be stands for lights, but like I said, just guessing.)
- speedcoder 5y agoWould this make "I Dreamed a Dream" from Les Miserables less moving https://youtu.be/RxPZh4AnWyk https://youtu.be/RxPZh4AnWyk ?
- MrYellowP 5y ago> Blink of an AI I know this post adds nothing, but that one's well worth being pointed out.