15 ms·
SHARP, an approach to photorealistic view synthesis from a single image
- brcmthrowaway 10mo agoSo this is the secret sauce behind Cinematic mode. The fake bokeh insanity has reached its climax!
- duskwuff 10mo agoAs well as their "Spatial Scene" mode for lock screen images, which synthesizes a mild parallax effect as you move the phone.
- IlikeKitties 10mo ago[flagged]
- calvinmorrison 10mo agoI understand AI for reasoning, knowledge, etc. I haven't figured out how anyone wants to spend money for this visual and video stuff. It just seems like a bad idea.
- accurrent 10mo agoSimulation. It takes a lot of effort today to bring up simulations in various fields. 3 D programming is very nontrivial and asset development is extremely expensive. If I have a workspace I can take a photo of and just use it to generate a 3d scene I can then use it in simulations to test ideas out. This is particularly useful in robotics and industrial automation already.
- jijijijij 10mo agoI don't see any examples of a 3D scene information usable for simulation. If you want to simulate something hitting a table, you need the whole table (surface) in space, not just some spatial illusion effect extrapolated from an image of a table. I also think modelling the 3D objects for simulation is the least expensive part of an simulation... the simulation is the expensive thing. I doubt this will be useful for robotics or industrial automation, where you need an actual spatial, or functional understanding of the object/environment.
- accurrent 10mo agoWith research like this you need to start somewhere. The fact we can get 3d information helps. There are people looking into making splats capture collision information [1]. I have worked on simulation and in my day job do a lot of simulation. While physics is oftem hard and expensive you only need to write the code once. Assets? You need to comission 3d artists and then spend hours wrangling file formats. Its extremely tedious. If we could take a photo and extract meshes Im sure we'd have a much easier time. [1] https://trianglesplatting.github.io/ https://trianglesplatting.github.io/
- re-thc 10mo agoDo people not spend on entertainment? Commercials? It's probably less of a bad idea than knowledge. AI giving a bad visual has less negatives than giving the wrong knowledge leading to the wrong decision.
- rv3392 10mo agoThis specific paper is pretty different to the kind of photo/video generation that has been hyped up in recent years. In this case, I think this might be what they're using for the iOS spatial wallpaper feature, which is arguably useless but is definitely an aesthetic differentiator to Android devices. So, it's indirectly making money.
- netsharc 10mo agoPhoto apps on phones (can you still call them cameras?) already have a lot of "AI" to enhance photos and videos taken. Some of it is technological necessity, since you're capturing something through a tiny hole, a lot of it is sexying it up to appeal to people, because hey, people would prefer a cinema-quality depiction of their memories rather than the reality...
- yodon 10mo ago> photorealistic 3D representation from a single photograph in less than a second
- arjie 10mo agoThis is incredibly cool. It's interesting how it fails in the section where you need to in-paint. SVC seems to do that better than all the rest, though not anywhere close to the photorealism of this model. Is there a similar flow but to transform either a video/photo/NeRF of a scene into a tighter, minimal polygon approximation of it. The reason I ask is that it would make some things really cool. To make my baby monitor mount I had to knock out the calipers and measure the pins and this and that, but if I could take a couple of photos and iterate in software that would be sick.
- necovek 10mo agoYou'd still need one real measurement at least: this might get proportions right if background can be clearly separated, but the absolute size of an object can be worlds apart.
- arjie 10mo agoThat's true. And there's lens correction and all that, but it would be nice to accelerate the CAD modeling.
- Geee 10mo agoThis is great for turning a photo into a dynamic-IPD stereo pair + allows some head movement in VR.
- SequoiaHope 10mo agoAh and the dynamic IPD component preserves scale?
- benatkin 10mo agoThat is really impressive. However, it was a bit confusing at first because in the koala example at the top, the zoomed in area is only slightly bigger than the source area. I wonder why they didn't make it 2-3x as big in both axes like they did with the others.
- yodon 10mo agoSee also Spaitial[0] which announced today full 3D environment generation from a single image [0]https://www.spaitial.ai/ https://www.spaitial.ai/
- andsoitis 10mo agoWhy are all their examples of rooms? Why no landscape or underwater scenes or something in space, etc.?
- jaccola 10mo agoConstrained environments are much simpler. I believe this company is doing image (or text) -> off the shelf image model to generate more views -> some variant of gaussian splatting. So they aren't really "generating" the world as one might imagine.
- boguscoder 10mo agoRequires email to view anything, that’s sad
- dag11 10mo agoI'm confused, does it actually generate environments from photographs? I can't view the galleries since I didn't sign up for emails but all of the gallery thumbnails are AI, not photos.
- jrflowers 10mo ago> I'm confused, does it actually generate environments from photographs? It’s a website that collects people’s email addresses
- avaer 10mo agoThe best I've seen so far is Marble from World Labs, though that gives you a full 360 environment and takes several minutes to do so.
- superfish 10mo ago"Unsplash > Gen3C > The fly video" is nightmare fuel. View at your own risk: https://apple.github.io/ml-sharp/video_selections/Unsplash/gen3c_aligned/-6ebJNtXtWs_0000-0001.mp4 https://apple.github.io/ml-sharp/video_selections/Unsplash/g...
- ghurtado 10mo agoSeth Brundle has entered the chat.
- Traubenfuchs 10mo agoEarly AI „everything turns into dog heads“ vibes. Beautiful.
- drcongo 10mo agoI miss those. Anyone know if it's still possible to get the models etc. needed to generate them?
- Traubenfuchs 10mo agoI wish there was an archive of all those melty dreamscapes. https://m.youtube.com/watch?v=DgPaCWJL7XI&t=1s&pp=2AEBkAIB0gcJCR4Bo7VqN5tD https://m.youtube.com/watch?v=DgPaCWJL7XI&t=1s&pp=2AEBkAIB0g... https://www.youtube.com/watch?v=X0oSKFUnEXc https://www.youtube.com/watch?v=X0oSKFUnEXc
- StilesCrisis 10mo agoGoogle was using them as wall mural artwork in one of the Sunnyvale offices. Very trippy.
- what-the-grump 10mo agoAll this work to recreate a WinAmp viz from 20 years ago :) ?
- tecleandor 10mo ago
- harhargange 10mo agoTMPI looks just as good if not better.
- jjcm 10mo agoDisagree - look at the sky in the seaweed shot. It doesn't quite get the depth right in anything, and the edges of things look off.
- shwaj 10mo agoAgreed. The head of the fly also seems to have weird depth.
- wfme 10mo agoHave a look through the rest of the images. TMPI has some pretty obvious shortcomings in a lot of them. 1. Sky looks jank 2. Blurry/warped behind the horse 3. The head seems to move a lot more than the body. You could argue that this one is desirable 4. Bit of warping and ghosting around the edges of the flowers. Particularly noticeable towards the top of the image. 5. Very minor but the flowers move as if they aren't attached to the wall
- ballpug 10mo ago[dead]
- tartoran 10mo agoImpressive but something doesn't feel right to me.. Possibly too much sharpness, possibly a mix of cliches, all amplified at once.
- a3w 10mo agoFor me, TMPI and SHARP look great. TMPI is consistently brighter, though, with me having no clue which is more correct.
- remh 10mo agoEnhance! https://www.youtube.com/watch?v=LhF_56SxrGk https://www.youtube.com/watch?v=LhF_56SxrGk
- mvandermeulen 10mo agoI thought this was going to be the Super Troopers version
- moondev 10mo agocuda gpu only https://github.com/apple/ml-sharp#rendering-trajectories-cuda-gpu-only https://github.com/apple/ml-sharp#rendering-trajectories-cud...
- matthewmacleod 10mo agoThis is specifically only for video rendering. The model itself works across GPU, CPU, and MPS.
- delis-thumbs-7e 10mo agoInterestingly Apple’s own models don’t work on MPS. Well, I guess you just have to wait for few years..
- diimdeep 10mo agoNo, model works without CUDA then you have .ply that you can drop into gaussian splatter viewer like https://sparkjs.dev/examples/#editor https://sparkjs.dev/examples/#editor CUDA is needed to render side scrolling video, but there is many ways to do other things with result.
- rcarmo 10mo agoFixed that: https://github.com/rcarmo/ml-sharp https://github.com/rcarmo/ml-sharp
- gs17 10mo agoThe gaussian splat output can be generated with CPU (this was honestly one of the easiest AI repos to get running).
- Leptonmaniac 10mo agoCan someone ELI5 what this does? I read the abstract and tried to find differences in the provided examples, but I don't understand (and don't see) what the "photorealistic" part is.
- eloisius 10mo agoFrom a single picture it infers a hidden 3D representation, from which you can produce photorealistic images from slightly different vantage points (novel views).
- avaer 10mo agoThere's nothing "hidden" about the 3d represenation. It's a point cloud (in meters) with colors, and a guess at the the "camera" that produced it. (I am oversimplifying).
- eloisius 10mo agoHidden in the sense of neural net layers. I mean intermediary representation.
- avaer 10mo agoRight. I just want to emphasize that this is not a NERF where the model magically produces an image from an angle and then you ask "ok but how did you get this?" and it throws up its hands and says "I dunno, I ran some math and I got this image" :D.
- uh_uh 10mo ago"Hidden" or "latent" in a context like this just means variables that the algo is trying to infer because it doesn't have direct access to them.
- ares623 10mo agoTakes a 2D image and allows you to simulate moving the angle of the camera with correct-ish parallax effect and proper subject isolation (seems to be able to handle multiple subjects in the same scene as well) I guess this is what they use for the portrait mode effects.
- avaer 10mo agoIs there a link with some sample gaussian splat files coming from this model? I couldn't find it. Without that that it's hard to tell how cherry-picked the NVS video samples are. EDIT: I did it myself, if anyone wants to check out the result (caveat, n=1): https://github.com/avaer/ml-sharp-example https://github.com/avaer/ml-sharp-example
- derleyici 10mo agoApple's Spatial Scene in the Photos app shows similar behavior, turning a single photo into a small 3D scene that you can view by tilting the phone. Demo here: https://files.catbox.moe/93w7rw.mov https://files.catbox.moe/93w7rw.mov
- Traubenfuchs 10mo agoIt‘s awful and often creates a blurry mess in the imaginated space behind the object. Photoshop content aware fill could do equally or better many years ago.
- diimdeep 10mo agoWorks great, model file is 2.8 GB, on M2 rendering took a few seconds, result is guassian .ply file but repo implementation requires CUDA card to render video, I have used one of webgl live renderers from here https://github.com/scier/MetalSplatter?tab=readme-ov-file#resources https://github.com/scier/MetalSplatter?tab=readme-ov-file#re...
- Dumbledumb 10mo agoIn Chapter D.7 they describe: "The complex reflection in water is interpreted by the network as a distant mountain, therefore the water surface is broken." This is really interesting to me because the model would have to encode the reflection as both the depth of the reflecting surface (for texture, scattering etc) as well as the "real depth" of the reflected object. The examples in Figure 11 and 12 already look amazing. Long tail problems indeed.
- deleted 10mo ago[deleted]
- yieldcrv 10mo agoI want to see with people
- BoredPositron 10mo agoThe paper is just a word salad and it's not better than previous sota? I might be missing a key element here.
- codebyprakash 10mo agoQuite cool!
- supermatt 10mo agoI note the lack of human portraits in the example cases. My experience with all these solutions to date (including whatever apple are currently using) is that when viewed stereoscopically the people end up looking like 2d cutouts against the background. I haven't seen this particular model in use stereoscopically so I can't comment as to its effectiveness, but the lack of a human face in the example set is likely a bit of a tell. Granted they do call it "Monocular View Synthesis", but i'm unclear as to what its accuracy or real-world use would be if you cant combine 2 views to form a convincing stereo pair.
- sorenjan 10mo agoThey're using their Depth Pro model for depth estimation, and that seems to do faces really well. https://github.com/apple/ml-depth-pro https://github.com/apple/ml-depth-pro https://learnopencv.com/depth-pro-monocular-metric-depth/ https://learnopencv.com/depth-pro-monocular-metric-depth/
- supermatt 10mo agoIm not sure how the depth estimation alone translates into the view synthesis, but the current implementation on-device is definitely not convincing for literally any portrait photographs I have seen. True stereoscopic captures are convincing statically, but don't provide the parallax.
- sorenjan 10mo agoGood monocular depth estimation is crucial if you want to make a 3D representation from a single image. Ordinarily you have images from several camera poses and can create the gaussian splats using triangulation, with a single image you have to guess z position for them.
- Someone 10mo agoFor selfies, I think iPhones with Face ID use the TrueDepth camera hardware to measure Z position. That’s not full camera resolution, but it will definitely help.
- pmontra 10mo agoSo Deckard got lucky that the picture enhancement machine allucinated the correct clue? But that was boundto happen 6 years ago, no AI yet.
- nashashmi 10mo agoI could not find any mention of it but does this use regenerative AI? I can’t imagine it able to accomplish anything like this without using a large graphical Model in the back.
- rcarmo 10mo agoWell, I got _something_ to work on Apple Silicon: https://github.com/rcarmo/ml-sharp https://github.com/rcarmo/ml-sharp (has a little demo GIF) I am looking at ways to approximate Gaussian splats without having to reinvent the wheel, but I'm a bit over my depth since I haven't been playing a lot of attention to those in general.
- 7moritz7 10mo agoThe example doesn't look particularly impressive to say the least. Look at the bottom 20 %
- rcarmo 10mo agoI just refactored the rendering and resampling approach. Took me a few tries to figure out how to remove the banding masks from the layers, but with more stacked layers and a bit of GPT-foo to figure out the API it sort of works now (updated the GIF) Keep in mind that this is not Gaussian splat rendering but just a hacked approximation--on my NVIDIA machine that looks way smoother.
- deleted 10mo ago[deleted]
- esperent 10mo agoI'm quite delighted that the gif banding artefacts make it look life the photi of a fire is flickering, and also highly impressed that the AI was able to recognize the fire as a photo within a photo and keep it in 2d.
- orthoxerox 10mo agoThe resulting animations feel more like "Live2D" than 3D.
- mhalle 10mo agoIt would be interesting to see how much better this algorithm would be with a stereo pair as input. Not only do many VR and AR systems acquire stereo, we have historical collections of stereo views in many libraries and museums.
- pluralmonad 10mo agoThis seems like what they have been doing with album covers on applemusic for a couple years.
- reactordev 10mo agoThis would be really fun to create stereoscopic videos with. Take a video input, offset x+0.5 or some coefficient, take the output, put them side by side (or interlaced for shutter glasses) and viola! 3D movies.
- alexgotoi 10mo agoApple dropping this is interesting. They've been quiet on the flashy AI stuff while everyone else is yelling about transformers, but 3D reconstruction from single images is actually useful hardware integration stuff. What's weird is we're getting better at faking 3D from 2D than we are at just... capturing actual 3D data. Like we have LiDAR in phones already, but it's easier to neural-net your way around it than deal with the sensor data properly. Five years from now we'll probably look back at this as the moment spatial computing stopped being about hardware and became mostly inference. Not sure if that's good or bad tbh. Will include this one in my https://hackernewsai.com/ https://hackernewsai.com/ newsletter.
- momojo 10mo agoI wonder if humans are any different. We don't have LIDAR in our eyes but we approximate depth "enough" with only our 2D input
- dTal 10mo agoWe also constantly move our heads and refocus our eyes. We can get a rough idea of depth from only a static stereo pair, but in reality we ingest vastly more information than that and constantly update our internal representation in real time.
- jakefromstatecs 10mo agoWe don't have 2d input, we have 3d input. We have two eyes that gives us depth by default.
- stronglikedan 10mo agoThat's cool and all, but it seems like only the first step in this, where they go from 2D photo all the way to fully animated (animatable?) characters: https://www.youtube.com/watch?v=DSRrSO7QhXY https://www.youtube.com/watch?v=DSRrSO7QhXY
- somethingsome 10mo agoLast time I tried depth pro it was not really metric, I wonder if this one is as they claim. If someone has some experience on that side I would be interested