7 ms·
Depth Pro: Sharp monocular metric depth in less than a second
- FactKnower69 2y ago300ms for inference is "fast" now? not even remotely usable in realtime
- habitue 2y agoDoes apple actually name their research papers "Pro" too? Like is there an iLearning paper out there?
- sockaddr 2y agoSo what happens in the far future when we send autonomous machines equipped with models trained on Earth life and structures to other planets, are they going to have a hard time detecting and measuring things? What happens when the model is tasked with detecting the depth of an object that’s made of triangular glowing scales and the head has three eyes?
- adolph 2y agoAssembly Theory
- deleted 2y ago[deleted]
- ClassyJacket 2y agoIt would be fairly easy to construct something like this it's never seen before to test it. Go try it. Cut some shapes out of paper or 3d print something weird. I'd suspect that a good model will infer what it can from universal depth cues (E.g. Bokeh and perspective) when present but not perform as well if they're as not there or as well as it does on familiar objects.
- EnigmaFlare 2y ago[dead]
- modeless 2y agoFalse color depth maps are extremely misleading. The way to judge the quality of a depth map is to use it to reproject the image into 3D and rotate it around a bit. Papers almost never do this because it makes their artifacts extremely obvious. I'd bet that if you did that on these examples you'd see that the hair, rather than being attached to the animal, is floating halfway between the animal and the background. Of course, depth mapping is an ill-posed problem. The hair is not completely opaque and the pixels in that region have contributions from both the hair and the background, so the neural net is doing the best it can. To really handle hair correctly you would have to output a list of depths (and colors) per pixel, rather than a single depth, so pixels with contributions from multiple objects could be accurately represented.
- jonas21 2y ago> The way to judge the quality of a depth map is to use it to reproject the image into 3D and rotate it around a bit. They do this. See figure 4 in the paper. Are the results cherry-picked to look good? Probably. But so is everything else.
- threeseed 2y ago> Figure 4 We plug depth maps produced by Depth Pro, Marigold, Depth Anything v2, and Metric3D v2 into a recent publicly available novel view synthesis system. We demonstrate results on images from AM-2k. Depth Pro produces sharper and more accurate depth maps, yielding cleaner synthesized views. Depth Anything v2 and Metric3D v2 suffer from misalignment between the input images and estimated depth maps, resulting in foreground pixels bleeding into the background. Marigold is considerably slower than Depth Pro and produces less accurate boundaries, yielding artifacts in synthesized images.
- incrudible 2y ago> I'd bet that if you did that on these examples you'd see that the hair, rather than being attached to the animal, is floating halfway between the animal and the background. You're correct about that, but for something like matte/depth-treshold that's exactly what you want to get a smooth and controllable transition within the limited amount of resolution you have. For that use case, especially with the fuzzy hair, it's pretty good.
- tedunangst 2y agoFunny that Apple uses bash for a shell script that just runs wget. https://github.com/apple/ml-depth-pro/blob/main/get_pretrained_models.sh https://github.com/apple/ml-depth-pro/blob/main/get_pretrain...
- brcmthrowaway 2y agoDoes this take lens distortion into account?
- dguest 2y agoWhat does this look like on an M. C. Escher drawing, e.g. https://i.pinimg.com/originals/00/f4/8c/00f48c6b443c0ce14b5199d805631b56.jpg https://i.pinimg.com/originals/00/f4/8c/00f48c6b443c0ce14b51... ?
- yunohn 2y agoLooks like a screenshot from the Monument Valley games, full of such Escher like levels.
- coder543 2y agoOn another thread, someone linked to this online demo where you can try it out: https://huggingface.co/spaces/akhaliq/depth-pro https://huggingface.co/spaces/akhaliq/depth-pro
- yread 2y agoHmm it doesn't crash or do anything weird but it just separates the background from the foreground. It looks like the whole foreground is at one distance and the background has a bit of a gradient (higher is further). I've never actually looked at the background that much before and it messes with you as well haha. So it's difficult to tell whether the model is right or wrong.
- qingcharles 2y agoAdobe's did the same: https://imgur.com/a/u87J9A9 https://imgur.com/a/u87J9A9
- qingcharles 2y agoAdobe's version of this tool couldn't figure it out at all (and it works almost flawlessly for complicated regular real world photos). Here's the depth map: https://imgur.com/a/u87J9A9 https://imgur.com/a/u87J9A9 (basically, it pretty much thought the whole thing was entirely flat with some distinction of the far-off background)
- cpgxiii 2y agoThe monodepth space is full of people insisting that their models can produce metric depth with no explanation other than "NN does magic" for why metric depth is possible from generic mono images. If you provide a single arbitrary image, you can't generate depth that is immune from scale error (e.g. produce accurate depth for a both an image of a real car and a scale model of the same car). Plausibly, you can train a model that encodes sufficient information about a specific set of imager+lens combinations such that the lens distortion behavior of images captured by those imagers+lenses provides the necessary information to resolve the scale of objects, but that is a much weaker claim than what monodepth researchers generally make. Two notable cases where something like monodepth does reliably work are actually ones where considerably more information is present: in animal eyes there is considerable information about focus available, let alone the fact that eyes are nothing like a planar imager; and phase-detection autofocus uses an entirely different set of data (phase offsets via special lenses) than is used by monodepth models (and, arguably, is mostly a relative incremental process rather than something that produces absolute depth).
- EnigmaFlare 2y ago[dead]
- reissbaker 2y agoYou wouldn't use monodepth for self-driving, but I think it's useful for producing natural-looking image transformations; after all, you're working with the same image onscreen that a human eye sees (both the computer and the human eye looking at an onscreen photo have access to the same number of bits of information), so if you can get an accurate read on what a human would visually process the image depth to be, you can use that for reasonable-looking changes. I'm a little surprised there isn't more research on producing depth from short videos, though, since iPhones take Live Photos and presumably could use more of the natural movement and shakiness to interpret more of the true depth in the scene... Presumably much better than processing a single still.
- cpgxiii 2y ago> You wouldn't use monodepth for self-driving... If you look where a lot of the money has gone into monodepth, self-driving or driver assistance isn't too far away... Actually, I tend to think self-driving is one of the few places you can make a case for monodepth, as a backup for failures in the rest of your depth sensing suite. You wouldn't want to use it as a primary sensor, but if you're a vehicle driving on a highway and you take critical damage to some of your sensors you do still have to keep driving, if only long enough to get safely off the road, and having something that only needs a single camera is very valuable.
- andrewmcwatters 2y agoNo mention of the effective range. Useful as a feature for SLAM on room scale? Maybe, probably useless at on-road scale.
- crancher 2y agoI’m guessing this is the tool behind the Vision Pro’s Photos.app’s 2D-to-Spatial feature which produces excellent results. It truly improves most photos significantly.
- tommiegannert 2y agoTagging along... If I wanted to do 3D reconstruction of a village from "smooth" street video, what's the best hobby tool today? Gaussian splatting had me amazed, but of course, I'd like a mesh, so there would have to be post-processing to optimize the splats into surfaces, and then reconstruct surfaces.
- jtxt 2y agoConsider meshroom, nerfstudio, (search alternatives too) ffmpeg to split to frames.
- briansm 2y agoJust for reference, the model is ~2GB.
- astrostl 2y agoThe example images all have prominent bokeh (blur from out of focus areas). I wonder if it works on images with narrow apertures or focus stacking where the entire image is in focus.
- isoprophlex 2y agoThe example images look convincing, but the sharp hairs of the llama and the cat are pictured against an out-of-focus background... In real life, you'd use these models for synthetic depth-of-field, adding fake bokeh to a very sharp image that's in focus everywhere. so this seems too easy? Impressive latency tho.
- JBorrow 2y agoI don't think the only utility of a depth model is to provide synthetic blurring of backgrounds. There are many things you'd like to use them for, including feeding into object detection pipelines.
- dagmx 2y agoOn visionOS 2, there’s functionality to convert 2D images to 3D images for stereo viewing. https://youtu.be/pLfCdI0mjkI?si=8K7rPHu558P-Hf-Z https://youtu.be/pLfCdI0mjkI?si=8K7rPHu558P-Hf-Z I assume the first pass is the depth inference here.
- deleted 2y ago[deleted]
- amluto 2y agoI’m not convinced that this type of model is the right solution to fake bokeh, at least not if you use it as a black box. Imagine you have the letter A in the background behind some hair. You should end up with a blurry A and most in-focus hair. Instead you end up with an erratic mess, because a fuzzy depth map doesn’t capture the relevant information. Of course, lots of text-to-image models generate a mess, because their training sets are highly contaminated by the messes produced by “Portrait mode”.
- qingcharles 2y agoI use the fake bokeh "lens blur" tool in Adobe Camera Raw every day, and 99% of the time it gets these kinds of problems correct. Every now and again I have to click the tool and adjust the depth map, but for the most part it is amazingly good. I don't know what ML model they use and how far behind this Apple SOTA research Adobe's product is.