6 ms·
I personally do not believe that depth generated purely from deep learning can be used as input to photogrammetry anytime soon. Photogrammetry works exceedingl
by Fission 7y ago
I personally do not believe that depth generated purely from deep learning can be used as input to photogrammetry anytime soon.
Photogrammetry works exceedingly well because the depth maps that they generate are quite precise and accurate, and mesh reconstruction usually assumes that these points are quite close to ground truth.
Deep learning approaches usually have medium accuracy but low precision, which causes the flickering and smooth surfaces that you see on the person. Even the background has flickering despite being computed through stereo, likely because the camera motion is primarily forward-backward (vs. more accurate side-to-side motion), the baseline is likely small, and the depth isn't globally optimized.
This type of research is super great for applications requiring lower accuracy, typically visual-only applications (e.g. selective blurring, faking stereo on a frame, etc.). But as an input to photogrammetry — probably not anytime soon, until the problems above get resolved.
- TaylorAlexander 7y agoInteresting. Perhaps my idea of this being inserted in to existing algorithms would not work. However I do ultimately seek a low accuracy “visually approximate” 3D scene that I could use for simulation purposes. I guess I could rephrase my desire as: I’d love to see this kind of approach used to train an end to end deep learning photogrammetry system. I feel like the parallel nature of neural nets as well as their ability to approximate results could result in a much less computationally intensive solution to my photogrammetry desires. (I want to train my four wheel drive robot to follow forest trails using the training method described in the “world models” research paper, which requires a simulation to work.)
- Fission 7y agoSome of my friends recently put out http://gibsonenv.stanford.edu/ http://gibsonenv.stanford.edu/ Full simulation with realistic 3D spaces, enables embodied agents to interact and learn from real-world spaces. Not forest trails, but a real world environment. If you really want to create a 3D model of forest trails, photogrammetry should be sufficient, because forest scenes are richly-textured.
- TaylorAlexander 7y agoYes I did come across Gibsonenv and it looks great for indoor scenes. As far as photogrammetry of forest trails, I found it to be very computationally intensive (taking a GCE 32 core instance 30+ hours using 90+GB of ram to compute a scene, only with errors that made it unusable). It felt very heavy handed and given all the great work I've seen in scene understanding using neural nets, it seems like deep learning would be a promising approach here. Maybe there is commercial photogrammetry software that has better pipelines, but I want to be able to compute my scenes on linux and use hundreds of images. I did my computation with OpenSFM and OpenMVS. Both wonderful projects for being free and open source. I did get a lot of great results. But I am convinced a simpler way is possible with deep learning.
- Fission 7y agoOpenSFM is quite out of date, so it's quite inefficient and rather inaccurate (e.g. exhaustive matching is O(n^2), and there are a lot of smarter ways that are closer to O(n)) Also, one of the main steps of mesh reconstruction is depth map generation. It typically takes anywhere from 30-75% of compute time for dense reconstruction, IF it's parallelized thru GPU. If you're using the CPU only to calculate depth maps, you're probably slowing yourself down by an order of magnitude. If you have a GPU, and use a better SFM-MVS solution, then you can quite easily reconstruct datasets of 1k-10k images within 24 hours.
- jasonjs 7y agoWhat would you recommend as a better SFM-MVS solution?
- deleted 7y ago[deleted]
- espes 7y ago> I personally do not believe that depth generated purely from deep learning can be used as input to photogrammetry anytime soon. 6d.ai uses depthnets in its mobile photogrammetry pipeline. demo: https://twitter.com/mattmiesnieks/status/1106722396889702406 https://twitter.com/mattmiesnieks/status/1106722396889702406
- Fission 7y agoHow do you know that 6D AI uses deep learning to predict depth maps? I'm very familiar with their work (they're doing a great job), but the demo video you linked appears to be a photogrammetric-based approach. You can tell because highly-textured surfaces are readily mapped, but low-texture regions remain unmapped, despite high coverage by the camera. Maybe they use learned features for things like persistent AR, but I'm quite certain that they do not use deep learning to predict depth maps ab initio.
- terminalhealth 7y agoCouldn't depth estimates and camera movement estimates based on NNs still substantially speed up the feature matching? Even if it isn't accurate, it seems like it should be able to reject like 90% of the candidate matches with high reliability.