5 ms·
3-D Depth Reconstruction from a Single Still Image (2007) [pdf]
- jmcmahon443 11y agotl:dr; Can combine monocular NN CV techniques with stereo techniques for good, cheap results. AKA: The future of SLAM.
- dheera 11y agoI do not think this kind of feature-based estimation of depth will be the future of SLAM. Point this algorithm at a scale model or a framed photograph and everything will go haywire. It's great for scene understanding, in which you point it at a photograph or artwork and you want it to understand what's going on in the photograph. But not for mission-critical mapping and navigation. Depth cameras, which are getting cheaper, will be the future of SLAM. SLAM algorithms for depth cameras are also a lot simpler to write. And with depth cameras, you're not estimating how far away objects are, you're actually measuring their distance. Data beats estimation.
- jmcmahon443 11y agoDepth/stereo cameras are the future, I agree. But the conclusion of this paper says that you can cascade the results of this algorithm with the stereo algorithm fairly easily.
- Xcelerate 11y agoAs an aside, I really enjoy it when people post articles like this on HN. Technical, but not too niche, and relevant to today's interests in technology.
- GrantS 11y agoThis is some very cool research, though many may be surprised to learn that it is an IJCV journal article from 2007, based on a conference paper from NIPS 2005. Source: http://www.cs.cornell.edu/%7Easaxena/learningdepth/ http://www.cs.cornell.edu/%7Easaxena/learningdepth/
- TheArcane 11y agoGah! I thought this was cutting edge. Is this research used in any recent works/research?
- Xcelerate 11y agoDefinitely. It's been cited 658 times since it was published: https://scholar.google.com/scholar?oi=bibs&hl=en&cites=18064180945750989807,18259224307089260510 https://scholar.google.com/scholar?oi=bibs&hl=en&cites=18064...
- jvhaarst 11y agoVideo is a nice watch, especially as this is from 2007, so the state of the art should be even better. https://www.youtube.com/watch?v=UZ7_ED9g4FY https://www.youtube.com/watch?v=UZ7_ED9g4FY
- pen2l 11y agoGoogle Tango (https://www.google.com/atap/project-tango/ https://www.google.com/atap/project-tango/) is able to do pretty impressive 3d reconstruction from what I hear. However it has a number of cameras, whilst linked technique can work with just one image (with enough training).
- leeoniya 11y agoreconstructing depth from 2+ offset images is vastly simpler than from a single frame.
- pen2l 11y agoYup, but I think there's a desire to get good 3d-reconstruction techniques for single images because in practice you often want to do reconstruction in extremely small environments where using multiple cameras is sometimes not feasible for various reasons.
- leeoniya 11y agoGoogle Camera can already do this from a single camera. It takes a bunch of photos as you move the phone around a bit and reconstructs the depth from multiple shots. http://googleresearch.blogspot.com/2014/04/lens-blur-in-new-google-camera-app.html http://googleresearch.blogspot.com/2014/04/lens-blur-in-new-...
- dr_zoidberg 11y agoThere's an OpenCV tutorial for those wishing to try it[0]. [0] http://docs.opencv.org/3.0-beta/doc/py_tutorials/py_calib3d/py_depthmap/py_depthmap.html#py-depthmap http://docs.opencv.org/3.0-beta/doc/py_tutorials/py_calib3d/...
- pen2l 11y agoI wish these groups would make their code available so others could play with it and test it out. Some do, but I wish more did. I would be curious to see the performance of this 3d depth reconstruction technique for non-rigid environments.
- jmcmahon443 11y agoI love open source, but there is a very good reason why academics generally don't release their code. They are scientists: they are best at testing, explaining, and documenting nature. Code is an engineer's job: making it as fast and robust as possible for the given application or with the latest tools. They want to make you work for it :)
- pen2l 11y agoWith all due respect, that is a stupid reason. I don't have time to implement every last thing. And I don't require robust code all the time -- it's very understandable for the code to not be robust if it's being provided by a small group with not a lot of funding or resources. Hell, if they make it open source I will contribute to it expecting nothing in return to make it better in any way I can. I really think that the reason the code isn't shared is journals don't encourage this type of thing (or make it easy). Lab websites live and die, so hosting it there is not a good idea. Very simply, the code should should be packed with the paper somehow. I hope some journal takes the lead soon in creating a system that encourages a bundled presentation of paper + code.
- jmcmahon443 11y agoAcademics aren't always concerned with making your job easy, they have lives too. They may have struggled a lot for their results and their code is still a living document, or they sell it to a company. Maybe their grant ran out and they got pulled into something else. Maybe they died. There's a lot of factors. Look at OpenRave. It's the result of an academic open sourcing his research. I just think at the end of the day, a research paper exists to validate human understanding of nature - not give you a boilerplate for implementing it.
- AndrewKemendo 11y agoAlways love seeing computer vision related stuff posted. Incidentally, we have improved on these techniques (given that this came out almost a decade ago!) for scale reconstruction with our mobile monocular SLAM system. The authors are now killing it worldwide (Obviously for Ng) working on applications.