4 ms·
Something that's maybe not clear from that page, but was mentioned in the Google IO keynote: this is not just stitching of multiple video streams. Google does
by bd 11y ago
Something that's maybe not clear from that page, but was mentioned in the Google IO keynote: this is not just stitching of multiple video streams.
Google does some heavy-duty machine learning / computer vision on their servers to extract 3D information from these video streams (they mentioned having depth data). Then they presumably re-generate seamless stereoscopic 360-degree video from this 3D model. That's how they can get stereo from mono data.
It's something like those hyperlapse videos from Microsoft, they also first do number crunching on lot of video data to generate 3D model and then use this to fill in missing data.
- waynecochran 11y agoYes. I wrote the code for the http://www.3d-4u.com http://www.3d-4u.com 3D camera. You definitely can not use simple still image stitching algorithms. Per pixel depth estimation (via dense stereo matching) and subsequent blending must be used. I would say more, but I had to sign a long NDA before I worked on this.
- joering2 11y agoThis looks like 180, not 360.
- waynecochran 11y agoYep -- same seam issues though.
- sebastianavina 11y agoi was expecting the cube to get solved once the page finished loading.. I'm dissapointer
- josu 11y ago> get stereo from mono data Why do you call it mono data? Given how close are each lense from another, I'm guessing, that anything far enough (0.5m?) from the cameras is beign recorded by at least two cameras at all times.
- jrowley 11y agoBecause simply stitching it will create one image/video stream from one perspective. If you have 2 eyes, you have stereo vision so can perceive depth. This technique allows for the synthesis of stereo video from a rig that doesn't directly capture stereo video. Pretty cool stuff
- ChuckMcM 11y agoThe point the GP was making was that you can treat each pair of cameras as a pair of 'eyes' for the stereoscopic analysis. And you get every combination of adjacent cameras as a pair so a ring of 8 cameras represents 8 distinct pairs (call them camera 'a' and camera 'a' + 1 clockwise) By processing those streams you should be able to derive high quality depth information from the combined video streams, and given the overlap in views re-compute a 'view' from any direction. A lot of early work was done using the streetview cameras as well to do street facing feature extraction using similar techniques but as those cameras were essentially single shot (discrete pictures over time rather than video) you end up with the "slewing" effect of moving from one spot to the next. Presumably Google will go out and do street view with this sort of setup and get a continuous terrain option so that you can really walk around as if you were there (caveat occlusion effects)
- bigiain 11y agoI don't know specs of the HERO4, but the GoPro HERO3 has a 170 degree view angle - so that 16 camera rig has a _significant_ overlap. I suspect you could generate both "adjacent camera" and "every second camera" stereoscopic pairs without too much difficulty, and you can probably even extract useable data out of the "every third camera" pairs and their ~30 degree view overlap. My guess is everything seen by that rig is viewed by 3 distinct pairs (a&a+1, a&a+3, and a&a+3), and that the entire scene from a 16 camera 170fov rig could be considered to have 16 a&a+1 pairs, 16 a&a+2 pairs and 16 a&a+3 pairs giving you three stereo views with different baselines for every point in the scene. I'm guessing there's a _lot_ of spatial information in there.
- 11y ago
- saluk 11y agoI had thought we would never get 360 + stereo. I wonder how much touch up work this will require and what kinds of scenes will provide problems with incorrect depth heuristics.