4 ms·
The distance threshold was determined by hand on a per scene basis, and then we find connected components in the graph of neighboring 3D points, so it's a bit o
by GrantS 12y ago
The distance threshold was determined by hand on a per scene basis, and then we find connected components in the graph of neighboring 3D points, so it's a bit of a hack, but there are two important things to note. First, we can get away with an overly permissive threshold because we also require that any two grouped points are also observed at the same time in at least N images -- meaning SIFT features were detected for both 3D points in the same image (so they therefore exist at the same point in time). This filters out lots of spurious groupings. Second, this approach was inspired by super-pixels, which is an OVER-segmentation of an image into groups of pixels that are somewhat coherent -- each super-pixel probably lives on the same semantic object in the world, but they are by no means complete. Still, it's massively better than reasoning about individual pixels (or individual points).
So we err on the side of dividing single buildings up into multiple semantic objects. If our data included detailed reconstructions of the streets between each building, then the whole thing might be connected and we'd need more criteria to separate them out -- we do automatically estimate a ground plane so that's one way: just ignore everything near the ground for grouping purposes.
There's slightly more detail in the paper:
http://www.cc.gatech.edu/~phlosoft/files/schindler10cvpr.pdf http://www.cc.gatech.edu/~phlosoft/files/schindler10cvpr.pdf