3 ms·
> I'm not sure if I fully understand, but consider that "a person with backpack" That's not the only information that you have to work with though. If I under
by drcomputer 12y ago
> I'm not sure if I fully understand, but consider that "a person with backpack"
That's not the only information that you have to work with though.
If I understand correctly, object recognition is trained from a data set. That means that every classification is a map from a cluster of images to a single label.
This cluster of images can be quantified in terms of proportionality, in relation to other objects identified in the scene. We are beginning with the assumption we can compute the identification of individual objects when they are grouped. The next thing to train on is groups of objects as a unit structure, where we infer and learn a proportion relation.
Consider a macro shot of an ant with a baseball in the background and perhaps someone's finger, and a shot of all three with only one dimension altered to the camera. In the first, all three objects appear to be the same size. In the other, object one and object two express a ratio. Then we can normalize vectors based on the 3 computed ratios (object a : object b). Distance is not solved, but relative distance is.
Then, if you can build a catalog of standardized measurements (common objects with a defined distance from a camera, like a face 3 feet away from a camera), then you can start training for actual distances.
I'm thinking about this like how I would reason about the objects in a raycaster: their definition in the machine versus what gets projected onto a 2d plane, and how moving the objects and camera affects the final projected plane. I'm certain I'm glossing over a lot of details that would actually be difficult in implementation, not to mention computational complexity.
> So there is less information in those locations than you might think.
I agree, but the only way I understand machine learning at all, as a computer scientist, is "gradual accumulation, definition, and relational organization of humanly defined atomic units". I don't think there are any magic tricks, just a lot of carefully constructed computation, even if it spans over generations of computers, groups, people.