3 ms·
Is it because of what the items are, specifically? Can you add an attribute to every potential item detection that includes a proportionality measure to someth
by drcomputer 12y ago
Is it because of what the items are, specifically?
Can you add an attribute to every potential item detection that includes a proportionality measure to something standardized?
Wouldn't this be a "simple" (that is, definable) calculation if every item that can be detected (meaning the item is already cataloged) has a standardized dimensional measurement in relation to the camera?
- karpathy 12y agoI'm not sure if I fully understand, but consider that "a person with backpack" tells you extremely little about the actual positions of those two items in the 2-D coordinate system of the image: It could be an extreme closeup where only the torso/head of the person is visible and a bit of the backpack, or it could be a tiny person anywhere in the image with a small blob sticking out from their back. The backpack could also be large, small, very occluded, etc. So there is less information in those locations than you might think.
- drcomputer 12y ago> I'm not sure if I fully understand, but consider that "a person with backpack" That's not the only information that you have to work with though. If I understand correctly, object recognition is trained from a data set. That means that every classification is a map from a cluster of images to a single label. This cluster of images can be quantified in terms of proportionality, in relation to other objects identified in the scene. We are beginning with the assumption we can compute the identification of individual objects when they are grouped. The next thing to train on is groups of objects as a unit structure, where we infer and learn a proportion relation. Consider a macro shot of an ant with a baseball in the background and perhaps someone's finger, and a shot of all three with only one dimension altered to the camera. In the first, all three objects appear to be the same size. In the other, object one and object two express a ratio. Then we can normalize vectors based on the 3 computed ratios (object a : object b). Distance is not solved, but relative distance is. Then, if you can build a catalog of standardized measurements (common objects with a defined distance from a camera, like a face 3 feet away from a camera), then you can start training for actual distances. I'm thinking about this like how I would reason about the objects in a raycaster: their definition in the machine versus what gets projected onto a 2d plane, and how moving the objects and camera affects the final projected plane. I'm certain I'm glossing over a lot of details that would actually be difficult in implementation, not to mention computational complexity. > So there is less information in those locations than you might think. I agree, but the only way I understand machine learning at all, as a computer scientist, is "gradual accumulation, definition, and relational organization of humanly defined atomic units". I don't think there are any magic tricks, just a lot of carefully constructed computation, even if it spans over generations of computers, groups, people.