4 ms·
Thanks for your reply. We are building the picture identification in-house, we have a crude working demo as a proof of concept that the tech is good enough. We
by cjr 14y ago
Thanks for your reply. We are building the picture identification in-house, we have a crude working demo as a proof of concept that the tech is good enough.
We've used the Amazon product API to build our databsae, there are ~80m products in there. We're starting one vertical at a time (the algo works best for highly textured products such as watches/book/cd covers so we've started there). We build an in-memory index of image-descriptors which is queried when trying to match a product. As we start seeing other product type, we'll expand beyond amazon.
As for scaling, the index is the one component that isn't trivial to parallelise. Once we get to a stage where a single box can't store it in memory, the first step would be to shard it, with different indexes per product category. Query images will be run through a product-categorisation algo to determine which indexes they need to be run against. This is how we think we'll approach this, from previous experience, but will investigate properly (if) we get large enough.
In terms of affecting page load, our plugin will load async, as will the results returned from the image matching algorithm, so it should be fairly negligible impact.