5 ms·
Elasticsearch Based Image Search Using RGB Signatures
- rcarmo 10y agoI wonder if pHash wouldn't make this a lot more effective. Anyone tried building an ES-usable distance function for pHash?
- sujitpal 10y agoHello, author of the post here. Thanks for all your valuable comments. I noticed that there was some talk about neural network driven models. In my follow up post to this one: http://sujitpal.blogspot.com/2016/06/comparison-of-image-search-performance.html http://sujitpal.blogspot.com/2016/06/comparison-of-image-sea... I describe how I use transfer learning to generate image vectors for my butterfly images from the Caffe reference model trained on ImageNet. There are some other approaches too, but probably not as interesting to this audience. @GrantS - thanks for the intro and the links, and for pointing out that I was using BoVW incorrectly. Not an image person, trying to get into it from a search/NLP background. @infinitone and @deckar01 - I measured the goodness of search by checking how many times the query butterfly was the #1 result. I agree its pretty hard to distinguish butterflies from one another, but I needed something to work with, and that was as good as any. I did not want one that crosses subject areas, such as flowers and cars, so a search for a red flower might bring back a red car. @chrischen - if you got rid of the bucketing portion in my pipeline, I think you might achieve what you are after. But if you are only looking for exact match on color, then hashing might be cheaper. @rcarmo - thanks for the reference to pHash, I think it might work better in near-duplicate detection kind of cases. Will check it out.
- GrantS 10y agoFor anyone interested in the computer vision side of this topic, the author here is using a variant of color histograms, which was state of the art around 1990 [1][2]. Since 2003, bag of visual words approaches have usually meant extracting SIFT-like features from a database of images, quantizing the features down to a list of thousands or millions of "words", and then treating the images like documents containing those "visual words" [3][4]. (Nothing wrong with the approach he's using [simple and fast], but the bag of words terminology in the article usually suggests a different class of approaches.) [1] https://staff.fnwi.uva.nl/r.vandenboomgaard/IPCV/_downloads/swainballard.pdf https://staff.fnwi.uva.nl/r.vandenboomgaard/IPCV/_downloads/... [2] https://www.cs.utexas.edu/users/dana/Swain1.pdf https://www.cs.utexas.edu/users/dana/Swain1.pdf [3] http://www.robots.ox.ac.uk/~vgg/publications/papers/sivic03.pdf http://www.robots.ox.ac.uk/~vgg/publications/papers/sivic03.... [4] http://www-inst.eecs.berkeley.edu/~cs294-6/fa06/papers/nister_stewenius_cvpr2006.pdf http://www-inst.eecs.berkeley.edu/~cs294-6/fa06/papers/niste...
- infinitone 10y agoForgive my ignorance, but is there something like word2vec but for images- like an image2vec? In terms of text processing, word2vec is one of the best approaches. So can't you describe the features of images using neural networks and then vectorize them using word2vec? Its been acouple years since my computer vision course (was my favorite course in university) but isn't SIFT a bit '99? Aren't there better methods now such as neural networks for feature description?
- markbnj 10y agoI'm also pretty ignorant here, but if I understand word vectors correctly they are a trained model that results in the ability to predict the next word in a sequence. I can imagine extending that idea to images, i.e. in such a way that the color of a target pixel can predict the colors of surrounding pixels. I wonder if tools like Photoshop and Gimp don't already have similar algorithms for some advanced effects. In any case, it seems several layers of abstraction down the stack from the problem of attaching meaning to images. Isn't that more what the BOW approach tries to do?
- GrantS 10y agoImages don't necessarily map directly onto a word2vec-like solution but the closest thing is pre-trained deep networks, for example: http://caffe.berkeleyvision.org/model_zoo.html http://caffe.berkeleyvision.org/model_zoo.html You're right that bag of words with SIFT is not state of the art, with deep learning dominating computer vision approaches these days.
- jkldotio 10y agoI think in addition to histograms and features like SIFT, SURF, and others like DAISY[0] that would permit searching with images as a query it would be beneficial just to use a neural network to text index using the classes at very top of a neural network, although you are right some features in layers just below that could be used too as I understand it. You could then use classical text indexing on the text, perhaps with a topic model like LDA. Then an image with a plane in it will be indexed by "plane" via the output of the neural network but would also come first, or in the top results, when using "flight" as the query via a topic model. Ditto for word2vec or para2vec over those words, the benefit being you can bring the knowledge of relations contained in the textual training data, Wikipedia or something else, to bear on the problem. I.e. a golf club and a baseball glove might not be correlated in the neural network that annotated the images but might be correlated in the text based knowledge model trained on Wikipedia and so a query of "sport" might bring both images up. The big players like Google already have caption generation that's capturing relationships between objects.[1] [0]http://scikit-image.org/docs/dev/auto_examples/features_detection/plot_daisy.html#sphx-glr-auto-examples-features-detection-plot-daisy-py http://scikit-image.org/docs/dev/auto_examples/features_dete... [1]https://arxiv.org/pdf/1411.4555.pdf https://arxiv.org/pdf/1411.4555.pdf
- infinitone 10y agoReally cool. But maybe its just me, aren't butterflies kind of hard to distinguish between each other. I feel as though his search result page- i couldn't really tell if it was good or not because they all looked kind of similar shape-wise. Only difference is color and even then its fairly little color.
- deckar01 10y agoThe strategy does seem to assume a highly normalized set of images that only differ in color and contour.
- willcodeforfoo 10y agoAnother open-source implementation using a similar concept and Elasticsearch is image-match: https://github.com/ascribe/image-match https://github.com/ascribe/image-match
- sandeepc 10y agoIf you're interested in searching photos with ES - I took a some what simpler approach focusing just on major colors in the image. But with some of the machine vision API google cloud etc. you could extend to other "features" Details: http://blog.sandeepchivukula.com/posts/2016/03/06/photo-search/ http://blog.sandeepchivukula.com/posts/2016/03/06/photo-sear...
- UncleChis 10y agoI have put quite a bit of similar effort to image retrieval using Elasticsearch before. While it is nice and convenient, what I found was that Elastic Search is too slow for larger scale (million of Images with dictionary of millions visual words in BoW model). May be there're some steps that I did not do right, but I gave up.
- softwaredoug 10y agoMy "Ghost in The Search Machine" talk builds a really naive image search demo (which Sujit uses for his starting point). You might enjoy that: https://www.elastic.co/elasticon/conf/2016/sf/opensource-connections-the-ghost-in-the-search-machine https://www.elastic.co/elasticon/conf/2016/sf/opensource-con...
- jerluc 10y agoIn my [limited] experience with CBIR and image based search, I found that using a color space with perceptual spatial qualities (such as one of the CIE LaB variants) to be more effective than a purely normalized geometric color space (such as RGB or HSV), as color similarities in the latter may not make much sense to a human.
- chrischen 10y agoDoes anyone have a solution to this where the input (search) is also a vector as opposed to a single color, but would still allow for exact color matches? Bucketing means that you can't get the granularity of a specific shade of a color.
- deleted 10y ago[deleted]