4 ms·
This is really awesome. For anyone who wants a naive, poor man's implementation of this, I have an unfinished Ruby gem that's a good starting point for you: htt
by joshdotsmith 13y ago
This is really awesome. For anyone who wants a naive, poor man's implementation of this, I have an unfinished Ruby gem that's a good starting point for you: https://github.com/JoshSmith/kaleidoscope https://github.com/JoshSmith/kaleidoscope
Here's how I got it to work (cribbed from my own README):
TL;DR: I used k-means clustering to segment a database of images into color bins for quick searching.
Using imagemagick, I ran histograms on images and converted their top n most frequent colors into Lab* color space for an approximate representation of human vision.
Colors were then matched to a user-defined set of colors using Euclidean distance, i.e. a "bin". I could choose any array of RGB values of arbitrary length.
I then stored hexadecimal values of the image's original color and the matched color, along with the frequency of that color within the image (for sorting based on frequency) and the Euclidean distance (for sorting by tolerance).
Then finding images close to a certain color was as simple as Photo.all.with_color('#993399') and order by frequency and Euclidean distance. Here's a photo of the results: https://github-camo.global.ssl.fastly.net/89cc87ac84cd3a1d1223e8f9d560e65eb8447ef6/687474703a2f2f636c2e6c792f696d6167652f336e3243313631373069306b2f53637265656e25323053686f74253230323031332d30322d30352532306174253230362e35362e3434253230504d2e706e67 https://github-camo.global.ssl.fastly.net/89cc87ac84cd3a1d12...
I might spend some time reverse-engineering Shutterstock's implementation, since it sounds way better than mine and clearly works at scale. But for my purposes, my own implementation worked just fine.
If you want help implementing it, feel free to reach out to me!
- clbecker 13y agoThat looks pretty cool. As far as scaling issues went, the bulk of them were addressed just by using Solr on pretty beefy hardware (for our full library of 30+ million images, we're running on Dell r510s w/ 24 cores and 128GB Ram). Of course, depending on your hardware limitations there's compromises you can make to increase speed / reduce memory usage - i.e. the fewer fields you need to index and sort on the better & the lower the cardinality of each field, the better too. Also, since our input only consisted of one slider, we were able to run all the distance calculations beforehand, and just store a score that represented the image's distance from the given slider position - something like that might also work for you, since you're using a limited palette of input colors too.
- kevinh 13y agoclbecker, your account has been shadowbanned so your replies won't show up on any posts. This is very unfortunate because it seems like you're the origin of this post. So, make a new account or something. Everyone else: If you want to see his comments, turn on the showdead option in your profile.
- deleted 13y ago[deleted]