5 ms·
Spoiler: The algorithms actually work by doing text analysis on the captions and other metadata, and no actual image analysis.
by waitwhat 15y ago
Spoiler: The algorithms actually work by doing text analysis on the captions and other metadata, and no actual image analysis.
- datageek 15y agothe algorithm needed to execute quickly. image analysis takes too long.
- kabdib 15y agoWhat's the bottleneck? Bandwidth for downloading images? Understanding of an algorithm that would do the job?
- marshallp 15y agoDifficulty in creating an algorithm. There are ways to get it done algorithmically, however, the challenge is in getting enough data. 30,000 images is too low. you would need a few million, then just simple machine learning algorithms would work. The latest machine learning techniques such as unsupervised deep learning might work, with millions of unlabeled images and the 30,000 labelled.
- petewarden 15y agoBandwidth and general cumbersomeness of dealing with larger amounts of data with starving-startup resources. I actually spent about a decade of my career focused on image processing, and while I love it's power, I knew how much of an engineering challenge it can be at massive scale. I need to do a blog post about this, since I know my choice is a bit surprising and needs explanation.
- lwat 15y agoTechnically I wouldn't call it a 'Photo quality algorithm' in that case
- Drbble 15y agoSpoiler 2: They achieved the low cost by exploiting Kaggle's army of machine learning practitioner suckers/volunteers/competitors.
- socialist_coder 15y agoPretty disappointing. The image analysis was the only reason I clicked.