5 ms·
How Google Cracked House Number Identification in Street View
- sjtgraham 13y agoAnyone else getting a lot of obvious door numbers in reCAPTCHA captchas lately?, I guess that's how Google trained these neural nets.
- RDeckard 13y agoYes.
- amjaeger 13y agowas just thinking that google can now fool captchas
- clhodapp 13y agoGoogle is also the one serving them, which does allow them to cheat.
- somesay 13y agoreCaptcha was always a free captcha service based on one simple idea: one part is the classical one, hopefully only readable by humans, the other one actually supports Google on doing OCR-like jobs. Previously it were undetected words from Google Books scans, now they are mostly using house numbers. More interesting is the other change: reCaptcha now tries to detects real users and then only generates a simple number captcha for the classical part. Likely they are using your Google Account cookie or Google Analytics for that.
- deleted 13y ago[deleted]
- ezrameanshelp 13y agoOne CAPTCHA to rule them all!
- mholt 13y agoI've been thinking that for a few months now. When reCAPTCHA was revised last year and turned into numbers, I immediately figured it was for Google to read the house numbers on street view images. Brilliant, really.
- VanillaCafe 13y agoBut in particular to create a seed training set (for the neural network) -- not as a standalone house number recognition system.
- malandrew 13y agoI hope that the training set gets released publicly. The more training sets openly available to all the better.
- mtrimpe 13y agoSorry for the downvote; I though you were off-topic before my brain processed the second sentence. Edit: thanks for the downvotes for apologizing for an errant downvote ... I guess ...
- blueskin_ 13y agoI noticed those when they first started appearing, my first thought was "must be Street View", followed immediately by wondering about the potential for poisoning some kind of machine learning effort with them (and yeah, I know, issue of scale, even a group of people dedicated to ruining the effort would be vastly outnumbered by background reCAPTCHA traffic).
- danielweber 13y ago4chan has had campaigns to get "penis" and a certain racial slur inserted into books. I was going to post links but then decided I didn't want to get fired today for my web searches.
- chippy 13y agoGooogle are being totally dishonest and misleading in their use of reCAPTCHA to digitize house numbers. "reCAPTCHA is a free CAPTCHA service that helps to digitize books, newspapers and old time radio shows." "Currently, we are helping to digitize old editions of the New York Times and books from Google Books." They are lying and they are not benefiting the public interest - which is what reCAPTCHA relies on - they are only benefiting their company and their own shareholders.
- userbinator 13y agoI just put in rubbish for that part. OCR'ing books isn't bad, but I really don't find the idea of helping Google with house numbers appealing at all.
- ezrameanshelp 13y agoWell I like being able to accurately view an address in street view, so you are helping me, too. Thanks?
- oxplot 13y agoYou say that as if you're not helping Google with anything else. Every search, every plus, every rating, every interaction helps Google with its business. And in majority of cases, users reap the benefits too and house numbers is one of those—better geographical knowledge is always good. Whether it can be abused is irrelevant to this specific matter. Every piece of information about anything can be abused.
- datawander 13y agoLike the "youtube cats" paper that only has 16% rate of success (which represented a huge improvement), this doesn't appear to be anything beyond what has already been done for text recognition, which is usually one of the first examples held up of a place Machine Learning has done exceedingly well with 20 years ago. This line gives it away. If they can remove that assumption I will be impressed, otherwise I would say they reinvented the wheel and probably could have used something off the shelf and got similar results. "To start off with, Goodfellow and co place some limits on the task at hand to keep it as simple as possible. For example, they assume that the building number has already been spotted and the image cropped so that the number is at least one third the width of the resulting frame. They also assume that the number is no more than 5 digits long, a reasonable assumption in most parts of the world."
- DannyBee 13y ago1. They are doing multi-digit recognition at once, which, as the paper says "To our knowledge, all previously published work cropped individual digits and tried to recognize those". So i don't understand your issue with the 5 digit limit, when everyone else is sticking to "1 digit at a time". 2. Spotting the building number, as the paper says, is taken care of by a different algorithm. I'm not sure why this is also a big deal, since spotting the building number is "not the hard part" in most cases. 3. Your assertion that they could have used something off the shelf seems directly contradicted by the fact that the paper says nobody has ever published a multi-digit simultaneous recognition paper. So i'm very curious what this "off the shelf" thing would be. Could you elaborate?
- darklajid 13y agoDisclaimer: I'm working in the OCR industry, which .. doesn't make me an expert on either state-of-the-art recognition algorithms nor do I know enough about neural networks to be dangerous. That said: The GP has a point, imo. I don't doubt that the paper describes something new and interesting (and all engines I work with during the day do segmentation/recognize character by character), but localizing a region of interest is usually the hard job for me. When I identify the right region and crop it/scale it/rotate it .. my job's "easy" and I can run a multitude of generally good OCR engines (off the shelf, if you will) and get decent results (maybe vote a bit, use engine A to segment and engine B and C to recognize the characters etc.) So .. ignoring the 'we trained a neural network' part (which makes me nod thoughtfully and mumble 'whatever they did there..'), which I _understand_ is the interesting thing here!, they did more or less what I do all the time. The preceding algorithm is what I do quite a bit less often and which in my environment is more interesting and often challenging. Then again, it can always be labeled as PEBKAC I assume :)
- magicalist 13y agoI think the actual paper (linked at the end) is much more informative and interesting than this coverage: http://arxiv.org/abs/1312.6082 http://arxiv.org/abs/1312.6082 (I wouldn't call it blogspam, because it looks like they interviewed the researchers, but the summary leaves something to be desired)
- aaronsnoswell 13y agoThe MIT guys generally do a pretty good job with their articles. Definitely not blog spam.
- danielweber 13y agoThis is unusual: an article about neural nets that actually uses neural nets.
- marc0 13y agoI must say I find this fascinating, esp two aspects: First, great idea to train on number sequences instead of single characters. That's certainly a lesson which is applicable in many other situation, too. Second, 11 levels and deep learning -- a bold approach, since (from what I know) the paradigm is that deep networks are generally not working well. They even found that performance improves with the levels (they just stopped at 11, probably because of resource limitations). As mentioned in the paper, this is probably due to the fact that the network ist trained with a huge dataset, so the size of the dataset really is the relevant factor.
- deleted 13y ago[deleted]