5 ms·
Like the "youtube cats" paper that only has 16% rate of success (which represented a huge improvement), this doesn't appear to be anything beyond what has alrea
by datawander 13y ago
Like the "youtube cats" paper that only has 16% rate of success (which represented a huge improvement), this doesn't appear to be anything beyond what has already been done for text recognition, which is usually one of the first examples held up of a place Machine Learning has done exceedingly well with 20 years ago.
This line gives it away. If they can remove that assumption I will be impressed, otherwise I would say they reinvented the wheel and probably could have used something off the shelf and got similar results.
"To start off with, Goodfellow and co place some limits on the task at hand to keep it as simple as possible. For example, they assume that the building number has already been spotted and the image cropped so that the number is at least one third the width of the resulting frame. They also assume that the number is no more than 5 digits long, a reasonable assumption in most parts of the world."
- DannyBee 13y ago1. They are doing multi-digit recognition at once, which, as the paper says "To our knowledge, all previously published work cropped individual digits and tried to recognize those". So i don't understand your issue with the 5 digit limit, when everyone else is sticking to "1 digit at a time". 2. Spotting the building number, as the paper says, is taken care of by a different algorithm. I'm not sure why this is also a big deal, since spotting the building number is "not the hard part" in most cases. 3. Your assertion that they could have used something off the shelf seems directly contradicted by the fact that the paper says nobody has ever published a multi-digit simultaneous recognition paper. So i'm very curious what this "off the shelf" thing would be. Could you elaborate?
- darklajid 13y agoDisclaimer: I'm working in the OCR industry, which .. doesn't make me an expert on either state-of-the-art recognition algorithms nor do I know enough about neural networks to be dangerous. That said: The GP has a point, imo. I don't doubt that the paper describes something new and interesting (and all engines I work with during the day do segmentation/recognize character by character), but localizing a region of interest is usually the hard job for me. When I identify the right region and crop it/scale it/rotate it .. my job's "easy" and I can run a multitude of generally good OCR engines (off the shelf, if you will) and get decent results (maybe vote a bit, use engine A to segment and engine B and C to recognize the characters etc.) So .. ignoring the 'we trained a neural network' part (which makes me nod thoughtfully and mumble 'whatever they did there..'), which I _understand_ is the interesting thing here!, they did more or less what I do all the time. The preceding algorithm is what I do quite a bit less often and which in my environment is more interesting and often challenging. Then again, it can always be labeled as PEBKAC I assume :)
- beagle3 13y agoSo ... as someone in the OCR industry, perhaps you could answer: I'm trying to OCR license plates from random videos (stable, but essentially random camera location that sees cars going by and stopping - and I would like to read their plates). I've tried every commercial offering under $4000/license, and -- even when manually localized and a single photo selected, I'm getting less than 95% on a single letter/digit, which translates to ~70% for full plates. (Humans get >99% for full plate on those photos). Where would you recommend I look next for a solution?
- darklajid 13y agoHard to tell. First: I'm a lowly developer here, right? I can't and won't sell you stuff. In addition: While I care a lot about my craft/developing, I don't care about the industry - I'm blind to most competition. Regarding your particular problem: No idea about international license plates (or the one you are interested in). It helps a lot to restrict engines to character sets. German license plates are roughly [1] ([A-Z]{1,3})-([A-Z]{1,2})(\d{1,4}) which helps a lot. Usually you try to combine 'dumb' OCR with datasets/fuzzy matches to rule out errors. Only you know if that is possible for your dataset. Depending on whether I understood your problem correctly we might again have the localization issue (which I complained about above): Find the licence plate, crop and rotate it. Bonus points if your images might contain multiple license plates and a human operator would 'obviously' see the right one.. Recognition itself should be okay: Limited character set, a limited number of fonts (here: One only) and hopefully decent binarization opportunities (here: black on white, background reflective). Feel free to shoot me a mail, details in my profile. 1: From memory, might be slighly inaccurate, sample only
- beagle3 13y agoThanks! It's actually Israeli license plates I'm working on right now (because that's the video data set that I got for training so far, though this project is going to be deployed mostly around Europe) - only digits, a standard xx-xxx-xx template, nice standard retro-reflective black-on-yellow - supposed to be the easiest possible case. And it is, for humans. But all the commercial OCRs I managed to find have abysmal performance. When I get some breathing time, I'm going to try the latest Tesseract again (when I last tried, it was v2, and its performance wasn't good).
- datawander 13y agoI stand corrected and wasn't aware of the multidigit claim.
- mtrimpe 13y agoIt seems like another practical application of the latest improvements on neural networks. You can watch this talk to get an idea for what they're about: http://www.youtube.com/watch?v=vShMxxqtDDs http://www.youtube.com/watch?v=vShMxxqtDDs
- yaroslavvb 13y ago"Off-the-shelf" approach didn't work, which is the reason we did this. And the simplifying assumptions are things we could get away with while solving the task at hand. Traditional OCR pipeline would be to use some heuristics to find line of text in an image, use some other heuristics to break line of text into candidate characters. Some candidate blobs may need to be merged to make a single character, so you use a separate character classifier pre-trained on correctly segmented characters to score the candidates, and then Viterbi/A* search on those scores to find the most likely interpretation of input. Many problems with this -- how do you tune the heuristics? How do you recover from error in an earlier stage of pipeline? How do you get character level ground truth from image/text pairs? With enough engineering time, you can solve those problems, but it's a lot of coding and tweaking. The point of the paper is that you can skip those steps and read OCR output directly off top layers of the network.
- datawander 13y agoAwesome to hear from you! Good points and I believe you're right on all of them. I thought there was more to it because the headline made me think that the general problem of taking a random StreetView location and finding the addresses in it was solved, which really set some high expectations. This work is equally impressive and good stuff. Hope to see this headline again in the future soon.