3 ms·
"Off-the-shelf" approach didn't work, which is the reason we did this. And the simplifying assumptions are things we could get away with while solving the task
by yaroslavvb 13y ago
"Off-the-shelf" approach didn't work, which is the reason we did this. And the simplifying assumptions are things we could get away with while solving the task at hand.
Traditional OCR pipeline would be to use some heuristics to find line of text in an image, use some other heuristics to break line of text into candidate characters. Some candidate blobs may need to be merged to make a single character, so you use a separate character classifier pre-trained on correctly segmented characters to score the candidates, and then Viterbi/A* search on those scores to find the most likely interpretation of input.
Many problems with this -- how do you tune the heuristics? How do you recover from error in an earlier stage of pipeline? How do you get character level ground truth from image/text pairs?
With enough engineering time, you can solve those problems, but it's a lot of coding and tweaking. The point of the paper is that you can skip those steps and read OCR output directly off top layers of the network.
- datawander 13y agoAwesome to hear from you! Good points and I believe you're right on all of them. I thought there was more to it because the headline made me think that the general problem of taking a random StreetView location and finding the addresses in it was solved, which really set some high expectations. This work is equally impressive and good stuff. Hope to see this headline again in the future soon.