3 ms·
Disclaimer: I was the original author of Tesseract.JS— though all the hard work nowadays is done by Jerome Wu. If you're interested in supporting the project, c
by antimatter15 7y ago
Disclaimer: I was the original author of Tesseract.JS— though all the hard work nowadays is done by Jerome Wu. If you're interested in supporting the project, consider backing the OpenCollective (https://opencollective.com/tesseractjs https://opencollective.com/tesseractjs)
By all of that, what I mean to say is that I've learned a decent amount of fun OCR trivia over the past few years.
Firstly, the engine that powers Google Cloud Vision is almost certainly an entirely independent code base from Tesseract built on neural networks. In fact, the most recent major version of Tesseract (version 4.0) was a sort of rewrite of the core of Tesseract to use bidirectional LSTMs to seem a bit more like the modern OCR pipeline that systems like GCV use.
The original Tesseract algorithm dates back a previous AI spring— in the 80s when neural networks were cool (before they were uncool, and then subsequently cool again). The core of the original algorithm involved fitting polygons to character shapes in order generate features which could be matched by a kind of rudimentary neural network.
One of the primary authors of Tesseract is Ray Smith (at Google)— who gave a presentation at some point a few years ago about the history of OCR— though I can't quite find a link to it at the moment.
OCR actually predates electronic computers. In 1929, someone had invented a machine that would take a piece of paper and shine a bright light on a single letter, and pass the letter through a carousel of letter masks, so that it could hit a (effectively single pixel) photo-sensor. When the carousel and the letter mask were in alignment with the printed letter, then the drop in brightness registered that a particular letter was seen!
OCR was used by the US Postal system for sorting mail as early as 1965, but it wasn't until 1976 that any system could reasonably support more than a certain number of hard-coded fonts (fun fact this was invented by Ray Kurzweil, the Singularity is Near guy).
- alexcnwy 7y agoFirst of all thank you for all your hard work. Major question: Why isn’t Tesseract using neural networks. I know it just introduced LSTM based models but they suck. Why is GCP vision text recognition API so much better than open source alternatives?!
- xanderjanz 7y agoThere are open source versions of everything done within a GCP API call, but it requires multiple machines and lots of data to build an NLP model to be as fast and accurate as GCP, and cloud computing is relatively new compared to OCR.
- bhl 7y agoTraining the model would be computationally intensive, but deploying that to use Tensorflow.js and predicting a single datapoint in the browser shouldn't be as much, right?
- est31 7y agoThere are ML models that are so computationally intensive that they can't reasonably run on the edge. AI accelerator chips obviously help move the line, but AI accelerators benefit the cloud, too. Furthermore, Models can be tens to hundreds of megabytes in size. Okay for the cloud, not okay for wasm running in the browser.
- beagle3 7y agoThere are? Can you give a list of pointers or what to look for? I was looking for an OCR that can do license plates while the car is moving, for a hobby project. The image quality is less than perfect, the lighting is never very good, and as the camera is mounted on my side window, all plates have a perspective transformation applied (e.g., topline and baseline are essentially never parallel) Tesseract fails miserably. Trying to help it, I have not found a good open source project that would consistently equalize color pictures to black-and-white - sometimes there's shadow on the plates that foils all simple attempts. And yet, GCV needs no parameters, and seem to do this perfectly on images I've tried. So, assuming I'm willing to put in the time - how do I build my own GCV -- even if it's just for the hobby use case of reading license plate (and the next stage: reading house numbers - which GCV does reasonably well, although it is a much much harder problem)
- mentat 7y ago
- xanderjanz 7y agoAlso AFAIK GCV uses techniques beyond better OCR that greatly help accuracy. It does image fixing, boundary detection, NLP, spell check, etc.