11 ms·
Show HN: Tesseract.js – Pure JavaScript OCR for 60 Languages
- yankyou 10y ago> Drop an English image on this page to OCR it! This looks great, and I'd really love to but > Uncaught ReferenceError: progress is not defined EDIT: works now!
- ckluis 10y agoWhat License? Doesn't mention it.
- atishay811 10y agoSays MIT in package.json
- talklittle 10y agoApache-2.0 They've added https://github.com/naptha/tesseract.js/blob/c26cae7ee956c399eeb992de0135e3af29b4edb5/LICENSE.md https://github.com/naptha/tesseract.js/blob/c26cae7ee956c399...
- slajax 10y agoPretty cool. I screen captured the text in the bottom right corner of the page and it had some issues. Here's a screenshot of what happened: http://io.kc.io/hkeM http://io.kc.io/hkeM
- mdani 10y agoLanguages list link is broken - getting 404 for the following https://github.com/naptha/tesseract.js/blob/master/tesseract_lang_list.md https://github.com/naptha/tesseract.js/blob/master/tesseract...
- bijection 10y agoThanks! Fixed. The actual link is https://github.com/naptha/tesseract.js/blob/master/docs/tesseract_lang_list.md https://github.com/naptha/tesseract.js/blob/master/docs/tess...
- employee8000 10y agoIs this at all affiliated with the already-existing tesseract OCR library? It doesn't seem to be from my cursory check so if not you need to rename your library, because you're ripping off their name. https://github.com/tesseract-ocr/tesseract https://github.com/tesseract-ocr/tesseract
- wibr 10y ago"Tesseract.js is a pure Javascript port of the popular Tesseract OCR engine." first sentence on http://tesseract.projectnaptha.com/ http://tesseract.projectnaptha.com/ linked from the github page
- bijection 10y agoIt's a wrapper around an Emscripten port of that library. See https://github.com/naptha/tesseract.js-core https://github.com/naptha/tesseract.js-core
- KiwiCoder 10y agoImpressive that this is pure JS, however trying an image cut from the page itself gave this result > Dropan Enghsh Wage on (Ms page to OCR m Should be > Drop an English image on this page to OCR it!
- bijection 10y agoAs another commenter mentioned, Tesseract.js won't perform very well on 'natural' images (e.g. the very light text you tried). It should work better if you feed it a screenshot of the black text at the top of the demo page though (Tesseract.js is a pure Javascript port etc...).
- lloeki 10y ago> Impressive that this is pure JS Well it's pure JS in that it's been running the C tesseract through emscripten. So in a way it's pure JS just as much as the original lib is pure assembly when compiled ;-)
- gentleteblor 10y agoI've always wanted to use Tesseract on .NET projects but it was always clumsy (wrappers). This looks like it'll make things easier. Thanks for putting this out.
- smarx007 10y agoI think a .NET wrapper would be more direct/elegant than using emscripten generated code (especially in a .NET project).
- gentleteblor 10y agoA few .NET wrappers exist...but it's always felt heavyweight for me (my use case is pretty rare). I am hoping this makes it trivially easy. We'll see how it goes.
- newtons_bodkin 10y agoHow long did this take to build?
- pyronite 10y agoThe text detection is lacking in comparison to Google's Vision API. Here is a real-life comparison between Tesseract and Google's Vision API, based on a PDF a user of our website uploaded. Original text [http://i.imgur.com/CZGhKhn.png http://i.imgur.com/CZGhKhn.png]: > I am also a top professional on Thumbtack which is a site for people looking for professional services like on gig salad. Please see my reviews from my clients there as well Google detects [http://i.imgur.com/pSJym1x.png http://i.imgur.com/pSJym1x.png]: > “ I am also a top professional on Thumbtack which is a site for people looking for professional services like on gig salad. Please see my reviews from my clients there as well ” Tesseract detects [http://i.imgur.com/wwbLU6g.png http://i.imgur.com/wwbLU6g.png]: > \ am also a mp pmfesslonzl on Thummack wmcn Is a sue 1m peop‘e \ookmg (or professmna‘ semces We on glg salad P‘ezse see my rewews 1mm my cuems were as weH
- bijection 10y agoAlthough Google's API is certainly better, Tesseract.js should work similarly if you increase the font size. Screenshots taken on 'retina' devices are around the smallest text it can handle well. Edit: A screenshot of the same text at a higher resolution: https://imgur.com/a/W7IGu https://imgur.com/a/W7IGu Tesseract.js output: https://imgur.com/a/niIfM https://imgur.com/a/niIfM "I am also a top professional on thumbtack which is a site for people looking for professional services like on gig salad. Please see my reviews from my clients there as well"
- jaytaylor 10y agoYour comment (zoomed in Chrome on Win 10): http://i.imgur.com/uuFhw90.png http://i.imgur.com/uuFhw90.png Tesseract.js analysis: Although Googie's API is certaihiy better, Tesseract.js should work simiiarly if you increase the font size. Screenshots taken on 'retiha’ devices are around the smailest text it can handie well. Edit: A screenshot of the same text at a higher resolution: httgs:[[imgurxomZaN/UGu Tesseract.js output: httgs://imguricom[a[hiIfM This is a neat toy, but not impressive compared to the results from tesseract-ocr/tesseract [0]: $ curl -s http://i.imgur.com/uuFhw90.png \ | tesseract stdin stdout Although Google's API is certainly better, Tesseract.js should work similarly if you increase the font size. Screenshots taken on 'retina' devices are around the smallest text it can handle well. Edit: A screenshot of the same text at a higher resolution: https:[ZimguncomlalWHGu Tesseract.js output: https:[[imgur.com[a[nilfM Notice how Tesseract.js results suffer from being unable to differentiate between n's and h's, i's and l's. [0] https://github.com/tesseract-ocr/tesseract https://github.com/tesseract-ocr/tesseract
- greenpizza13 10y agoExcited about this... but the OCR quality seems to be very bad. Maybe it's not optimized for recognizing black text on a white background. For example, I took a screenshot of this comment and ran it through the demo and got this: Excited ehent this... but the OCR enenty Seems te be very bad. Maybe it's het Dptimized far recngnizing black text an e white heckgmnhe. EDI example, 1 tank e Screenshnt at this cement ehe teh it. thmneh the den» ehd get this: It seems to recognize the bounding boxes just fine but mangles the words.
- bijection 10y agoDid you try increasing the font size a bit? On a retina macbook (so effectively ~2x bigger font) I get: Excited about this... but the OCR quality seems to be very bad. Maybe it's not optimized for recognizing black text on a white background. For example, I took a screenshot of this comment and ran it through the demo and got this:
- mrcactu5 10y agoTesseract is not specific to JavaScript right? I do recall there being a version for Python
- someonewithpc 10y agoNo, tesseract is a C++ library; this is a wrapper for an Emscripten port of that library.
- mrcactu5 10y agoTesseract is not specific to JavaScript right? I do recall there being a version for Python
- jameslk 10y agoFor all those claiming issues with reading text from a screen shot of this page, note that this is more an issue with the original Tesseract library, not this library (which appears to wrap Tesseract compiled through Emscripten). I remember having a similar issue when I used the original Tesseract. The quick hack I found to fix it was to rescale any small text input images 3x first before feeding it to Tesseract. I'm sure there's more intelligent solutions to mitigate that problem.
- dunham 10y agoYeah, in the past, I've had to scale my own scans/photos to get good results. The tesseract github site mentions this: "Tesseract works best on images which have a DPI of at least 300 dpi, so it may be beneficial to resize images." - https://github.com/tesseract-ocr/tesseract/wiki/ImproveQuality https://github.com/tesseract-ocr/tesseract/wiki/ImproveQuali...
- zhte415 10y agoDoes this include taking a text and for example, when viewing it, 'wiping' the text in the logical native language order? For languages that don't employ much whitespace, this would be nice.
- AgentME 10y agoWhy the promise-like interface? If it returned a promise with a this-returning progress method monkey-patched onto it, then you could use it otherwise like a regular promise: Tesseract.recognize(myImage) .progress(function(message){console.log(message)}) .then(function(result){console.log(result)}) .catch(function(err){console.error(err)}); or Tesseract.recognize(myImage) .progress(function(message){console.log(message)}) .then( function(result){console.log(result)}, function(err){console.error(err)} ); I guess I just still have bad memories of jQuery's old almost-like-real promises. I'd rather never have to think ever again about whether I'm dealing with a real promise or one that's going to surprise me and break at run-time because I tried to use it like a real one.
- bijection 10y agoIf you want to use a real Promise, you can wrap the call to recognize in Promise.resolve: Promise.resolve(Tesseract.recognize(myImage)).then(result => console.log(result))
- goatslacker 10y agoI've been using this library to read screenshots of Pokemon Go to automatically calculate Individual Values for each Pokemon[1] It's worked great on desktop, but on mobile safari where it matters most the library causes the browser to crash :( 1: https://github.com/goatslacker/pokemon-go-iv-calculator/blob/master/web/components/PictureUpload.js https://github.com/goatslacker/pokemon-go-iv-calculator/blob...
- methyl 10y agoConsider doing it server-side
- deleted 10y ago[deleted]
- xigency 10y agoTo anyone screen capturing small fonts as a demonstration, or capturing digital text especially at a small resolution, I don't believe that that is the purpose of this OCR library. (As a specialized problem, that might be easier to solve depending on the typeface.) A much better example that works quite well is a picture of someone holding a book: http://i.imgur.com/3JWs64x.jpg http://i.imgur.com/3JWs64x.jpg Magic . Read this to yourself. Read it silently Don't move your lips. Don’t make a suund Listen to yourself. Listen without hearing What a wonderfully weird thing, huh? NOW MAKE THIS PART LOUD! SCREAM IT IN YOUR MIND! DROWN EVERYTHING OUT. Now, hear a whisper. A tiny whisper. New, read this next line with your best crotchety— old-man voice: “Hello there, sonny. Does your town have apost 0 Awesome! Who was that? Whose voice was that? It sure wasn’t yours! How do you do that? How?! Must be magic. Problems with this text: misspelled 'sound' as 'suund', didn't recognize the word 'anything', and mis-recognized 'a post office' as 'apost 0'. Not bad. Especially since two of three mistakes are on the edge of the page.
- minism 10y agoThe old man voice was spoken in my mind as Deckard Cain.
- holografix 10y agoI stayed a while and I listened
- jbhatab 10y agoit was instantaneous Deckard Cain for me.
- Ph0X 10y agoIsn't this an issue with the algorithm? Can someone try and see how it would perform if you simply upscale the image using normal bicubic interpolation? And if it performs much better, I feel like that should be a preprocessing option to scale up the image since it seems to do so poorly on small resolutions.
- maaaats 10y agoDoes it block while it works and do the work in several setTimeouts or how do they get it to report progress without freezing everything?
- bijection 10y agoTesseract.js uses webworkers in the browser.
- deleted 10y ago[deleted]
- jamesville 10y agoAnyone have this working in a pre-made cluster type environment? Trying to OCR 100k+ pages on my desktop is not very fun. Also, has anyone had any luck with handwriting recognition? Need to convert some handwritten time records, so really just numbers & signatures.
- iplaw 10y agoHOW is there not a better, almost 100% accurate OCR tool? I routinely (daily) need to OCR PDF files. The PDF files are not scans. They are PDF files created from a Word file. The text is 100% clear, the lines are 100% straight, and the type is 100% uniform. And, yet, Microsoft and Google OCR spits out gibberish that is full of critical errors. From a problem solving perspective, this seems like an incredibly easy problem to solve in this exact use case. That is, PDFs generated from text files. Identify a uniform font size (prevent o-to-O and o-to-0 errors), identify a font-family (serif, sans-serif, narrow to particular fonts), and OCR the damn thing. And yet, the output is useless in my field.
- nabla9 10y agoWhy do you use OCR and not PDF to text conversion?
- angry-hacker 10y agoProbably because the pdf is just a big image file? If I understand correctly. Otherwise it should be just copy paste from pdf.
- pitaj 10y ago> The PDF files are not scans. They are PDF files created from a Word file. I am unsure as to why he can't just copy / paste.
- iplaw 10y agoApologies. The PDFs that we deal with are digital-native, but do not have embedded text and are not searchable. I simply want to OCR the PDF and spit the text into a Word/text file. I don't even care about perfect formatting, that's easy to fix. I do care about perfect OCR. That's crucial.
- iplaw 10y agoRight. It's an image PDF generated from a text file, so there are no digital-to-analog-to-digital errors introduced. These files should be perfect OCR candidates, but everything that I've found is full of errors, missing portions of sentences, rearranged fragments, etc.
- deleted 10y ago[deleted]
- jaytaylor 10y agoFor those who may be interested; I threw together a quick proof-of-concept in Go for exposing tesseract via a web API: https://github.com/jaytaylor/tesseract-web https://github.com/jaytaylor/tesseract-web
- SmokyBourbon 10y agoI would love to see a Chrome extension for extracting data from a page that uses both DOM navigation and OCR to produce reliable results.
- sanketbajoria 10y agoAwesome
- userbinator 10y agoTesseract was one of the best publicly-available CAPTCHA solvers when I was playing around with that stuff a few years ago; I remember somewhere in the neighbourhood of 90%+ accuracy on ReCAPTCHA, no wonder they've changed those considerably since then to make it difficult even for humans.
- zelon88 10y agoDoes this mean I can implement Tesseract on my home server without using php's shell_exec to perform magic on my files? I can just use Jscript instead? Cool! My current HRCloud2 project could benefit greatly if I ever get around to it. Currently I make the php interpreter jump through hoops and move stuff all over the place to OCR images and docs. This could save a ton of time and shift the processing to the client instead of my server.
- bijection 10y agoYep, this is completely client side :) You can even host the external language files yourself as described in the readme: https://github.com/naptha/tesseract.js#local-installation https://github.com/naptha/tesseract.js#local-installation
- mgalka 10y agoAwesome! The ability to OCR video in a browser opens up so many interesting possibilities.
- artf 10y agoSorry guys, probably a stupid question (googled quickly, doesn't worked), but does this kind of stuff involve ML? Do I need to train it?
- daliwali 10y agoThe title and description are very misleading: this is technically "pure JavaScript" but the JS is compiled from the original C++ library of the same name using emscripten. I think "pure JS" would imply that all of its sources are written in JS which is not the case here. It's mostly the C++ code doing the actual work, with a little JS wrapper on top.
- z3t4 10y agoMore instructions, like how to train it, would be nice.
- codemode 10y agoIs it true, that original implementation of tesseract exexuted from commandline is faster than javascript translated version?
- codemode 10y agoAccording to my tests this is true, but for curiosity can anyone get equivalent or better speed with tesseract.js? This is nice but I don't need client side processing so is there any reason to pick up tesseract.js?
- tofupup 10y agoneat
- niutech 10y agoHow does it compare with Ocrad.js?