9 ms·
Principally, yes. However, the approach may be more nuisanced than that. If I were you, I would first pick a character recognition engine (which might have alre
by bhaavan 10y ago
Principally, yes. However, the approach may be more nuisanced than that. If I were you, I would first pick a character recognition engine (which might have already been well trained) to convert the image to text. Once the text is there, that might serve as a better feature to classify the content. Furthermore, I had recommend converting words in the text to word-embeddings/ vectors using a suitable Glove or Word2Vec dataset similar to your content.
While there are many benefits of end-to-end training, I don't think it might be best suited for this case. This is because we already know that the only useful feature in the document is the text, and not the contours or textures. Theefore wasting neurons in your neural network to learn the wastefulness of this features is just waste of resources. Furthermore, you benefit from even a larger corpus of learned data, which the char recognition engine has been trained on.
- starik36 10y agoOCR is my current approach. I am not really happy with it. The quality of OCRing leaves much to be desired, probably due to the documents themselves being haphazardly handled by the court personnel. OCR itself is a pretty CPU intensive activity and takes a significant time to complete for many documents. Thus, I was looking for a more advanced approach.
- jononor 10y agoWhat do you want to recognize? OCR is a better understood problem than a general neural net, so I think it's likely easier to improve its quality that to superseded the quality with image-based recognition.
- starik36 10y agoIdeally, I would like to get all the information from the page. Phone numbers, who is suing whom, case caption, etc... With OCR, you get bits and pieces of information, but because I don't know what the type of the document is, it is difficult to determine where, structurally speaking, this information resides on a page. If I could use AI to determine the type of the doc, I would know the structure of the document and I could then use OCR to pinpoint specific information on the page. Most court documents are created from templates.
- jononor 10y agoIf your OCR is unable to recognize the characters on a full page, I don't think it when scanning a region either? Unless using full resolution of a full image is somehow too much for the algorithm in use. But then I'd just subdivide the entire image into regions, and scan them all independently. This is also a trivially parallelizable task, so you can throw many servers at it, if the time to get results is an issue.
- barrkel 10y agoI think you'd have a better time doing your OCR on AWS - spread the CPU intensive activity across multiple machines. Even if the resulting text has errors, if you're doing further document classification, it would be better to do it on the text including the errors, than on the original images.
- perfmode 10y agoYou don't need perfect character recognition. It's just gotta be good enough. The way you determine good enough is by completing the pipeline and measuring the result.
- johncip 10y ago> OCR itself is a pretty CPU intensive activity and takes a significant time to complete for many documents. Leaving the quality part aside -- this job itself is easy to parallelize in that you can split it up by document or by page. Open option is to run each job in Lambda asynchronously, with the input being a URL to the page or the full document, and have the job call back to you with the text of the page (or put it on S3 as a text file, or add it to a message queue, or whatever works). Regarding splitting: we've been using a python wrapper + pdfium for splitting PDFs into page images on Lambda, with excellent results. To make the Lambda function, you'll either have to build e.g. Tesseract such that it fits into a 50MB zip, or download it while the Lambda function executes. LambCI has a set of docker containers that they've made for simulating lambda, and the "lambda:build" container makes building things easy and repeatable: https://github.com/lambci/docker-lambda https://github.com/lambci/docker-lambda. In a pinch, you can build on an Amazon Linux EC2 instance and it should work on Lambda, but you will have to be more careful about dynamic linking. As another option: I'm not sure if it's been mentioned, but you can also try a ready-made OCR service before packaging up Tesseract, like this one: https://algorithmia.com/algorithms/ocr/SmartOCR https://algorithmia.com/algorithms/ocr/SmartOCR. So anyway, the performance part has good solutions, at least. For fixing the accuracy: I know next to nothing about approximate string matching, but perhaps it would then be possible to do a fuzzy search over the text using something similar a Levenshtein automaton: https://en.wikipedia.org/wiki/Levenshtein_automaton https://en.wikipedia.org/wiki/Levenshtein_automaton. You may also want to take a look at this: https://en.wikipedia.org/wiki/Bag-of-words_model https://en.wikipedia.org/wiki/Bag-of-words_model More broadly, I'm sure that there are text-based document classification methods that are robust against sloppy OCR. It may just take some research on the main approaches people take to document classification -- it's not my area, but my understanding is that this is typically approached with statistical methods. Otherwise your spam filter would get defeated by typos.
- LeanderK 10y agoi don't know, maybe different approaches could be combined. Maybe the layout provides a clue for some types of court documents? You could calculate the probability function of prediction a certain type right (or just use the outputs of the NN, that depends on the problem) as a confidence value and only do the OCR as a last resort. Disclaimer: pretty new to ML
- starik36 10y agoThat is exactly the approach I had in mind, except I would use the knowledge of the document type (as determined by AI) to guide OCR to specific sections of the page to get information from it. But I know next to nothing about AI and ML - that's why I was asking this question.
- LeanderK 10y agowell, the problem is data. NN needs a lot (depending on the problem thousands or millions of samples). It will probably get very difficult to get this much data needed to train an ML algorithm the location of the relevant OCR text. Often ML problems way more experimental than "normal" coding. I would first try modelling it as a classification-problem and just do some cross-entropy validation to check the performance of the model. If it's useful, go with it, if not back to the drawing board. You will need some serious computing power, so either buy some GPUs or use the cloud. You could train a random forest based on the inputs of OCR and NN, if you want to get total ML. You would gain some interpretability (i don't know whether thats important, but i would guess it might) I am sorry that I can't give you a more concrete answer, these are just ideas. They are probably wrong. Like i said, i am a beginner and also don't really know the problem. Edit other idea: If you know the location of the relevant OCR-text, you could use the following approach: Use the NN for classification. It will return probability-like values for every category. Take the top 2 (or 3, or every top until they add up to 70 percent...idk). Then do some OCR for every category you have to check. If one is positive you have your result, if not run the others.
- iaml 10y agoI would argue that layout of the text on page could be a useful feature in this case.