3 ms·
the problem with extracting information is not just limited to getting OCR results. the bigger problem while building something like this is extracting the fiel
by jumpskiphop 7y ago
the problem with extracting information is not just limited to getting OCR results. the bigger problem while building something like this is extracting the fields and understanding the structure of the document automatically. using some python OCR libraries, you'd probably get text results for a drivers license or a passport separately and process these results on separate rules written for each. with deep learning a non-template solution seems possible which will figure out which ID it is, where the name, address, relevant numbers are and put them in a structure.
- headbansown 7y agoIn the case of U.S. drivers licenses, there are standards for the 2D barcode that would make it very straightforward to parse: https://www.aamva.org/uploadedFiles/MainSite/Content/SolutionsBestPractices/BestPracticesModelLegislation(1)/BarCodeDataEncodingReqmtsBestPractice.pdf https://www.aamva.org/uploadedFiles/MainSite/Content/Solutio...