8 ms·
Base64.ai – Extract text, data, photos and more from all types of docs
- alierkurt 6y agoBase64.ai is a cloud API that can extract data, photos, and signatures from all types of documents. We have prebuilt models for IDs, driver licenses, passports, visas, invoices, and many more document types. The integration is only a single API call.
- Radim 6y agoAli, you may want to add some "About us" page. Sending such sensitive data "into the cloud" is no joke, for any company.
- alierkurt 6y agothanks for the feedback. as replied earlier, Base64.ai neither stores the images you sent, nor the extracted data. We provide the power and extensibility of the cloud without the risks of a data breach. Base64.ai complies with GDPR requirements too. Our SOC-2 compliance report details the extend of security measures we take for your data. Happy to share the report for your review under MNDA.
- markdown 6y agoA dollar per page? :O
- aabhay 6y agoThat’s more expensive than manual data entry!
- pseudosavant 6y agoAnd has a 1 second response. That is worth something.
- mike_d 6y agoYeah, this will never take off unless they can get pricing below 1c per call. Manual data entry for an _entire page_ of text is about 15c, or 10c at volume.
- kazinator 6y agoMaybe the plan is to compete on latency? Say someone wants to regularly extract content from the same kind of document, and wants it fast, like 500 milliseconds.
- purplecats 6y agoperhaps its cheaper than paying for 401k/benefits etc
- djohnston 6y agogenerally people doing these sorts of tasks arent full time employees, rather contractors
- woadwarrior01 6y agoThe pricing makes me wonder if it's an AAI (Artificial Artificial Intelligence) service?
- eejjjj82 6y agowith this pricing model I'd expect they're just reselling something like the GCP OCR APIs, most likely with some domain specific value adds
- alierkurt 6y agoyes but we have to cover the costs as well :) we're flexible on pricing based on the volumes. what you see there differs based on the needs but we always aim to find a common price point for all parties.
- alierkurt 6y agoWe have startup plans that start free and runs at 10 cents/page after volume discounts. We also offer prices in local currencies. Happy to work on a deal that works for you. We are a pure AI company, i.e. there is no human-in-the-loop. We are and strive to be more accurate than manual labor, and our processing time is 1 second rather than minutes-to-hours. Also our AI is naturally unbiased and does not discriminate.
- markdown 6y ago> We are a pure AI company, i.e. there is no human-in-the-loop. In that case, shouldn't it a fraction of a penny rather than a whole dollar? Automation is supposed mean lower costs.
- m3nu 6y agoMaybe a kind to artificial AI with lots of manual verification and templates? Hence the price.
- llarsson 6y agoIt says that it does not store the submitted data. If true, then it's essentially just a trained model that we get to invoke for a dollar per API call to get the output from.
- oebilgen 6y agoThis is accurate. We train every document models upfront and make it available via the API. We believe well-trained, high-quality models don't need retraining, just like humans don't need to re-learn reading every day.
- robarr 6y agoWhat about the liabilities of sharing data with a third party? Your are sending all kind of data to a third party processor. Edit: I am not being critical, I am really asking.
- lionkor 6y agoAll I can find on this is > Base64.ai SOC 2 compliancecertifies our bank-level security standards. Our API does not store your data to prevent possible data breaches. All API traffic must be authenticated and encrypted over HTTPS. Sounds... Good enough? I mean, for what it is, it sounds like it's at least trying.
- nemoniac 6y agoIn Europe that falls far short of the requirements of GDPR law for personal data.
- jgtrosh 6y agothat's not a question with an answer, it's a negative point for any third party processor such as this one.
- alierkurt 6y ago@robarr that's a fair question to ask. Briefly, Base64.ai neither stores the images you sent, nor their extracted data. We provide the power and extensibility of the cloud without the risks of a data breach. Base64.ai complies with GDPR requirements too. Our SOC-2 compliance report details the extend of security measures we take for your data. Happy to share the report for your review under MNDA.
- voiper1 6y agoI looked into OCR a while ago for some hundreds of thousands of pages of PDF. All hosted offerings would end up costing quite a bit. After looking at options and few tests, I figured I'd use https://github.com/jbarlow83/OCRmyPDF https://github.com/jbarlow83/OCRmyPDF It converts the PDF to an image for Tesseract and then recreates the PDF with the text copy-able. It won't identify the address part of a driver's license, but that wasn't necessary for this project.
- Darkphibre 6y agoInteresting! I've been thinking of running OCR on video frames. I'd also like to do speech-to-text extraction for searching my archives later (have about 4TB of video to trawl through, and desire text-based search capabilities). It's an interesting space to explore, but everything's been moving to web-service at a cost-prohibitive model.
- voiper1 6y agoShould be able to use ffmpeg[0] to extract a single frame each second/keyframe (doubtful it's worth doing every single frame) and then pass it to tesseract. For speech to text.. if english, try mozilla's deepspeech? https://github.com/mozilla/DeepSpeech https://github.com/mozilla/DeepSpeech Might be fun to try. [0] https://stackoverflow.com/questions/27568254/how-to-extract-1-screenshot-for-a-video-with-ffmpeg-at-a-given-time/27573049 https://stackoverflow.com/questions/27568254/how-to-extract-...
- Darkphibre 6y agoYup, was planning to use ffmpeg (or, more likely, OpenCV), and a subset of the frames. Thanks so much for the tip on DeepSpeech!
- ghgr 6y agoFor speech-to-text extraction you can try Silero [1]. Free software (AGPL-3.0 License), fast, highly accurate and extremely simple to deploy (I have no affiliation with them). [1] https://github.com/snakers4/silero-models https://github.com/snakers4/silero-models
- hartem_ 6y agoBase64.ai has nailed time to value for customers. It’s pretty straightforward to integrate with and their extensive list of models makes it really easy to process a wide variety of document types. We used it Bardeen.ai and couldn’t have been happier. Kudos for a great service!
- deleted 6y ago[deleted]
- opheliate 6y agoSomewhat confused by the naming choice here. Naming your company after something as fundamental as base64 encoding seems to inevitably lead to confusion down the line.
- arthurcolle 6y agoDoesn't matter. Facebook will buy them for 2B in another year and a half and will get folded into the mix
- oebilgen 6y agoSorry you find it confusing. Our vision is to provide AI services for everything in base64 format; images, videos, sounds, etc.
- tzfld 6y agoThe demos are not working for me. Not finish processing.
- je42 6y agojust tried the Android app. very slow. didn't return any result for simple basic text card (black text on white paper).
- je42 6y agoanother text. some warranty info of some product in multiple languages was recognized as "drivers license": "First Name": "400 MHz ~2433,5MH:" "Issuing authority": "0MHz~2833.5MH:" may be the requirements about the documents the system can accurately recognize need to be explained in the app.
- oebilgen 6y agoIf you want our AI to learn warranty info documents, we're happy to work together. Let's meet http://base64.ai/meeting http://base64.ai/meeting
- solarkraft 6y agoStay tuned for my new biotech startup, http.ai. What was the process resulting in this name?
- oebilgen 6y agohttp.ai is a cool name too! Our vision is to provide AI services for everything in base64 format; images, videos, sounds, etc.
- Farbklex 6y agoDoes your solution have any unique features or benefits in comparison to existing solutions like Acuant, MicroBlink or Regula? Those already classify various documents and extract the data pretty well. https://www.acuant.com/idscan-data-capture-software/ https://www.acuant.com/idscan-data-capture-software/ https://microblink.com/products/blinkid https://microblink.com/products/blinkid https://api.regulaforensics.com/ https://api.regulaforensics.com/
- oebilgen 6y agoThey are good too, but we have products and services that match their offerings at a fraction of their cost. We also offer products that they don't provide. Our AI is capable of analyzing sound data (speech to text). It is extensible to add your custom forms and document types. We provide a cloud API and RPA components for UiPath, Bardeen and other RPA providers. We built Base64.ai so that you won't need a new vendor for new document types and platforms. Happy to meet over Zoom if you want to learn more https://base64.ai/meeting https://base64.ai/meeting
- emmelaich 6y agoAlso filingdb - https://filingdb.com/b/pdf-text-extraction https://filingdb.com/b/pdf-text-extraction
- RamRavi 6y agoThanks for sharing this informative content, Great work. To crack Scrum master interview: https://leanpitch.com/blogs/scrum-master-interview-questions https://leanpitch.com/blogs/scrum-master-interview-questions
- _joel 6y agoReally confusing name
- cochne 6y agoSad that it doesn't seem work for HTML! Maybe I will try taking a screenshot... Otherwise cool though, looks very promising.
- m3nu 6y agoIt's not really working. Tried 2 English PDF invoices. Normal format. One came back empty, the other only had the amount right. I'm assuming they only trained on some specific documents (passport of country X, etc) and all others don't work. If someone processes the same document all the time, then my invoice2data project may work better and is open source. It's based on Regx, rather than machine learning: https://github.com/invoice-x/invoice2data https://github.com/invoice-x/invoice2data
- ZeroCool2u 6y agoWe've been working with Google Cloud on a very difficult data extraction problem for about 6 months now. Seeing very impressive results with their DocumentAI service. One of my teammates is planning to try this out on some of our data this afternoon though!
- oebilgen 6y agoThank you! We're here to help. Please pick a time in our calendar https://base64.ai/meeting https://base64.ai/meeting