Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
constantinum
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
91.
▲
by
constantinum
2y ago
I see a lot of comments on hallucination risk and the accumulation of non-traceable rotten data. If you are curious to try a better non-llm-based OCR, try LLMWhisperer. https://pg.llmwhisperer.unstract.com/
92.
▲
by
constantinum
2y ago
There is one with Langchain+pydantic+llmwhisperer https://unstract.com/blog/comparing-approaches-for-using-llm...
93.
▲
by
constantinum
2y ago
If you're looking for better accuracy and table layout preservation, give LLMWhisperer and Docling a try! Both keep tables tidy with a Markdown-like structure.
94.
▲
by
constantinum
2y ago
Tested it with the following documents: * Loan application form: It picks up checkboxes and handwriting. But it missed a lot of form fields. Not sure why? * Edsger W. Dijkstra’s handwritten notes(from Texas univ archive) - Parsing is good.*
95.
▲
by
constantinum
2y ago
The tool doesn't use any LLMs for processing/parsing the data. It parses and converts into raw text. The final output(raw text) of the parsing is then fed to LLMs for data extraction. e.g. Extracting data from insurance, banking,
96.
▲
by
constantinum
2y ago
The primary issue with LLMs is hallucination, which can lead to incorrect data and flawed business decisions. For example, Llamaparse( https://docs.llamaindex.ai/en/stable/llama_cloud/llama_parse... ) uses LLMs
97.
▲
by
constantinum
2y ago
Give LLMWhisperer a try. Here is a playground for testing https://pg.llmwhisperer.unstract.com/
98.
▲
by
constantinum
2y ago
This is not open-source. It has high accuracy and it is faster too. All you need is to point your documents to the API.
99.
▲
by
constantinum
2y ago
For instace Llamaparse( https://docs.llamaindex.ai/en/stable/llama_cloud/llama_parse... )uses LLMs for pdf text extraction, but the problem is hallucination. e.g > https://github.com/run-llam
100.
▲
by
constantinum
2y ago
Non-fiction as audiobooks
101.
▲
by
constantinum
2y ago
> The "best" models just made stuff up to meet the requirements. They lied in three ways: > The main difficulty of the is project lies in correctly identifying page zones; wouldn't it be possible to properly find the zo
102.
▲
by
constantinum
2y ago
Answering an important question: “Why is it difficult to extract meaningful text from PDFs?” https://unstract.com/blog/pdf-hell-and-practical-rag-applica...
103.
▲
by
constantinum
2y ago
I will try it with some complex layout PDFs or documents with tables. These documents have real business use cases for automation — insurance, banking, etc. Anyone here who wants to convert PDF documents or scanned images as it is preservin
104.
▲
by
constantinum
2y ago
The chapter where there is a comparison of techniques for structured data extraction is insightful.[1] Does anyone wants to explore more on the structured data extraction techniques, do refer to this piece [2] [1] https://www.sou
105.
▲
by
constantinum
2y ago
Whatsapp for friends and family Slack for work Telegram for apartment community email for very very close friends iMessage for those using ios
106.
▲
by
constantinum
2y ago
https://eagle.cool/ - image curation app Raycast Notability
107.
▲
by
constantinum
2y ago
Ahrefs or SEMRush Toss a coin and choose between the two.
108.
▲
by
constantinum
2y ago
GA4 is like a free trial offering to upscale to Looker and google cloud products(paid)
109.
▲
Ts_server: A web server proposing a REST API to large language models
(bellard.org)
2 points
by
constantinum
2y ago
|
1 comments
110.
▲
by
constantinum
2y ago
Reading from the comments, some of the common questions regarding document extraction are: * Run locally or on premise for security/privacy reasons * Support multiple LLMs and vector DBs - plug and play * Support customisable schemas *
111.
▲
Scaling Document Data Extraction with LLMs and Vector Databases
(timescale.com)
1 points
by
constantinum
2y ago
|
0 comments
112.
▲
by
constantinum
2y ago
The problem with using LLMs for OCR is hallucinations. It makes it impossible to use in business use cases such as insurance, banking and health/medical — which demands high accuracy or predictable inaccuracy rate. Not to mention hand
113.
▲
by
constantinum
2y ago
Unstract can be a good starting point. https://github.com/Zipstack/unstract Refer this > https://unstract.com/blog/extract-table-from-pdf/ And this > https://unstract.com
114.
▲
by
constantinum
2y ago
Any metric without segmentation might fall in this category. Conversion rate Win rate Mql to sql rate Total visitors
115.
▲
Next-Gen Virtual Office App.Remote Work Reimagined
(teracy.io)
3 points
by
constantinum
2y ago
|
0 comments
116.
▲
by
constantinum
2y ago
haskell
117.
▲
by
constantinum
2y ago
Unstract, if you are into automating document processing > https://github.com/Zipstack/unstract?tab=readme-ov-file#-eco...
118.
▲
by
constantinum
2y ago
Songs of Earth > https://www.youtube.com/watch?v=HTzJws8GbUQ&ab_channel=Stran...
119.
▲
by
constantinum
2y ago
John Cleese on Creativity In Management > https://www.youtube.com/watch?v=Pb5oIIPO62g&ab_channel=Video...
120.
▲
by
constantinum
2y ago
Ghost
More ›