Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
constantinum
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
constantinum
1y ago
1. Bill Cunningham New York 2. General Magic(2018) 3. The Armstrong Lie 4. Icarus (2017 film) 5. Man on Wire 6. Baraka & Samsara 7. BBC Planet Earth 8. Finding Vivian Maier
32.
▲
Reducto Raises $108M to Shape the Future of AI Document Intelligence
(reducto.ai)
2 points
by
constantinum
1y ago
|
0 comments
33.
▲
by
constantinum
1y ago
anyone looking for an ocr or text pre-processor that maintains the layout(tables, forms) try LLMWhisperer > https://pg.llmwhisperer.unstract.com/
34.
▲
Iconfinder will permanently close on November 15, 2025
(support.freepik.com)
4 points
by
constantinum
1y ago
|
0 comments
35.
▲
by
constantinum
1y ago
Other players: 1. Trellis (YC W24) 2. Roe AI (YC W24) 3. Omni AI (YC W24) 4. Reductor (YC W24) Other players(extended): 1. Unstract: Open-source ETL for documents ( https://github.com/Zipstack/unstract ) 2. Datalab: Make
36.
▲
by
constantinum
1y ago
Mobile phones from the past 15 years. Film cameras. Vinyl records (you have to see the look on today's kids' faces when they hear music from playing records).
37.
▲
by
constantinum
1y ago
Photography, music, books, cooking
38.
▲
by
constantinum
1y ago
In a way only marketeers think hubspot does very good product marketing. (i’m a user of hubspot for 8 years or so)
39.
▲
by
constantinum
1y ago
Another common stack that is commonly use is Langchain + Pydantic https://unstract.com/blog/comparing-approaches-for-using-llm...
40.
▲
by
constantinum
1y ago
There is also Unstract open-source. Structured data extraction + ETL. https://github.com/Zipstack/unstract
41.
▲
by
constantinum
1y ago
Thanks for the tip.
42.
▲
by
constantinum
1y ago
1. https://cal.com/ 2. https://cap.so/ 3. https://n8n.io/ 4. https://plausible.io/ 5. https://www.papermark.com/ 6. https://www.tooljet.ai/
43.
▲
by
constantinum
1y ago
Getting past 200 pages is the tough part. Hope I’ll get through. Also, getting used to so many characters with unfamiliar Russian names is slowing things down. Let's see. Any tips and tricks for reading the magnum opus? Would help!
44.
▲
by
constantinum
1y ago
War and peace - third attempt
45.
▲
by
constantinum
1y ago
Financial Times and NewYorker
46.
▲
by
constantinum
1y ago
TL;DR "With all the pieces on the board, the key to Romania’s Olympiad success is three-fold: put the best students in the same classrooms, put the best teachers with the best students, and then incentivize schools, teachers, and stude
47.
▲
by
constantinum
1y ago
There are a few other reasons why PDF parsing is Hell! > https://unstract.com/blog/pdf-hell-and-practical-rag-applica...
48.
▲
by
constantinum
1y ago
Other tools worthy of mention that help with OCR'ing PDF/Scans to markdown/layout-preserved text: LLMWhisperer(from Unstract), Docling(IBM), Marker(Surya OCR), Nougat(Facebook Research), Llamaparse.
49.
▲
Retab: The developer starter pack for document processing
(retab.com)
2 points
by
constantinum
1y ago
|
0 comments
50.
▲
by
constantinum
1y ago
Other PDF parsing woes include: 1. Identifying form elements like check boxes and radio buttons. 2. Badly oriented PDF scans 3. Text rendered as bezier curves 4. Images embedded in a PDF 5. Background watermarks 6. Handwritten documents PD
51.
▲
by
constantinum
1y ago
There is also Unstract(open-source) that helps process structured data extraction. Key differences: 1. Unstract has a Pre-processing layer(OCR). Which converts documents into LLM readable formats.(helps improve accuracy, and control costs)
52.
▲
by
constantinum
1y ago
Reminds me of Team Lab Borderless in Tokyo, Japan
53.
▲
Understanding why deterministic output from LLMs is nearly impossible
(unstract.com)
2 points
by
constantinum
1y ago
|
0 comments
54.
▲
by
constantinum
1y ago
just curious? Can we integrate Apple Music Subscription into LongPlay? Or This works only when you have your own collection of music that you own?
55.
▲
by
constantinum
1y ago
I'm getting the same loading loop too.
56.
▲
Lumigator: The Dev Tool for AI Model Evaluation
(mozilla.ai)
1 points
by
constantinum
1y ago
|
0 comments
57.
▲
Any-agent: A single interface to use and evaluate different agent frameworks
(mozilla.ai)
2 points
by
constantinum
1y ago
|
2 comments
58.
▲
by
constantinum
1y ago
LLMs are not yet there for complex and diverse document parsing use cases, especially at an enterprise scale (processing millions of pages). Some of the reasons are: Complex layouts, nested tables, tables spanning multiple pages, checkboxes
59.
▲
by
constantinum
1y ago
this is great!
60.
▲
by
constantinum
1y ago
Thank you
More ›