Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ses425500000
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
ses425500000
2y ago
I’m sorry that I didn’t know that detail, thank you so much for letting me know! I’ll read AGPL-3.0 license more carefully and check if it’s okay with MIT. If not, I’ll fix license or change model. really appreciate your help!
2.
▲
by
ses425500000
2y ago
Thanks a lot! Yeah, theoretically the pipeline handles math and special symbols fine, and from my testing it worked well. But I didn’t test much on other languages or encodings, so if there’s any weird behavior, please let me know and I’ll
3.
▲
by
ses425500000
2y ago
Haha not exactly like predicting actual questions. Just trying to find patterns or what topics show up often. I made this to help my study, didn’t think people would care this much.
4.
▲
by
ses425500000
2y ago
Yeah, prompt injection is good point. For now, I try separate instruction and data by using JSON format, and run it in sandbox. Not perfect maybe, but I will try add small explanation in README so people can check it better.
5.
▲
by
ses425500000
2y ago
Yeah, hallucination part was also one thing I was worry about. So I make LLM only run after OCR step, and I put simple check to not change correct text. I will try to show real examples and hallucination rate too. Thanks for feedback! This
6.
▲
by
ses425500000
2y ago
Haha good catch! I’m 19 and from Korea, so I’ve been using an LLM to help with replies since my English isn’t perfect yet. But I designed and built the project myself (with help from some open models/tools) — just wanted to communicate
7.
▲
by
ses425500000
2y ago
Yep — this project uses a pre-trained DocLayout-YOLO model released under an open license by the original authors. No additional datasets were used for training. All sample data in the repo is either synthetic, publicly available, or user-g
8.
▲
by
ses425500000
2y ago
Thanks for the insightful comment! You’re absolutely right — organizing extracted data into a coherent, semantically meaningful structure is critical for high-quality ML training. Right now, the pipeline focuses on generating OCR outputs op
9.
▲
by
ses425500000
2y ago
Great question — I’m using traditional OCR engines for the initial text extraction (e.g., MathPix, Google Vision), but then I apply generative AI models in a second stage to refine the output. This includes removing noisy or irrelevant elem
10.
▲
by
ses425500000
2y ago
Thanks! Yes — I’m definitely planning to update and refine the project over time. This initial release is mostly a working prototype to demonstrate the full pipeline logic, and I’ll continue improving stability, modularity, and usability. A
11.
▲
by
ses425500000
2y ago
Thanks for sharing — Marker is a great tool, especially for human-readable formatting! In contrast, this project focuses less on preserving the visual layout for human readers, and more on extracting structured semantic data for machine lea
12.
▲
by
ses425500000
2y ago
Yep — some components currently rely on external APIs (e.g. OpenAI, MathPix), primarily for stability and ease of deployment during early release. But I’m planning to support fully local inference in the future to eliminate API key dependen
13.
▲
by
ses425500000
2y ago
Yeah — I ran into that exact problem during early testing. The prompt has since been adjusted to prevent GPT from auto-translating non-English text (Korean, Japanese, etc.). If it still misbehaves in any edge cases, feel free to open an iss
14.
▲
Show HN: OCR pipeline for ML training (tables, diagrams, math, multilingual)
(github.com)
170 points
by
ses425500000
2y ago
|
38 comments