Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mrkn1
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
Show HN: Local CPU OCR for images, PDFs, webpages
(github.com)
3 points
by
mrkn1
4mo ago
|
0 comments
32.
▲
by
mrkn1
4mo ago
https://github.com/kouhxp/fftext
33.
▲
by
mrkn1
5mo ago
For 100% local CPU fact checking, I made this: https://news.ycombinator.com/item?id=48301003
34.
▲
by
mrkn1
5mo ago
A lot of my queries are summarize/explain/fact check, and these are covered 100% on my CPU locally [0], reducing frontier model reliance [0] https://news.ycombinator.com/item?id=48301003
35.
▲
Show HN: Local CPU model for fact-checking, summarizing, explaining text
(github.com)
6 points
by
mrkn1
5mo ago
|
0 comments
36.
▲
by
mrkn1
5mo ago
just released new version that implements your ideas
37.
▲
by
mrkn1
5mo ago
new version released is up to 3x faster on CPU. Let me know!
38.
▲
Show HN: CPU transcription for YouTube/TikTok/X, now 3x faster and diarization
(github.com)
3 points
by
mrkn1
5mo ago
|
0 comments
39.
▲
Advancing Mathematics Research with AI-Driven Formal Proof Search
(arxiv.org)
1 points
by
mrkn1
5mo ago
|
0 comments
40.
▲
by
mrkn1
5mo ago
thank you being thorough clipboard: rn input is treated like any other source, so text gets written to ./textsnaps/clipboard_ocr.txt, and stdout just prints that path. Nothing goes back to the clipboard in this version (stay tuned
41.
▲
by
mrkn1
5mo ago
Great question. I'm not familiar with docling-serv but pretty different beasts from what I gathered. Docling is a heavier pipeline (actually uses GPU).textsnap is the opposite: single-file CLI, small VLM running on plain CPU cores, one
42.
▲
by
mrkn1
5mo ago
thanks! yapsnap is audio to text, and textsnap is image to text. Both have been daily use cases for me for a while. And yes, the feedback on yapsnap encouraged me to also release textsnap on github
43.
▲
Show HN: CPU-only fast OCR for screenshots, images, PDFs, webpages
(github.com)
4 points
by
mrkn1
5mo ago
|
9 comments
44.
▲
by
mrkn1
5mo ago
Reading this means a lot, thank you! Even faster version in the coming days, stay tuned and PRs welcome!
45.
▲
by
mrkn1
5mo ago
Added Diarization / Speaker Separation that is fast and CPU only. Thank you all for the great feedback and support. PRs welcome! yapsnap "https://www.youtube.com/watch?v=NzKJ-xO-VhE" --diarize SPEAKER_00
46.
▲
by
mrkn1
5mo ago
done! new version separates speakers on CPU fast
47.
▲
by
mrkn1
5mo ago
done! just pushed a new version with CPU diarization
48.
▲
by
mrkn1
5mo ago
Kroko's website says benchmarks aren't formalized yet. FWIW, this url says 5% WER for English [0]. though it doesn't specify the dataset, so not directly comparable to Parakeet's 6.32 on the Open ASR Leaderboard Best way
49.
▲
by
mrkn1
5mo ago
in the roadmap!
50.
▲
by
mrkn1
5mo ago
will work on it, that would be neat. I love pyannote but not happening on CPU at reasonable speeds lol
51.
▲
by
mrkn1
5mo ago
thank you! good use case, what hetzner box specs have you chosen?
52.
▲
by
mrkn1
5mo ago
small, ONNX-optimized models designed specifically for low-latency CPU streaming, so it avoids overhead of large transformer arch and GPU memory transfers
53.
▲
by
mrkn1
5mo ago
thanks! making a note of the feature request
54.
▲
by
mrkn1
5mo ago
Good point! I haven't found a faster way to consume info than reading. But depends on the type of learner you are (visual, auditory, hands-on/interactive, etc)
55.
▲
by
mrkn1
5mo ago
thank you!
56.
▲
by
mrkn1
5mo ago
Just download the model for your preferred language, all hosted on the Kroko-ASR collection here: https://huggingface.co/Banafo/Kroko-ASR/tree/main Right now you have Dutch, French, Portuguese, Spanish, Germa
57.
▲
by
mrkn1
5mo ago
Yes fair point, asr cached and exposed. I meant to draw the line more on fetchable or not.
58.
▲
by
mrkn1
5mo ago
thanks for running it Niraj. I see something similar on my machine, which still surprises me every time lol
59.
▲
by
mrkn1
5mo ago
I appreciate the perspective! higher ceiling than I'd put on it, but if it gets there awesome. PRs welcome!
60.
▲
by
mrkn1
5mo ago
Youtube has transcripts on most videos, not all. The others don't expose them. If you mean the "transcript APIs" for TikTok/IG/X, they are all transcribing audio like yapsnap does. If you have a way to pull native o
More ›