Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kapitalx
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
Show HN: I Built Grid View for Switchboard – Claude Code CLI Manager
(github.com)
2 points
by
kapitalx
6mo ago
|
0 comments
2.
▲
by
kapitalx
7mo ago
Fixed some bugs and much more stable now. https://github.com/doctly/switchboard/releases/tag/v0.0.8
3.
▲
by
kapitalx
7mo ago
Thank you. Yeah I looked at a few, another one is air.dev. The problem is they are recreating the actualy coding pane. Switchboard runs terminal directly.
4.
▲
Show HN: A desktop app for managing Claude Code sessions
(github.com)
5 points
by
kapitalx
7mo ago
|
3 comments
5.
▲
by
kapitalx
1y ago
https://doctly.ai We're building Doctly.ai - PDF Extraction with AI. We started out with document conversions to Markdown but quickly realized that most use cases were for JSON conversion. We recently launched our "Ext
6.
▲
by
kapitalx
1y ago
Check it out at https://doctly.ai
7.
▲
Show HN: Extractor Studio – The Fastest Way to Build PDF Extractors
(medium.com)
10 points
by
kapitalx
1y ago
|
2 comments
8.
▲
by
kapitalx
1y ago
This is approximately the approach we're taking also at https://doctly.ai , add to that a "multiple experts" approach for analyzing the image (for our 'ultra' version), and we get really good results. And
9.
▲
by
kapitalx
2y ago
To be fair, they didn't include themselves at all in the graph.
10.
▲
by
kapitalx
2y ago
In addition, gemini Pro 2.5 does really well with bounding boxes, but yeah not open source :(
11.
▲
by
kapitalx
2y ago
If you're limited to open source models, that's very true. But for larger models and depending on your document needs, we're definitely seeing very high accuracy (95%-99%) for direct to json extraction (no markdown in between
12.
▲
Show HN: We OCR'ed 60k pages of the JFK files with AI
(doctly.ai)
11 points
by
kapitalx
2y ago
|
3 comments
13.
▲
by
kapitalx
2y ago
I'll dig deeper into your code, but scanning your post does look like your are addressing this. That's great. If I do find anything, I'll share with you for comments before I publish the post.
14.
▲
by
kapitalx
2y ago
Exactly. You still have to be explicit in order to remove bias. Either by sorting the keys, or looking up specific keys. For arrays, I would say order still matters. For example when you capture a list of invoice items, you should maintain
15.
▲
by
kapitalx
2y ago
Great list! I’ll definitely run your benchmark against Doctly.ai (our PDF-to-Markdown service) specially as we publish our workflow service, to see how we stack up. One thing I’ve noticed in many benchmarks, though, is the potential for bia
16.
▲
by
kapitalx
2y ago
Customers are willing to pay for accuracy compared to existing solutions out there. We started out in need of an accurate solution for a RAG product we were building, but none of the solutions we tried were providing the accuracy we needed.
17.
▲
by
kapitalx
2y ago
Looks to be API only for now. Documentation here: https://docs.mistral.ai/capabilities/document/
18.
▲
by
kapitalx
2y ago
We've been getting great results with those aswell. But ofcourse there is always some chance of not getting it perfect, specially with different handwritings. Give it a try, no credit cards needed to try it. If you email me (ali@doctly
19.
▲
by
kapitalx
2y ago
Yes I used the API. They have examples here: https://docs.mistral.ai/capabilities/document/ I used base64 encoding of the image of the pdf page. The output was an object that has the markdown, and coordinates for
20.
▲
by
kapitalx
2y ago
Great question. The language models are definitely beating the old tools. Take a look at Gemini for example. Doctly runs a tournament style judge. It will run multiple generations across LLMs and pick the best one. Outperforming single gene
21.
▲
by
kapitalx
2y ago
Haha for sure. Naming isn't just the hardest problem in computer science, it's always hard. But at some point you just have to pick something and move forward.
22.
▲
by
kapitalx
2y ago
We'll definitely be doing more tests, but the results I got on the complex tests would result in a lower score and might not be worth the extra cost of the judgement itself. In our current setup Gemini wins most often. We enter multipl
23.
▲
by
kapitalx
2y ago
We need to update the examples on the front page. Currently for things that are considered charts/graphs/figures we convert to a description. For things like logos or images we do an image tag. You can also choose to exclude them.
24.
▲
by
kapitalx
2y ago
This is a good idea. We should publish a benchmark results/comparison.
25.
▲
by
kapitalx
2y ago
From my testing so far, it seems it's super fast and responded synchronously. But it decided that the entire page is an image and returned `` with coordinates in the metadata for the image, which is the entire
26.
▲
by
kapitalx
2y ago
Co-founder of doctly.ai here (OCR tool) I love mistral and what they do. I got really excited about this, but a little disappointed after my first few tests. I tried a complex table that we use as a first test of any new model, and Mistral
27.
▲
by
kapitalx
2y ago
Love the website design
28.
▲
by
kapitalx
2y ago
This applies to PDFs with lots of text. The smaller the text on the page, the more impacted it is.
29.
▲
Why OpenAI Models Struggle with PDF Extraction(and Why Gemini Fairs Much Better)
(medium.com)
5 points
by
kapitalx
2y ago
|
2 comments
30.
▲
by
kapitalx
2y ago
I'm the founder of https://doctly.ai , also pdf extraction. The hallucination in LLM extraction is much more subtle as it will rewrite full sentences sometimes. It is much harder to spot when reading the document and sounds
More ›