Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
throwaway4496
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
throwaway4496
1y ago
You're wrong. There is nothing inherent in "rendering" that means "raster or pixels". You can render PDFs or any format into any format you want, including XML for example. In fact, in majority of PDFs, a large part
32.
▲
by
throwaway4496
1y ago
You just defined it, "business value creation", from there, you narrow it down, short term vs long term, strategic vs revenue generating, and so on. The idea that you can't measure productivity is align with people who think
33.
▲
by
throwaway4496
1y ago
If you actually read what I have been saying and commenting, you would realise how silly your comment is.
34.
▲
by
throwaway4496
1y ago
Your mistake is in thinking that computers "see the image", second, you somehow think the output of OCR is different from a PDF engine that renders it into structured data/text.
35.
▲
by
throwaway4496
1y ago
You're thinking "rendering structured data" means parsing PDF as text. That is just wrong. Carefully read what I said. You render the PDF, but into structured data rather than raster. If you still get letters in reverse when
36.
▲
by
throwaway4496
1y ago
Sir, some of our cars breaks down every now and then, so we push them, because it happens every so often and we want to avoid it, we have implemented a policy of pushing all cars instead of driving them at all times. This removes the proble
37.
▲
by
throwaway4496
1y ago
There is also PDF to HTML, PDF to Text, MuPDF also has PDF to XML, both projects along with a bucketful of other PDF toolkits have PDF to PS, and there is many many XML, HTML, and Text outputs for PS. Rastering and OCR'ing PDF is like
38.
▲
by
throwaway4496
1y ago
It is a bit more involved, we have a rule engine that is fine tuned over time and works on most of invoices, there is also an experimental AI based engine that we are running in parallel but the rule based Engine still wins on old invoices.
39.
▲
by
throwaway4496
1y ago
Yes, and don't for a second think this approach of rastering and OCR'ing is sane, let alone a reasonable choice. It is outright absurd.
40.
▲
by
throwaway4496
1y ago
Okay, this sounds like "because some part of the road is rough, why don't we just drive in the ditch along the road way all the way, we could drive a tank, that would solve it"?
41.
▲
by
throwaway4496
1y ago
Computers "don't understand" things. They process things, and what you're saying is called layoutinng which is a key part of PDF rendering. I do understand for someone unfamiliar with the internals of file formats, parsi
42.
▲
by
throwaway4496
1y ago
You can use an existing readymade renderer to render into structured data instead of raster.
43.
▲
by
throwaway4496
1y ago
But you have to render the PDF to get an image, right? How do you go from PDF to raster?
44.
▲
by
throwaway4496
1y ago
None of that changes the fact that to get a raster, you have to solve the PDF parsing/rendering problem anyways, so might as well get structured data out instead of pixels so that it now another problem (OCR).
45.
▲
by
throwaway4496
1y ago
We process invoices from around the world, so more PDF generators than I care to count. It is hard a problem for sure, but the problem is the rendering , you can't escape that by rastering it, that is rendering. So it is absurd to pre
46.
▲
by
throwaway4496
1y ago
I would hire someone who understands PDFs instead of doing the equivalent of printing a digital document and scanning it for "digital record keeping". Stop everything and hire someone who understands the basics of data processing
47.
▲
by
throwaway4496
1y ago
Jesus Christ. What other approaches did you try?
48.
▲
by
throwaway4496
1y ago
No model can do better on images than structured data. I am not sure if I am on crack or you're all talking nonsense.
49.
▲
by
throwaway4496
1y ago
I do PDF for a living, millions of PDFs per month, this is complete nonsense. There is no way you get better results from rastering and OCR than rendering into XML or other structured data.
50.
▲
by
throwaway4496
1y ago
Hard no. LLMs aren't going to magically do more than what your PDF rendering engine does, rastering it and OCR'ing doesn't change anything. I am amazed at how many people actually think it is a sane idea.
51.
▲
by
throwaway4496
1y ago
But all those problems exist when rendering into a surface or rastering. I just don't understand how one thinks, this is a hard problem, let me make it harder by solving the problem into another kind of problem that is just as hard as
52.
▲
by
throwaway4496
1y ago
You can, if you have clear objectives.
53.
▲
by
throwaway4496
1y ago
Cryptographic proof of job experience? Please explain more. Sounds interesting.
54.
▲
by
throwaway4496
1y ago
This is the parallel of some of the dotcom peak absurdities. We are in the AI peak now.
55.
▲
by
throwaway4496
1y ago
How is it reasonable to render the PDF, rasterize it, OCR it, use AI, instead of just using the "quality implementation" to actually get structured data out? Sounds like "I don't know programming, so I will just use AI&q
56.
▲
by
throwaway4496
1y ago
So you parse PDFs, but also OCR images, to somehow get better results? Do you know you could just use the parsing engine that renders the PDF to get the output? I mean, why raster it, OCR it, and then use AI? Sounds creating a problem to us