Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jazzido
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
jazzido
10y ago
Also see Walnut.io: https://thewalnut.io/app/release/34
32.
▲
When Documents Become Databases – R Wrapper for Tabula PDF Table Extractor
(blog.ouseful.info)
2 points
by
jazzido
10y ago
|
0 comments
33.
▲
What's Wrong with Open-Data Sites –and How We Can Fix Them
(blogs.scientificamerican.com)
4 points
by
jazzido
10y ago
|
0 comments
34.
▲
by
jazzido
10y ago
"All Watched Over by Machines of Loving Grace", an amazing 3-part documentary about computers, society, the shortcomings of systems-thinking. Also by Adam Curtis.
35.
▲
Mondrian-Rest: A REST Interface for Mondrian ROLAP Server
(github.com)
1 points
by
jazzido
11y ago
|
0 comments
36.
▲
by
jazzido
11y ago
Postgres users might find Teodor Sigaev's smlar [1] extension useful for this kind of thing. It implements cosine and other similarity metrics. [1] http://sigaev.ru/git/gitweb.cgi?p=smlar.git;a=summary
37.
▲
by
jazzido
11y ago
Can you elaborate on what scaling issues you've run into with Neo? I'm using it for my thesis project so I don't have very strict scalability requirements, but would love to know.
38.
▲
by
jazzido
11y ago
I tried really hard to use Titan for my project, but the tooling and the query language is —in my view— just atrocious. Setting up a running instance, trying to import my moderately complex property graph data into it, and navigating Tink
39.
▲
by
jazzido
11y ago
the little known csvfix ( https://bitbucket.org/neilb/csvfix ) is pretty good too.
40.
▲
by
jazzido
11y ago
PDFBox 1.8 less-than-great rendering engine forced us to include a separate library for that purpose only. Moving to PDFBox 2.0 is also on our roadmap. But the text extraction API in 2.0 has changed a lot too, so porting our engine would re
41.
▲
by
jazzido
11y ago
Hi. Can you share the PDF with us on our issue tracker? ( https://github.com/tabulapdf/tabula/issues ) We'd be happy to take a look at it
42.
▲
by
jazzido
11y ago
Hi. Tabula author here. We use JPedal for rendering pages as images. For parsing, we use Apache PDFBox. In the near future, we plan to render the PDFs client side with Mozilla's PDF.js