Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
adbarba
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
The Rick Roll programming language
(github.com)
1 points
by
adbarba
3y ago
|
0 comments
2.
▲
Collection of stand-alone Python machine learning recipes (2021)
(github.com)
84 points
by
adbarba
3y ago
|
8 comments
3.
▲
by
adbarba
3y ago
Thanks! Here is what I put together in the docs, you could basically preprocess/render/filter the webpages with the software of your choice and then pass the result to trafilatura: https://trafilatura.readthedocs.io
4.
▲
by
adbarba
3y ago
Concerning tooling I'd say you have two different worlds, JavaScript and Python, each with a series of tools to tackle such tasks. It's not easy to compare them directly because of varying software environments and I haven't
5.
▲
by
adbarba
3y ago
Regarding content extraction it's more accurate than newspaper3k (especially for languages other than English) and it entails more information: metadata, text, and comments. It works out of the box in most cases so no need to write a p
6.
▲
by
adbarba
3y ago
Author here, nice to see the package on the HN's front page this morning and thanks for the kind words! Just created an account to participate in the discussion, I'll try to answer your questions.