4 ms·
One of the cases when AI not needed. There is very good working algorithm to extract content from the pages, one of implementations: https://github.com/buriy/py
by LeonidBugaev 2y ago
One of the cases when AI not needed. There is very good working algorithm to extract content from the pages, one of implementations: https://github.com/buriy/python-readability https://github.com/buriy/python-readability
- asadalt 2y agooh AI is optional here. I do use readability to clean the html before converting to .md.
- haddr 2y agoSome years ago I compared those boilerplate removal tools and I remember that jusText was giving me the best results out of the box (tried readability and few other libraries too). I wonder what is the state of the art today?
- jot 2y agoThis is worth having a look at: https://mixmark-io.github.io/turndown/ https://mixmark-io.github.io/turndown/ With some configuration you can get most of the way there.
- IanCal 2y agoHow do you achieve the same things without AI here using that tool?
- chrisweekly 2y ago"How do you do it without AI" is a question I (sadly) expect to see more often.
- IanCal 2y agoFeel free to answer then, how do you do the same functions this does with gpt(3/4) without AI? Edit - This is an excellent use of it, a free text human input capable of doing things like extracting summaries. It does not seem to be used at all for the basic task of extracting content, but for post filtering.
- cactusfrog 2y agoI think “copy from a PDF” could be improved with AI. It’s been 30 years and I still get new lines in the middle of sentences when I try to copy from one.
- IanCal 2y agoThat's a great use case, you might be able to do this if you've got a copy and paste on the command line with https://github.com/simonw/llm https://github.com/simonw/llm In between. An alias like pdfwtf translating to "paste | llm command | copy"
- genewitch 2y agoi've long assumed that is a "feature" of PDF akin to DRM. Making copying text from a PDF makes sense from a publisher's standpoint.
- hombre_fatal 2y agoMeh, it’s just the “how does it work?” question. How content extractors work is interesting and not obvious nor trivial. And even when you see how readability parser works, AI handles most of the edge cases that content extractors fail on, so they are genuinely superseded by LLMs.
- foundzen 2y agohow is it compared to mozilla/readability?
- asadm 2y agoit uses readibility but does some additional stuff like relink images to local paths etc., which I needed
- foundzen 2y agoI have had challenges with readability. The output is good for blogs but when we try it for other type of content, it misses on important details even when the page is quite text-heavy just like blog.
- asadalt 2y agoyeah that’s correct. i put a checkbox to disable readability filter if needed…
- fbdab103 2y agoI was honestly expecting it to be mostly black magic, but it looks like the meat of the project is a bunch of (surely hard won) regexes. Nifty.
- nyokodo 2y ago> I was … expecting it to be mostly black magic, but … the meat of the project is a bunch of … regexes Wait, regexes are the epitome of black magic. What do you consider as black magic?
- fbdab103 2y agoMacros? Any situation where code edits other code? Sure, I could not write a regex engine, but the language itself can be fine if you keep it to straightfoward stuff. Unlike the famous e-mail parsing regex.
- jot 2y agoLast time I tried readability it worked well with articles but struggled with other kinds of pages. Took away far more content than I wanted it to.