5 ms·
> having to have DOM knowledge to select the paths is less inclusive Right click on the element in Firefox, click "inspect element" and it shows you the unique
by moehm 5y ago
> having to have DOM knowledge to select the paths is less inclusive
Right click on the element in Firefox, click "inspect element" and it shows you the unique selector for that element.
- ricardo81 5y agoAre you aware that some websites use randomised css ids and classes? Have to wonder why. Google search results being an example.
- playpause 5y agoThere’s usually some way to target what you want through descendent/sibling, tag name, and attribute selectors.
- ricardo81 5y agoYes, if you're determined enough you can scrape anything. Just not sure what the new thing here is.
- detaro 5y agowho has claimed there is something fundamentally new here?
- ricardo81 5y agoPresenting something on the home page of HN would suggest something novel. I must be missing the point because of the downvotes. Happy to be enlightened. Pick CSS selectors and scrape something? That's quintessentially scraping.
- onli 5y agoI'd say the novel part of the project is the Github actions/Gitlab CI integration. I haven't seen that used for RSS generation yet. That the selectors are stored in a config file is also unusual for these types of feed generation projects. Though Show HNs do not need to be novel to get to the top. And RSS related projects seem to go to the frontpage simply because of the fondness HN has for the technology. Also, don't forget https://xkcd.com/1053/ https://xkcd.com/1053/ :)
- moehm 5y agoYes, I am. And this project is obviously not designed to handle that. (It can't even do pagination.) But it can bring feed capabilities to simple, timeline like sites. Sites like most frameworks produce. And with the dev tools you can quickly find the needed dom path. It's limited, but easy to use. If it doesn't work, you need a real scraper which is an order of magnitude more complex. (I maintain some of them as well.)