5 ms·
I'm sorry but there is nothing new here? This seems like a backwards step if anything. Usually when web scraping, I can just load in the HtmlAgilityPack (c#),
by hacker_9 8y ago
I'm sorry but there is nothing new here? This seems like a backwards step if anything.
Usually when web scraping, I can just load in the HtmlAgilityPack (c#), point it at a URL then write some functional code to extract the necessary data.
Even better, I'll examine the website in Fiddler and hope they have a data-view separation going on, and be able to just intercept the json file they load instead.
Worse case scenario I need to dynamically click on buttons etc, but this can usually be handled by selenium, or if they detect that just roll a custom implementation of CefSharp (again not hard, just download the nuget, and it lets you run your own custom javascript).
A new, more limited, language (with no IDE tools) is not the way to go. If anything a better web scraper just make the above processes I mentioned more seamless, for example combining finding/selecting of elements in chrome with codegen.
- curiousgal 8y ago> I'll examine the website in Fiddler Isn't that overkill compared to just using the built-in browser devtools Network tab?
- hacker_9 8y agoI find the Fiddler UI easy to use, plus you can add plugins such as converting the request straight to code. Up to you what tool you use of course.
- curiousgal 8y agoBoth Firefox and Chrome have a Copy As Curl feature which is really useful. I agree with you though, the UI sucks.
- ziflex 8y agoThe main advantage of this over your approach with HtmlAgilityPack is that Ferret can handle dynamic web pages - those that are rendered with JS. And also, it can emulate user interactions. But anyway, thanks for your feedback :)
- hacker_9 8y agoRight but in that case that implies a view-model separation, so you can usually just access the data file directly, which is usually json.
- ziflex 8y agoI'm sorry I do not fully understand what you mean. Imagine, that you need to grab some data from SoundCloud and, also imagine, they do not have a public API :) How would you do it without launching a browser?
- StavrosK 8y agoHe means that you look at the private API and use that.
- kbenson 8y agoI think what's being noted is that when the data comes as a data structure on the page, or as data passed back from an XHR request, you can just use that data directly and there's less page scraping to be done. This is generally how dynamic pages are created, out of a shipped data structure and rules to create the page out of it. If you have the data structure, it's generally much easier to parse than the page generated from it. That said, for pages that use a background request to fetch the data, this can be useful, as that data used to build the page isn't always kept around as a data structure (at least not one easily accessible) afterwards. That is if accessing the endpoint of the background request isn't feasible for some reason.
- hobofan 8y agoEven without a documented API, there is often one at play using websites. Example: Website is available at "http://example.com/books/1234" http://example.com/books/1234". When loading it, you see that it fires a request to "http://example.com/api/booksdata/1234" http://example.com/api/booksdata/1234" to load the data that popluates the page. So now you don't have to use a slow browser that loads everything, but can just use your normal http client (for all the ids you know).