3 ms·
I am really curious how do people actually evaluate scrapers? There are so many options and I am dizzy just trying to read them... Also wondering how does the
by churros_train 2y ago
I am really curious how do people actually evaluate scrapers? There are so many options and I am dizzy just trying to read them...
Also wondering how does the OP think about comparing themselves and standing out in the marketplace of seemingly bazillion options
- marcell 2y agoTry it out and let me know if you like it :). If there's a bug I'll fix it! More specifically, FetchFox is targeting a specific niche of scraping. It focuses on small scale scraping, like dozens or a few hundred pages. This is partly because, as a Chrome extension, it can only scrape what the user's internet connection can support. You can't scrape thousands or millions of pages on a residential connection. But a separate reason is, I think that LLM's open up a new market and use case for scraping. FetchFox lets anyone scrape without coding knowledge. Imagine you're doing a research project, and want data from 100 websites. FetchFox makes that easy, whereas with traditional scraping you would have needed coding knowledge to scrape those sites. As an example, I used FetchFox to research political bias in the media. I was able to get data from hundreds of articles without writing a line of code: https://ortutay.substack.com/p/analyzing-media-bias-with-ai https://ortutay.substack.com/p/analyzing-media-bias-with-ai . I think this tool could be used by many non-technical people in the same way.
- churros_train 2y agoAh thats really interesting! How do you evaluate large scale cloud scraping services, since its operations are entirely hidden from you? Personally I am looking into options in this area, are you planning to offer a cloud based version of this at some point/could you tell which existing ones are good if not?
- marcell 2y agoIf you don’t really care about the “morals” of the proxy services, just sign up for a few and see which one is reliable and has good cost. I used Luminati before, they have rebranded to Bright Data. Another is Oxy Labs. I do want to offer a cloud version. If it’s something you’d be interested in, please email and maybe you’d be a good early user for it. You’d get 1:1 attention as one of our early users. Email is on the site and in my profile.
- Malidir 2y agowhy would these non technical people even use a tool when they can go to a internet connected llm and say 'go to this site and get this info' e.g. Mr John Smith is a journalist, find his ten most recent articles via locating his personal website, news sites and social media. so wondering if your tool will be obsolete in a years time?
- marcell 2y agoI can’t predict the future but right now you can use ChatGPT to scrape maybe one or two pages at a time, but it’s harder to do a dozen or a few hundred. So that’s the niche I’m going for. Try it out and see if you like it. Curious if you think it’s better or worse than ChatGPT for scraping dozens of pages.
- deleted 2y ago[deleted]
- mhuffman 2y ago>I am really curious how do people actually evaluate scrapers? It is not easy to evaluate scrapers unless you have had to deal with lots of poorly written websites in your life. Just using it on a few highly structured well-maintained sites can be impressive but if you are using it to acquire data from many websites or large websites things get hairy fast. Most scrapers today are some combination of extracting xpaths and reducing them to the loosest common form, parsing semantic (or easy to identify, like links) or highly structured content that has discoverable patterns, and LLMs. The actual best way to scrape a site is to determine if they are populating the data you want with API calls and replicate those. People are usually more reluctant to completely change back-end code but will make subtle breaking changes (to your scraper) to front-end code all the time. For example small structural or naming changes. This has become more problematic since people have been moving to more SSR and semi-SSR injection. There can also be a problem with discovering all the pages on a site if it doesn't have poorly designed or implemented paging or search. Some of the worst sites to scrape are large WP sites that have obviously been through a few developers. If you really want to test a scraper find some of those and they will put it to the test. Cloudflare is another issue. Not necessarily an issue with this plugin, but because so many sites use it, you typically have to spin up multiple automated headless browsers using residential proxies for any type of large-scale scraping. Some things that LLM does shine at related to scraping is interpreting freeform addresses, custom tables (meaning no TR,TD, but just divs and CSS to make it look like a table, and lists that are also just styled divs. Often there are no tags, attributes, keywords, or generalized xpath that will help you depending on how the developer put it together. Surprisingly there is a pretty old library from Microsoft of all places called Prose if you can still find it (they keep updating but using the same name for different things and trying to inject AI) that is really good at pattern matching and prediction and is small, fast, and free and generally great at building a generalized scraper. Only drawback is I believe the only one I could find was .NET at the time.