6 ms·
Show HN: I'm making an AI scraper called FetchFox
Hi! I'm Marcell, and I'm working on FetchFox (https://fetchfoxai.com https://fetchfoxai.com). It's a Chrome extension that lets you use AI to scrape any website for any data. I'd love to get your feedback.
Here's a quick demo showing how you can use it to scrape leads from an auto dealer directory. What's cool is that it scrapes non-uniform pages, which is quite hard to do with "traditional" scrapers: https://youtu.be/wPbyPSFsqzA https://youtu.be/wPbyPSFsqzA
A little background: I've written lots and lots of scrapers over the last 10+ years. They're fun to write when they work, but the internet has changed in ways that make them harder to write. One change has been the increasing complexity of web pages due to SPAs and obfuscated CSS/HTML.
I started experimenting with using ChatGPT to parse pages, and it's surprisingly effective. It can take the raw text and/or HTML of a page, and answer most scraping requests. And in addition to traditional scraping thigns like pulling out prices, it can extract subjective data, like summarizing the tone of an article.
As an example, I used FetchFox to scrape Hacker News comment threads. I asked it for the number of comments, and also for a summary of the topic and tone of the articles. Here are the results: https://fetchfoxai.com/s/cSXpBs3qBG https://fetchfoxai.com/s/cSXpBs3qBG . You can see the prompt I used for this scrape here: https://imgur.com/uBQRIYv https://imgur.com/uBQRIYv
Right now, the tool does a "two step" scrape. It starts with an initial page, (like LinkedIn) and looks for specific types of links on that page, (like links to software engineer profiles). It does this using an LLM, which receives a list of links from the page, and looks for the relevant ones.
Then, it queues up each link for an individual scrape. It directs Chrome to visit the pages, get the text/HTML, and then analyze it using an LLM.
There are options for how fast/slow to do the scrape. Some sites (like HN) are friendly, and you can scrape them very fast. For example here's me scraping Amazon with 50 tabs: https://x.com/ortutay/status/1824344168350822434 https://x.com/ortutay/status/1824344168350822434 . Other sites (like LinkedIn) have strong anti-scraping measures, so it's better to use the "1 foreground tab" option. This is slower, but it gives better results on those sites.
The extension is 100% free forever if you use your OpenAI API key. It's also free "for now" with our backend server, but if that gets overloaded or too expensive we'll have to introduce a paid plan.
Last thing, you can check out the code at https://github.com/fetchfox/fetchfox https://github.com/fetchfox/fetchfox . Contributions welcome :)
- smcin 2y agoInteresting. How long did it take to figure out how to do this with ChatGPT? > "By scraping raw text with AI, FetchFox lets you circumvent anti-scraping measures on sites like LinkedIn and Facebook. Even the the complicated HTML structures are possible to parse with FetchFox."
- marcell 2y agoThe ChatGPT part is pretty easy actually. You can just dump text and HTML and ask it a question, and it usually answers correctly. The trickier part is “everything else” to make the extension work.
- smcin 2y agoEven the parsing of obfuscated HTML + CSS + dynamic JSON content?
- marcell 2y agoSurprisingly yes, most of the time. I’ve put in a few optimizations: 1. Remove all <style> and <svg > tags. These rarely add value, and can dramatically increase token counts. 2. For the “crawl” step, I exclusively pull out <a> tags and only look at those. The “extract” step looks at full HTML 3. For now, it only looks at the first 50k text characters, and the first 120k HTML characters. This is to stay within token limits. The last part will be what I focus on improving in the next version.
- dillondoyle 2y agoCould go the google way, capture an image screenshot of state, ocr, then parse it. They keep throwing it in my url bar. I refuse to click (big warning it sends to google's servers)
- FrenchDevRemote 2y agohow do you deal with the fact that some basic pages can have tens of thousand of tokens?