5 ms·
Turn entire websites into LLM-ready data
- jshchnz 2y agoThis is cool, seems like a nice way to easily add context for ChatGPT at the least
- altdataseller 2y agoWhat user agent does it use and do you obey robots.txt? * apparently not as I tested it with LinkedIn.com which blocks most non-google/bing bots and it crawls it OK
- super256 2y ago> FireCrawl is built to navigate common web scraping challenges, including reverse proxies, rate limits, and caching They probably ignore robots.txt
- thawab 2y ago[dead]
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- nyolfen 2y agohttps://gist.github.com/yoki/fd2a3642b059703529d158de8233bbc9 https://gist.github.com/yoki/fd2a3642b059703529d158de8233bbc...
- cpeffer 2y agoIt crawls webpages (finds subdirectories), handles JS blocking with fallbacks to headless browsers, and does this all concurrently. If only that script worked for every website. But, alas, it does not.
- deleted 2y ago[deleted]
- pryelluw 2y agoDoes this respect any anti AI related scraping rules set forth by website owners?
- duskwuff 2y agoIt doesn't even respect general robots.txt exclusions, so... I doubt it.
- padolsey 2y agoDo you have pay-as-you-go pricing? I feel that’s always missing from things like this. Cool otherwise.