4 ms·
A web scraping API: https://scrapingfish.com https://scrapingfish.com We started with a simple web scraping solution for real estate market that was up and run
by mateuszbuda 4y ago
A web scraping API: https://scrapingfish.com https://scrapingfish.com
We started with a simple web scraping solution for real estate market that was up and running in just a couple of days. We used it to track prices of apartments in our area aggregated across multiple websites.
Then, as we saw value in this, we expanded data scraping to other cities and types of properties and released the product to external users. We had a few paying customers after a couple of months.
As we wanted to include more websites to collect data from, we run into significant problems of being blocked. In result, we started investigating how to overcome different mechanisms that websites use to prevent automated traffic from web scrapers.
It turns out that one of the the most important factors is to use good quality proxy which provides IP addresses shared with other real users and change them frequently. So, we started building our own proxy infrastructure powered by 4G proxies and implemented an API on top of it. And this is how we created Scraping Fish API for web scraping.
Now, we can offer a reliable solution for scraping even the most demanding websites like Instagram or Facebook.
Here is the full story of our product on IndieHackers: https://www.indiehackers.com/product/scraping-fish https://www.indiehackers.com/product/scraping-fish
- emadda 4y agoThis looks cool. When you say "4G proxies", are these devices that route their traffic over 4G connections via the mobile network? So you are using the IP ranges of telecommunications companies rather than clouds? Also, how do you protect against customers doing illegal things with the IP's that are assigned to your business?
- mateuszbuda 4y agoOn our blog, we give away how to create a home-scale version of our 4G proxy: https://scrapingfish.com/blog/byo-mobile-proxy-for-web-scraping https://scrapingfish.com/blog/byo-mobile-proxy-for-web-scrap... We have some domains blacklisted and periodically monitor logs for websites which our customers scrape but so far not a single case of illegal scraping. We've also decided to go with a small $2 purchase instead of free trial account with no credit card required. If someone contacts us with their use case and request for a free test account, we're happy to create one for them but it has to go through us. This way we're not a good choice for people willing to do illegal stuff with our API.
- is_true 4y agoHave you ever thought about sharing revenue with the owner of the sites? I asked on the guys that runs scrapping bee but he just ignored me.
- mateuszbuda 4y agoInteresting point. As far as I know, it's not that easy. We have a client who contacted the owners of the website he needs to scrape and offered to pay them for access to their data but they were not interested. I know that they are not their competitor and they're not doing anything that would harm their revenue in any way. I assume that they figured it was not worth the legal, organisational and other formal hassle to share their data. Let's say that we want to pursue this idea anyway. We would have to reach our to every website owner that our clients scrape and figure out how to share the revenue. You cannot just transfer your money to another company. You need to sign a contract and in many cases have it approved by legal and compliance teams.