Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mateuszbuda
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
by
mateuszbuda
3y ago
Our front is software but on the backend we’ve built a mobile proxy pool and continue to work with hardware to improve it. More details on our blog: https://scrapingfish.com/blog/byo-mobile-proxy-for-web-scrap...
32.
▲
State of Web Scraping 2023
(scrapingfish.com)
4 points
by
mateuszbuda
3y ago
|
0 comments
33.
▲
by
mateuszbuda
3y ago
That’s why people use mobile proxies which rotate IPs to scrape Instagram, Facebook, TikTok, etc: https://scrapingfish.com/blog/scraping-instagram
34.
▲
by
mateuszbuda
3y ago
Thanks! Internally we use more sophisticated system required by larger scale to support higher volume and concurrent connections from multiple clients, but for smaller scale limited to one person or a small team, everything described on our
35.
▲
by
mateuszbuda
3y ago
Yes, but we have so many plans and actually many of them are unlimited. We also track data usage and can either add extra data package or a SIM card gets excluded from our proxy pool.
36.
▲
by
mateuszbuda
3y ago
Actually, it’s a feature :)
37.
▲
by
mateuszbuda
3y ago
Good question. Here we explain, with a few simplifications, how we source our IPs: https://scrapingfish.com/blog/byo-mobile-proxy-for-web-scrap...
38.
▲
by
mateuszbuda
3y ago
For https://scrapingfish.com/ it took us about 5-6 months from idea to $2k/month. It’s still not our main source of income. It was fun to build especially that it involved working with hardware.
39.
▲
by
mateuszbuda
4y ago
In this particular case, GPT can help you mostly with parsing the website but not with the most challenging part of web scraping which is not getting blocked. In this case, you still need a proxy. The value from using web scraping APIs is a
40.
▲
by
mateuszbuda
4y ago
This may be difficult as for half of the products scraped from Walmart, sugar is the main ingredient: https://scrapingfish.com/blog/scraping-walmart
41.
▲
by
mateuszbuda
4y ago
This is definitely possible. I’m not sure if we’re going to have time for this as we’re occupied by work on Scraping Fish but we shared the code for scraping nutrition facts data from Walmart on github: https://github.com/pa
42.
▲
by
mateuszbuda
4y ago
Here is my post from last year which didn’t get much attention on HN: https://news.ycombinator.com/item?id=33507260 It analyzes how much sugar is in the food based on nutrition facts data scraped from Walmart. It also shows
43.
▲
by
mateuszbuda
4y ago
It means that our IPs are ethically sourced. You can read more here: https://scrapingfish.com/how-ips-for-web-scraping-are-source...
44.
▲
by
mateuszbuda
4y ago
Scraping Fish - a web scraping API powered by custom-build, ethical, mobile proxy pool: https://scrapingfish.com/
45.
▲
by
mateuszbuda
4y ago
Here is my story, which may provide some inspiration. My co-founder was searching for an apartment to purchase and found that all the tools he was using had weaknesses and none of them met his needs. We decided to create an aggregator for r
46.
▲
by
mateuszbuda
4y ago
What is the best setup for Safari on mac to block ads, block YouTube ads, block trackers, block cookies popups and automatically deny cookies? I like Orion but it is simply broken on many websites, especially some forms or buttons do not wo
47.
▲
by
mateuszbuda
4y ago
Great project! With your DIY attitude, if you need to build your own infrastructure for web scraping, here is a tutorial for mobile proxy setup which might be helpful: https://scrapingfish.com/blog/byo-mobile-proxy-for-
48.
▲
by
mateuszbuda
4y ago
I don't want to spread negativity here but such success stories from indie hackers or side projects always remind me that most of them don't make any money: https://scrapingfish.com/blog/indie-hackers-revenue
49.
▲
by
mateuszbuda
4y ago
You’re right that your not alone in this. More than 50% of products built by indie hackers don’t make any money at all: https://scrapingfish.com/blog/indie-hackers-revenue
50.
▲
by
mateuszbuda
4y ago
Interesting point. As far as I know, it's not that easy. We have a client who contacted the owners of the website he needs to scrape and offered to pay them for access to their data but they were not interested. I know that they are no
51.
▲
by
mateuszbuda
4y ago
On our blog, we give away how to create a home-scale version of our 4G proxy: https://scrapingfish.com/blog/byo-mobile-proxy-for-web-scrap... We have some domains blacklisted and periodically monitor logs for websites
52.
▲
by
mateuszbuda
4y ago
A web scraping API: https://scrapingfish.com We started with a simple web scraping solution for real estate market that was up and running in just a couple of days. We used it to track prices of apartments in our area aggregated
53.
▲
by
mateuszbuda
4y ago
On a related topic, here is an article which presents interesting findings based on nutrition data scraped from Walmart food products: https://scrapingfish.com/blog/scraping-walmart It turns out that sugar is the main
54.
▲
Scraping Google SERP with Geolocation
(scrapingfish.com)
2 points
by
mateuszbuda
4y ago
|
0 comments
55.
▲
Scraping Walmart to Estimate Share of Sugar in Food
(scrapingfish.com)
5 points
by
mateuszbuda
4y ago
|
1 comments
56.
▲
by
mateuszbuda
4y ago
Web scraping is considered shady business by many people mostly because of how IP address are obtained by proxy providers. More details on this here: https://scrapingfish.com/how-ips-for-web-scraping-are-source...
57.
▲
by
mateuszbuda
4y ago
Scraping Fish uses puppeteer/playwright (headless browsers) connected to our mobile proxy pool under the hood.
58.
▲
by
mateuszbuda
4y ago
The distribution of IP addresses depends to a large extent on 4G provider. Some of them reuse IPs and you can get the same IP after you change network mode 4G > 3G > 4G. The same applies to IP addresses allocation space. You can reque
59.
▲
Build Your Own Mobile Proxy for Web Scraping
(scrapingfish.com)
185 points
by
mateuszbuda
4y ago
|
31 comments
60.
▲
by
mateuszbuda
4y ago
I have a script for concurrent web scraping: https://github.com/mateuszbuda/webscraping-benchmark It takes a file with urls and scrapes the content. For more demanding websites it can use web scraping API that handles
More ›