Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jakubbalada
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
jakubbalada
7y ago
> A growing catalog of up-to-date scrapers for popular websites would put of lot of freelancers out of work. I would invest in this. Check out Apify store ( https://apify.com/store ). It's built exactly for that pur
2.
▲
Web scraping in 2018 – forget HTML, use XHRs, metadata or JavaScript variables
(medium.com)
4 points
by
jakubbalada
9y ago
|
1 comments
3.
▲
by
jakubbalada
9y ago
You can use services like Anti-captcha [1] We have a public API on Apify for that [2] [1] https://anti-captcha.com/mainpage [2] https://www.apify.com/petr_cermak/anti-captcha-recaptcha
4.
▲
by
jakubbalada
9y ago
Typically depends on what you do with scraped data. There is an interesting recent US court ruling [1] [1] http://www.reuters.com/article/us-microsoft-linkedin-ruling/...
5.
▲
by
jakubbalada
9y ago
(co-founder here) We have hundreds of customers, half of them on recurring subscriptions, half just one-time customers paying for crawler configurations.
6.
▲
by
jakubbalada
9y ago
You might try Apifier for that, we've recently scraped more than 150k reviews for 27k restaurants in London. Here's a community crawler you can use: https://www.apifier.com/community/crawlers/Yonny/b
7.
▲
by
jakubbalada
10y ago
It's also hard to get direct access to the data. But you're right it's a hard sell to enterprises although we have some (e.g. real estate developer creating pricing maps)
8.
▲
by
jakubbalada
10y ago
Both - developers on a free plan using own RSS for sites without one and business people (mainly startups) building their products on top of Apifier. Typical use is an aggregator that needs common API for all partners who are not able to pr
9.
▲
by
jakubbalada
10y ago
We see a lot of users who needs data from the web or APIs for sites which doesn't have one. Just not all of them can code and we have to scale custom development.
10.
▲
by
jakubbalada
10y ago
Disclaimer: I'm a co-founder of Apifier [1]. It's not an open source, but free up to 10k pages per month. And it can handle modern JS web applications (your code runs in a context of crawled page). You can for example scrape API k
11.
▲
by
jakubbalada
10y ago
If you want to scrape many websites on a daily basis, have a look at https://www.apifier.com as an alternative. Disclaimer: I'm a cofounder there
12.
▲
by
jakubbalada
11y ago
At least in our batch we felt we should really only focus on our product. That's why we've got the grant - to be able to work on our startup full-time.
13.
▲
by
jakubbalada
11y ago
(Apifier co-founder here) It will be less beneficial now than it was in the first YCF batch. Other fellows really motivated us during the group office hours and other events. And still - main thing is to focus on a product which you can do
14.
▲
by
jakubbalada
11y ago
You're right, price per request would be easier for estimation. But as you can use JavaScript, you can scrape whole website with just one page request (see the SFO flights example). In other words, our costs doesn't correlate with
15.
▲
by
jakubbalada
11y ago
Of course not, API will be available soon - in a week or two. If you have some other feature requests, please let us know, we need to help with prioritization.
16.
▲
by
jakubbalada
11y ago
Yes, by default we respect robots.txt. There is a switch to disable it - on your own responsibility. We don't fully respect Crawl-delay, but minimum delay between requests for all our crawlers is set to 2000ms. We don't publish ou