Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mnmkng
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
mnmkng
4y ago
Hey! Crawlee uses the libraries from our fingerprint suite internally. https://github.com/apify/fingerprint-suite#performance It has an A rating in the BotD (fingerprint.js) detection. Now we're working on improvi
32.
▲
by
mnmkng
4y ago
Thanks, that's our experience exactly and that's why we built the library this way. It's not uncommon to switch from HTTP to headless back to HTTP in the lifecycle of a project as the website evolves or as you find better way
33.
▲
Are website terms of use enforced?
(blog.apify.com)
2 points
by
mnmkng
4y ago
|
0 comments
34.
▲
Apify starts a web scraping academy
(apify.com)
1 points
by
mnmkng
5y ago
|
0 comments
35.
▲
by
mnmkng
5y ago
While I agree that Scrapy is a great tool for beginner tutorials and easy entry into scraping, it's becoming difficult to use it in real world scenarios because almost all the large players now employ some anti-bot or anti-scraping pro
36.
▲
by
mnmkng
5y ago
The author here. I actually am a lawyer. I just wanted to make it clear that I'm not "your" lawyer and therefore can't give advice on your specific use-case. Thanks for feedback. Might want to clarify that.
37.
▲
by
mnmkng
5y ago
Sadly not a lot of people share this opinion.
38.
▲
Is Web Scraping Legal?
(blog.apify.com)
33 points
by
mnmkng
5y ago
|
16 comments
39.
▲
Show HN: Web scraping focused HTTP client for Node.js
(github.com)
3 points
by
mnmkng
5y ago
|
1 comments
40.
▲
by
mnmkng
5y ago
Hey everyone, we built a special-purpose web scraping client for Node.js. When scraping with pure HTTP clients, you want to blend in with the regular traffic as much as you can. This means your request signature needs to look like a browser
41.
▲
by
mnmkng
6y ago
I think of web scraping as nothing less than automation of human work. There really is no hacking or unauthorized access involved. It's either me, using a browser, to read something publicly accessible on the internet, or it's my
42.
▲
by
mnmkng
6y ago
Not necessarily. It is true that most websites today are JavaScript heavy. However, they are server-side rendered more often than not. Mostly for performance reasons. Also, not all search engines are as good as Google at indexing dynamic JS
43.
▲
by
mnmkng
6y ago
I have not tried the Headless Chrome Crawler personally, but try the Apify SDK out https://github.com/apify/apify-js if the Headless Chrome crawler does not scale well enough. We use it to scrape billions of pages ever
44.
▲
by
mnmkng
6y ago
Actually, if you're scraping at any scale above a hobby project, most of your web scraping hours would now be spent on avoiding bot detection, reverse engineering APIs and trying to make HTTP requests work where it seems only a browser
45.
▲
by
mnmkng
6y ago
Could you explain a bit more about how you run the React DevTools in a headless Chrome? As far as I know, headless Chrome can't run extensions.
46.
▲
by
mnmkng
6y ago
Hey everyone, maintainer of the Apify SDK here. As far as we know, it is the most comprehensive open-source scraping library for JavaScript (Node.js). It gives you tools to work with both HTTP requests and headless browsers, storages to sav