Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jardah
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
jardah
8y ago
Good point. Haven't seen a single detection library do this, but at least now I know, that I still need to work on alternative solution. Thanks
2.
▲
by
jardah
8y ago
The webdriver property is as far as we know the only one that stays different if you use non-headless chrome with puppeteer. Rest can be handled by use of non-headless chrome as mentioned in the article. But you are right, after reading thr
3.
▲
by
jardah
8y ago
That is kinda sad to hear. The approach should always be to go through the path of least resistance and smallest effect on the website. So for example, if a company has API that can be used instead of scraping their website, then it's
4.
▲
by
jardah
8y ago
Yea, it was a very general example, since there is at least one rule that is based on rate limiting too, and this 300/IP limit is what have seen on average.
5.
▲
by
jardah
8y ago
Yes it is, and since it's IP based, it's even easier if you are for example working from an office and there are multiple people using google. But that is why they only show recaptcha, you fill it in and you will get extemption co
6.
▲
Bypassing website anti-scraping protections
(kb.apify.com)
235 points
by
jardah
8y ago
|
118 comments
7.
▲
by
jardah
9y ago
Amazon is unfortunately not using any metadata information for reviews (probably to prevent easy scraping for competing companies). You can only get it from from html (At least from what I can see).
8.
▲
by
jardah
9y ago
Depends on whether we access the website from a proxy that is known by the WAF. But for most websites it's just a single normal request. If it's an issue in the future we could make browser extension, that will do the analytic on
9.
▲
by
jardah
9y ago
Just a quick update: Thank you for using it and playing around with it. Looking at the usage and results I found a quite a lot of things to improve. Which is great, since it's hard to develop something like this without real usage data
10.
▲
by
jardah
9y ago
Aha! I see, it shows data based on POST request from FORM on this page http://www.dsden93.ac-creteil.fr/spip/spip.php?page=annu1d so if you provide just a link to the results page without the POST data then it will sho
11.
▲
by
jardah
9y ago
Yes, that is probably the problem, when I looked for the text it returned: [ 0:{ "selector":".bloc-blanc > p:nth-child(1)" "text":" 0 école(s) correspondent à votre recherche " } ]
12.
▲
by
jardah
9y ago
When I open the link in my browser it shows "0 école(s) correspondent à votre recherche" and no table, probably what happens to the analyzer too.
13.
▲
by
jardah
9y ago
:D yep the documentation needs a lot of work. It started as a test of an idea, then slowly became a usable tool and the code was getting incrementaly more complex without me event noticing. I only added the readme on github yesterday and th
14.
▲
by
jardah
9y ago
Actually the list of assets shouldn't be that hard. Looking at pinterest the xhr requests for images are loaded immediately when page is open, so potentialy it then it's catched in onRequest function (only now I'm aborting th
15.
▲
by
jardah
9y ago
This tool basicaly performs the simplest data loading, it opens the webpage, then waits till most xhr requests are done, wait's a second (tio give JS time to manipulate DOM) and then loads data from the page. This way, it has what user
16.
▲
by
jardah
9y ago
Some general authentication (like separate input fields for your login credentials on the website) could be potentialy done (but very unsafe for user of the tool, since you would be sending us your credentials as plaintext). But authenticat
17.
▲
by
jardah
9y ago
It's why I'm using proxies, every request is routed through different proxy address and the application as whole is rate limited. So hopefully I'm not making too much traffic on yelp. They are just a perfect example because t
18.
▲
by
jardah
9y ago
Yep only microdata. I completely forgot about RDFa. I'm immediately writing RDFa to my todo list. It would be a great addition.
19.
▲
by
jardah
9y ago
Oh... Clearly I need to work on my UX skills, I will improve that in next iteration.
20.
▲
by
jardah
9y ago
Thank you, still lot's of things to improve (for example 404 handling) but it's great to see positive feedback.
21.
▲
by
jardah
9y ago
Nope, there is no caching now, every run of the tool has a single instance and writes the output into separate file. I'm using it to test stability of cloud when multiple users are using it and to test proxies. It would not be much of
22.
▲
by
jardah
9y ago
Sadly not for now. Our company has a solution for that (for some websites), but currently this tool does not have this functionality, since I wanted it be as simple as possible. Maybe in the future.
23.
▲
by
jardah
9y ago
Good idea and I would implement that if I used an API from server to get the response. But currently I'm at the same time testing stability of Apify "Actor" solution and proxies, so for my case it's good that there are r
24.
▲
by
jardah
9y ago
Jaroslav, yes, I'm the author. Did you notice any problems or ways how I can improve it?
25.
▲
by
jardah
9y ago
I'm still testing it and improving it (there are so many different websites with different responses...), so If you have any comments I'm looking forward to what you think about it.
26.
▲
Show HN: Web scraping page analyzer
(apify.com)
166 points
by
jardah
9y ago
|
50 comments