4 ms·
Does anyone know how they do this scraping from a technical standpoint. The articles allude to it being the same as data Google/Bing spiders, which can clearly
by opaque 9y ago
Does anyone know how they do this scraping from a technical standpoint. The articles allude to it being the same as data Google/Bing spiders, which can clearly access more data that average internet IP for making their result summaries. I had assumed big sites whitelisted specific crawler IP ranges or User-Agents for the search giants. Do they somehow spoof this?
- revelation 9y agoI don't think they do any such thing, if anything they are rotating IPs/user agents to avoid being limited or blocked. Google requires sites to send the crawler the same content as someone clicking a link on a Google results page would see, so even if some sites get creative covering it up with blurred boxes and similar dark patterns, the data is there in the markup.
- opaque 9y agoI haven't checked the markup, but if you try and hit a linkedin profile page you just get forwarded to a login page. Perhaps if you don't follow the forwarding? Not sure how this complies with google's requirements, I suspect if you're big enough you get a custom arrangement. However, that doesn't explain how hiQ are getting the data.