4 ms·
I don't know about archive.is, but 12ft.io does identify as google to bypass paywalls afaik
by Yujf 3y ago
I don't know about archive.is, but 12ft.io does identify as google to bypass paywalls afaik
- strunz 3y ago12ft.io also doesn't work or is disabled for many sites that archive.is still works on
- hda111 3y agoMaybe because the creator of 12ft.io isn't anonymous
- janejeon 3y agoWouldn't sites be able to see that requests from 12ft.io isn't coming from Google's IPs?
- dpifke 3y agoYes. Google recommends using reverse DNS to verify whether a visitor claiming to be Googlebot is legitimate or not: https://developers.google.com/search/docs/crawling-indexing/verifying-googlebot https://developers.google.com/search/docs/crawling-indexing/... You can also verify IP ownership using WHOIS, or by examining BGP routing tables to see which ASN is announcing the IP range. Google also publishes their IP address ranges here: https://www.gstatic.com/ipranges/goog.json https://www.gstatic.com/ipranges/goog.json
- nora-puchreiner 3y agohttps://search.google.com/test/rich-results?url= https://search.google.com/test/rich-results?url= operates from legit Googlebot IPs so it allows anyone to get the paywalled content even archive.is fails to fetch (from theinformation.com, for example)
- rahimnathwani 3y ago"Google recommends using reverse DNS to verify..." This is almost right. They recommend two steps: 1. Use reverse DNS to find the hostname the IP address claims to have. (The IP address block owner can put any hostname in here, even if they don't own/control the domain.) 2. Assuming the claimed hostname is on one of Google's domains, do a forward DNS lookup to verify that the original IP address is returned. The second step is the important one.