4 ms·
They are not actually bypassing firewalls - therefore I think they are on ethically good grounds. Those sites show their full text for web crawlers - only not t
by postexitus 11mo ago
They are not actually bypassing firewalls - therefore I think they are on ethically good grounds. Those sites show their full text for web crawlers - only not to humans. Basically, archive.is and the folk simulate that through various means. Headless browsers, better agent injection etc.
- mr_mitm 11mo agoI don't think that's true. If it was that simple, there would be browser plugins or other apps that would replicate that behavior. Do you know of any?
- aspenmayer 11mo agoPerhaps not in the same way as described above, but BPC exists. https://en.wikipedia.org/wiki/Bypass_Paywalls_Clean https://en.wikipedia.org/wiki/Bypass_Paywalls_Clean
- postexitus 11mo agoI always assumed so - because Google can index them full text. It used to be the case that you could see those full snapshots in Google as cache - this was before the sites strong armed Google to remove those snapshots from being accessible, then archive.* folk rose to power. You can test this yourself for searching for a unique quote on those sites and still getting hits in Google. But you are right - why could this not be achieved with a plugin then - don't know.