4 ms·
Archive.org works surprisingly well as a general purpose web proxy. Just prefix the URL, e.g., http://example.com http://example.com, with https://web.archive.
by x3blah 6y ago
Archive.org works surprisingly well as a general purpose web proxy. Just prefix the URL, e.g., http://example.com http://example.com, with https://web.archive.org/save/ https://web.archive.org/save/, e.g., https://web.archive.org/save/http://example.com https://web.archive.org/save/http://example.com
The aesthetic intrusiveness of the archive.org header and footer are minimal since I use a text-only browser that has no Javascript engine.
Sometimes I get "This url is not available on the live web or can not be archived." However this happens for only a surprisingly small minority of websites.
Rarely I find that /save is unsuccessful in which case I can still find past copies using something like
curl -o 1.txt "https://web.archive.org/cdx/search/cdx?url=http://www.example.net&fl=timestamp,original" ;
sed -i '/^[12][0-9]* h/!d;/^[12][0-9]* h/{s/^/http:\/\/web.archive.org\/web\//;s/ /\//;s/\r//;}' 1.txt
The limitation with past copies versus /save is that archive.org will not usually crawl past page one on websites with many successive pages, e.g., http://example.com/?page=2 http://example.com/?page=2, http://example.com/?page=3 http://example.com/?page=3, etc.
Has anyone ever considered mirroring archive.org, or parts of it, to other geographic locations.
Could this be done. Why or why not.
- Springtime 6y agoI use a custom browser keyword search to find existing archived pages before saving one, personally. Eg: ar <URL> for: https://wayback.archive.org/web/*/%S I'd imagine it would be useful for IA to implement some message for scenarios where a page has already been saved within a certain timespan and provide both a link to the already saved version and offer to save again. As this would mitigate mass savings of an identical page that can occur when some popular link is accidentally shared with the /save/ URL instead of the static URL or when it's a popular page that people want to archive. Archive.is displays such a message (to the effect of, 'this page was archived <date>, if it looks outdated click save') and also redirects to the most recent copy.
- x3blah 6y agoThat's a good point. I mainly use it for browsing websites that change daily as well as ones with many successive pages, e.g., 1, 2, 3, etc. that IA does automatically not crawl. Thus, not many existing copies if any. HAproxy changes the Host header and modifies the URL. I can either use the text-only browser's http-proxy option or I can direct the request to the web.archive.org backend by adding a custom HTTP header to the request. If I am not mistaken, the so-called "modern" browsers do not have built-in capability to add headers.