5 ms·
It seems a lot of people havent heard of it, but I think its worth plugging https://perma.cc/ https://perma.cc/ which is really the appropriate tool for somethi
by basch 8mo ago
It seems a lot of people havent heard of it, but I think its worth plugging https://perma.cc/ https://perma.cc/ which is really the appropriate tool for something like Wikipedia to be using to archive pages.
mroe https://en.wikipedia.org/wiki/Perma.cc https://en.wikipedia.org/wiki/Perma.cc
- jsheard 8mo agoDoes Wikipedia really need to outsource this? They already do basically everything else in-house, even running their own CDN on bare metal, I'm sure they could spin up an archiver which could be implicitly trusted. Bypassing paywalls would be playing with fire though.
- toomuchtodo 8mo agoArchive.org is the archiver, rotted links are replaced by Archive.org links with a bot. https://meta.wikimedia.org/wiki/InternetArchiveBot https://meta.wikimedia.org/wiki/InternetArchiveBot https://github.com/internetarchive/internetarchivebot https://github.com/internetarchive/internetarchivebot
- jsheard 8mo agoYeah for historical links it makes sense to fall back on IAs existing archives, but going forward Wikipedia could take their own snapshots of cited pages and substitute them in if/when the original rots. It would be more reliable than hoping IA grabbed it.
- toomuchtodo 8mo agoNot opposed, Wikimedia tech folks are very accessible in my experience, ask them to make a GET or POST to https://web.archive.org/save https://web.archive.org/save whenever a link is added via the Wiki editing mechanism. Easy peasy. Example CLI tools are https://github.com/palewire/savepagenow https://github.com/palewire/savepagenow and https://github.com/akamhy/waybackpy https://github.com/akamhy/waybackpy Shortcut is to consume the Wikimedia changelog firehose and make these http requests yourself, performing a CDX lookup request to see if a recent snapshot was already taken before issuing a capture request (to be polite to the capture worker queue).
- jsheard 8mo agoI didn't know you can just ask IA to grab a page before their crawler gets to it. In that case yeah it would make sense for Wikipedia to ping them automatically.
- extraduder_ire 8mo agoThere's a /save/<url> endpoint that archives the page you point it at. You can see a text box for it on the right, if you go on the waybackmachine's homepage. I used it yesterday.
- RupertSalt 8mo agoSpammers and pirates just got super excited at that plan!
- toomuchtodo 8mo agoThere are various systems in place to defend against them, I recommend against this, poor form against a public good is not welcome.
- ferngodfather 8mo agoWhy wouldn't Wikipedia just capture and host this themselves? Surely it makes more sense to DIY than to rely on a third party.
- deleted 8mo ago[deleted]
- huslage 8mo agoWhy would they need to own the archive at all? The archive.org infrastructure is built to do this work already. It's outside of WMF's remit to internally archive all of the data it has links to.
- Gander5739 8mo agoThis already happens. Every link added to Wikipedia is automatically archived on the wayback machine.
- snigsnog 8mo agoArchive.org are left wing activists that will agree to censor anything other left wing activists or large companies don't want online.
- Maken 8mo agoLike what?
- snigsnog 8mo agoKiwifarms is an example: https://old.reddit.com/r/DataHoarder/comments/x95gd5/internet_archive_breaks_from_previous_policies_on/ https://old.reddit.com/r/DataHoarder/comments/x95gd5/interne... Anyone can request anything be removed and they may honor the request: https://help.archive.org/help/how-do-i-request-to-remove-something-from-archive-org/ https://help.archive.org/help/how-do-i-request-to-remove-som... they say nothing about only removing things illegal in the US or anything like that, meaning they can and will remove things based on personal judgements about whether it should be archived.
- Maken 8mo agoUnlisting (not even removing) doxing information about living people is being a left wing activist? Is there where the Oberton window lies now?
- AlexeyBelov 8mo agoAnd you're another disruptive "N days old" account. Troll somewhere else.
- deleted 8mo ago[deleted]
- raincole 8mo ago> Does Wikipedia really need to outsource this? I hope so. Archiving is a legal landmine.
- IshKebab 8mo agoOf course they do. If Wikipedia did it themselves they'd immediately get DMCA'd and sued into oblivion. > Bypassing paywalls would be playing with fire though. That's the only reason archive.today was used. For non-paywalled stuff you can use the wayback machine.
- deleted 8mo ago[deleted]
- ronsor 8mo agoIt costs money beyond 10 links, which means either a paid subscription or institutional affiliation. This is problematic for an encyclopedia anyone can edit, like Wikipedia.
- toomuchtodo 8mo agoWikimedia could pay, they have an endowment of ~$144M [1] (as of June 30, 2024). Perma.cc has Archive.org and Cloudflare as supporting partners, and their mission is aligned with Wikimedia [2]. It is a natural complementary fit in the preservation ecosystem. You have to pay for DOIs too, for comparison [3] (starting at $275/year and $1/identifier [4] [5]). With all of this context shared, the Internet Archive is likely meeting this need without issue, to the best of my knowledge. [1] https://meta.wikimedia.org/wiki/Wikimedia_Endowment https://meta.wikimedia.org/wiki/Wikimedia_Endowment [2] https://perma.cc/about https://perma.cc/about ("Perma.cc was built by Harvard’s Library Innovation Lab and is backed by the power of libraries. We’re both in the forever business: libraries already look after physical and digital materials — now we can do the same for links.") [3] https://community.crossref.org/t/how-to-get-doi-for-our-journal/3163 https://community.crossref.org/t/how-to-get-doi-for-our-jour... [4] https://www.crossref.org/fees/#annual-membership-fees https://www.crossref.org/fees/#annual-membership-fees [5] https://www.crossref.org/fees/#content-registration-fees https://www.crossref.org/fees/#content-registration-fees (no affiliation with any entity in scope for this thread)
- RupertSalt 8mo agoIf the WMF had a dollar for every proposal to spend Endowment-derived funds, their Endowment would double and they could hire one additional grant-writer
- nine_k 8mo agoIf the endowment is invested so that it brings very conservative 3% a year, it means that it brings $4.32M a year. By doubling that, rather many grant writers could be hired.
- ouhamouch 8mo ago[dead]
- culi 8mo agoThe 3 listed alternatives there seem to have nothing to do with digital archiving. Here's a better alternative to g2 that doesn't login-wall you: https://alternativeto.net/software/freezepage/ https://alternativeto.net/software/freezepage/
- Computer0 8mo agoI switched to Perma.cc earlier this week and have had a mixed experience to say the least. I think image heavy pages just error out completely, while still charging me such as: https://www.in.gov/nircc/planning/highway/traffic-data/intersectionarterial-data/ https://www.in.gov/nircc/planning/highway/traffic-data/inter... and reddit blocks their agent seemingly. It is open source though.