6 ms·
this is so sad. i understand the protest, but it feels like burning books. is all that prior information gone? or is the protest temporary way to get the commun
by modzu 3y ago
this is so sad. i understand the protest, but it feels like burning books. is all that prior information gone? or is the protest temporary way to get the community migrated away but content will be restored for historical purposes?
- PaulHoule 3y agomore like building a new library
- elijaht 3y agohow do I access the content that was at /r/startrek?
- samtho 3y agoAfter burning the current one, books and all. Granted this had to do with the fact that unlike any other agreed upon standard protocol, Reddit posts are not accessible except by indirect means, i.e. you can’t download the contents of a community the same way you would with a git repository or an email server and migrate it elsewhere.
- toomuchtodo 3y agoIf you can enumerate all of the threads in a subreddit, each thread is a blob of JSON when you tack `.json` on to the end. Tangentially, https://old.reddit.com/r/DataHoarder/comments/12ucc9z/downloading_an_entire_subreddit_in_2023/ https://old.reddit.com/r/DataHoarder/comments/12ucc9z/downlo... and https://old.reddit.com/r/redditdev/comments/c93rdt/how_do_i_get_json_data_of_a_specific_post_without/ https://old.reddit.com/r/redditdev/comments/c93rdt/how_do_i_... A sibling comment mentions ArchiveTeam, which ends up in the Wayback Machine. Some work to be done around tools to make that corpus more readily available for consumption and perhaps backfill. Lots of existing tooling to query the Internet Archive's CDX servers to understand what coverage looks like and retrieve archived content.
- samtho 3y agoCrawling is indirect. This isn’t a protocol like IMAP where each object exists inherently as specified by the protocol. The idea of a “thread” or “user” does not exist in HTTP, only “documents” do. Everything at this layer is made up by (and at the disposal of) the operator and is not standardized.
- toomuchtodo 3y agoRarely do optimal conditions exist unfortunately.
- modzu 3y agolets take a moment to remember that reddit was cofounded by the inventor of rss.
- malermeister 3y agoThere's an archive available going all the way till March: https://archive.org/details/pushshift-reddit-2023-03 https://archive.org/details/pushshift-reddit-2023-03 They could probably import the data to lemmy.
- jorblumesea 3y agoSocial media is an endless process of migrating to a new place It's writing a new chapter in your internet story.
- jrm4 3y agoNot to be a jerk about it, but this is why you shouldn't have stored the books in a warehouse owned by a company that could lock the doors at any time. This is the opposite of why the 'net was invented.
- EA-3167 3y agoWhat was the other option, realistically? There are barely other options today, never mind 15 years ago.
- fsflover 3y agoI guess it's fine that 15 years ago the users chose Reddit due to the lack of good alternatives. But today people should be aware of the problem and go to decentralized platforms.
- EA-3167 3y agoOnce you've built a community and a "library" though, as we see, it's hard to up and move it. At least, until the guy who owns the library threatens to burn it all down.
- jrm4 3y agoThe whole damn point of the internet is that the "library" never ever ever actually MUST be "owned."
- krapp 3y agoSince when? Every website on every server going back to the very first one at CERN has been owned by someone.
- jrm4 3y agoSeems like it should be relatively clear that when I say "owned" here, I mean "strongly subject to the whims of a possibly capricious owner." Like, sure, university libraries are "owned" by the school, but one can generally reliably treat them as "public."
- BasedAnon 3y agobooks rot over time, and when there are billions of books being produced daily it's not reasonable to expect all of it to stick around
- waboremo 3y agoThey can transfer the information if they want, API changes don't go through for another couple of weeks and archives do exist.
- camel-cdr 3y agoAnd there is also the pushshift torrent backup.
- neurostimulant 3y agoIf a link you click on a Google search result is dead, chance that the Internet Archive has it. Just put the dead link in the wayback machine and find out.