5 ms·
ArchiveTeam did manage to scrape ~98.6% of the user profiles before it went down. It was around 16 hours from completion when Google pulled the plug. It was do
by neic 8y ago
ArchiveTeam did manage to scrape ~98.6% of the user profiles before it went down. It was around 16 hours from completion when Google pulled the plug.
It was done with the distributed scraper 'Warrior' using a massive amount of small cloud instances to spread the load to 1000s of IP adresses. The dataset is ~1.45 PB.
http://tracker.archiveteam.org/googleplus/ http://tracker.archiveteam.org/googleplus/
- empyrical 8y agoI don't know if ArchiveTeam's bots can be logged in as Google Accounts, but I can still (for the time being at least!) browse around on a GSuite account, and an old grandfathered-in "Google Apps for Business" account
- orev 8y agoI really wonder what is so valuable about saving everything on the Internet. For all of human history until now, every little human interaction was fleeting and only meant something to the people involved. Things that were significant were preserved when the people involved decided they were important. Now we are trying to keep every one of those insignificant things, for what purpose? Training an AI is all I can think of.
- stronglikedan 8y agoI wonder too. I was berated on here for mentioning that I deleted my reddit comments when I quit reddit, as if I was somehow stealing (my own content) from the users. Just weird IMO.
- badsectoracula 8y agoWell, as a Reddit user i dislike these "delete old comments" scripts since i very often read old discussions about topics i find interesting and when i hit the occasional "deleted by script foo, do it yourself too because reddit evil" i get irritated since they make following that old discussion hard. It is your account of course, but i still dislike the practice since it goes counter to the entire purpose of a discussion site.
- toomuchtodo 8y agoSeveral sites ingest and store Reddit comments forever. Scripts that overwrite old comments are ceremony for the user, nothing more.
- stronglikedan 8y ago> are ceremony for the user, nothing more. Only a small subset of users would be interested in those sites, and they likely don't come up in searches nearly as much as the original reddit post, so there's much more value to the deletion scripts than you are giving credit for.
- toomuchtodo 8y agoIronic considering you comment here and HN does not allow editing or deleting comments after a two hour window to protect discussion integrity.
- stronglikedan 8y agoWhat do find ironic about it? It's two entirely different types of discourse, so it's an apples-to-oranges comparison.
- NoodleIncident 8y agoA difference in the content of the two platforms doesn't mean that the comments on the two sites are so dissimilar that they can't be compared. In particular, they remove the forum requirement of quoting the comment you're replying to, by allowing you to respond to someone's comment directly below it in the middle of an existing thread. It would be exactly as inconvenient to readers of old Reddit and HN threads if that context was deleted automatically 6 months after the fact, HN just doesn't let you do it.
- stronglikedan 8y ago> doesn't mean that the comments on the two sites are so dissimilar that they can't be compared. Actually, it does. I discuss things of significance here, where the information may be useful to someone in the future, but I only commented on reddit for entertainment purposes (jokes, witty remarks, etc.) Also, creepy people don't comb through past HN comments looking to call you out in current conversations like they frequently do on reddit, and almost always out of context and spun to fit their agenda. While entertaining, it started getting annoying after a while, so in my opinion, that's why they can't have nice things, so to speak.
- est31 8y ago> Now we are trying to keep every one of those insignificant things, for what purpose? What seems unimportant and insignificant to us now might become very interesting in the future. People in the future might consider the arrival of the internet as the beginning of a new age. Furthermore, some things may seem obvious to a person born in our times, but might not be obvious to people born in the future. So many works, including even movies from the 20th century are lost now. E.g. the original version of Metropolis used during its premiere screening. Diogenes wasn't just a guy living in the town square, he also created some written works. They are lost now.
- bunderbunder 8y agoI'm less worried about the curiosity of people decades or centuries in the future than I am in the privacy interests of former Google+ users right now. Many of them chose that social network specifically because it was supposed to offer them better privacy protections. Which isn't a litmus test for whether or not they'd be OK with having their profile scraped and archived, but it is at least suggestive. It would be polite if ArchiveTeam were to now contact everyone whose profile they have scraped, and ask for permission to retain that data. And then delete all the profiles for which they didn't get affirmative consent.
- meruru 8y agoI can't feel sorry for someone who goes to Google for better privacy protections. I mean, really?
- dredmorbius 8y agoThe content IA are archiving was public and publicly accessible.
- bunderbunder 8y agoWell, for one, ArchiveTeam is unaffiliated with Internet Archive. Judging by the website, it's more closely associated with 4chan or Encyclopedia Dramatica or something. For two, being public and publicly accessible doesn't mean it isn't gauche to scrape it. It's kind of like how nobody sticks a sign that says "take one only" over the bowl of mints at a restaurant; it's assumed you just know that it's not cool to stick a whole handful of them into your pocket.
- ocdtrekkie 8y agoArchaeologically, I think people tend to be interested in a lot of things about society that don't get well-preserved. Social media offers potentially a time capsule to the lives of ordinary citizens today in the distant future. That being said, I intend to ask the Internet Archive to remove my personal G+ profile.
- deathhand 8y agoContent that is created is valuable to someone, somewhere. This is probably my favorite Google Plus post: https://webcache.googleusercontent.com/search?q=cache:9FEWJomYipIJ:https://plus.google.com/%2BJeanBaptisteQueru/posts/dfydM2Cnepe+&cd=1&hl=xx-hacker&ct=clnk&gl=us https://webcache.googleusercontent.com/search?q=cache:9FEWJo...
- kubbity 8y agoGreat post. Thanks!
- badsectoracula 8y agoMore often than not, what today sounds insignificant can become significant tomorrow - if not for everyone, then at least for some. Personally i have dug a lot on old websites stored at the web archive trying to find information and files (especially patches, older software and amateur games that few knew about) seemingly lost. For me the Archive (not just the web archive but all the archive.org projects) is as important as Wikipedia (and donate to both).
- rootsudo 8y agoAlso linking events and destroying plausible deniability.
- icebraining 8y agoBefore, people knew those were fleeting, and so they knew they had to record them if they wished to keep them. Current platforms give an idea of permanence, so people stopped worrying about preserving what they care about, and when one closes both are lost.
- orev 8y agoIf you look at the way most people use social media, they absolutely have no concept that they provide permanence. It’s exactly the opposite. People saving images solely to Flickr or videos on YouTube is a different usage pattern than the ephemeral likes and comments they make on the social platforms.
- adrusi 8y agoHistorically, information was more likely to be saved by accident. If something was published in print, that means that copies of it were distributed nationally or globally, found in hundreds of homes or libraries. If the publisher went out of business, people would still have physical publications lying around. And if something came to be of historical interest 50 years later, it could be preserved forever. If a web publication or community goes out of business, where's that data going to be 50 years later?
- lostphilosopher 8y agoThree additional reasons: We don't necessarily know what's "important" in advance. Archiving this content can be a nice public service for people who put content on the internet, don't do a good enough job preserving it themselves, and then want to access it later. I've heard "check the wayback machine" as a solution on HN for people whose blog got taken down unexpectedly from some provider and they didn't have a back up solution. Perhaps the same as the first reason, but some of this content can turn out to be useful as evidence. It's not uncommon for public figures to take down unflattering messages they posted on social media platforms and then deny having ever made them.
- vorpalhex 8y agoCrazy person trying to save everything on the Internet here (at least, a few sections of it.) It was actually when the Trump administration basically gutted and removed the whole of the EPA's website that caused me to take notice. I realized that while my written notes from a decade ago were still around and usable, every digital note I'd made more than a year or two ago was gone. Our knowledge culture has changed from memorizing facts to knowing how to get to facts. Information also no longer flows down from a "chosen few" who have the means to publish, but instead from everyone at just about all times. Archiving is also no longer a horrifically complicated or expensive hobby. There's no reprinting books on archive paper or storing them in helium. Instead you can simply buy a stack of hard drives and start bulk storing data. Since getting involved, I've come to realize that data goes offline frequently, with no warning - extensively due to bad copyright claims. I can't really fix the copyright process (at least, not very quickly), but I can help save the data and help make it available to others in meaningful formats.
- amelius 8y agoIt's a kind of "dragnet conservation".
- hobs 8y agoA lot of what we previously know about societies comes from things like midden piles, literal trash heaps of discarded crap nobody wanted.
- swiley 8y agoWhen I as younger I always used to love reading the silly things people archived. It gave me a little glimpse of what college was going to be like and helped me deal with being so isolated.
- dredmorbius 8y agoFor a lot of people and groups, the Internet Archive is their only hope at preserving at least some of their Google+ content. I've been working with many of them over the past six months. Google+ had millions of active users of a wide range of technical skills, but by far the majority at the lower end of the scale. Even the ones with technical chops were often limited by budget, network bandwidth, costs, or reliability, or other factors, in what they could do. Google+'s promise was to host text, image, and video content with Google's deserved reputation for high reliablity. The company itself is not, as many other failed social and online media services were, going out of business. It's simply decided to exit this particular activity. The first realisation I had of the problem Google+ Communities faced was when someone commented on the subreddit I'd created about the G+ shutdown that they were wondering how they were going to move 400,000 users and content to a new home. I didn't even know how many communities G+ had. The online information I'd found said about 5 million, but on checking it was actually 7.9 million as of late November, 2018, and over 8.1 million by January 2019. Through Loysoft (Friends+Me), I got a summary dump of all 8.1 million communities and some overview characteristics (members, posts and posting dates), and could finally start getting a handle on how many significant communities there were. That resulted in an extract of about 100,000 communities of 100+ members and posting activity within the previous 30 days, which was used for outreach and migrations. I made that freely available to anyone looking to assist in migrations. But a lot of communities were missed, Google itself shut down the G+ Communities serving Community owners and moderators, with no notice, making outreach all but impossible. Is there a lot of crap data out there? Yes, there is. But there are also some gems, and the task of sorting between the two in advance of archiving it is more effort than simply archiving everything and making it available. And the Internet Archive has set itself the mission of total archival, where possible. So there's that. The questions are ones we're discussing though at the PlexodusReddit: https://old.reddit.com/r/plexodus https://old.reddit.com/r/plexodus
- cblades 8y ago>Things that were significant were preserved when the people involved decided they were important Sounds like you answered your own question.
- jrochkind1 8y agoThe other side of this is for that the past centuries things that were on _paper_ (letters, photographs, newsletters, zines, personal notes/journals, whatever) could be just left sitting around, and _some_ portion of them would still be there and legible decades+ later for historical purposes. When it's digital... if you don't keep feeding it, it's gone forever. There has been some coverage of historians and archivists concerned about this, here's just the first random thing i found googling: https://abcnews.go.com/Technology/digital-era-end-history/story?id=97367 https://abcnews.go.com/Technology/digital-era-end-history/st... But yes, professional archivists increasingly recognize you can't save _everything_ digital, we couldn't afford it. Figuring out what to save is a challenge. But if we save nothing, we're not gonna have much history to look at.
- dredmorbius 8y agoArchive Team are absolute heros. The Google+ Mass Migration community learned of them and their "googleminus" project in January, and worked to help give information on the crawl, amount of data, and particulars of G+. arkiver and Fusl in particular have been absolutely amazing in what they've accomplished. They also managed to pull in 94.5% of all Google+ Communities, which should provide the ability to view posts by Community (they're otherwise scattered among user posts). We're still assessing how much of that was login-page redirects in the last hour or two of the crawl, but it's amazing work. I'd managed to send of about 80k larger, recently-active (100+ members, <30 day activity) over the past few weeks, with a final grab about 18 hours ago, using the Internet Archive's "save" URL. If you ever need to use that it's: https://web.archive.org/save/<URL> Where you replace "<URL>" with whatever it is you're trying to save, including the protocol string, say, this HN post: https://web.archive.org/save/https://news.ycombinator.com/item?id=19556665 That can be scripted, and my submissions used a bog-simple Bash script and xargs to plow through 100k submissions (20k appear to have been dead) in about 90 minutes, on very modest hardware. Also: the Internet Archive (and Archive Team) run off volunteers and donations. You can help, and please do. https://archive.org/donate/ https://archive.org/donate/ https://www.archiveteam.org/index.php?title=Donate https://www.archiveteam.org/index.php?title=Donate (Not affiliated, but very grateful to them.)
- dymk 8y agoDoes the ArchiveTeam coordinate with services like Google+ for archival? Are the techniques used (such as spreading the load across 1000s of separate IPs) used for anything other than getting around Google's integrity services e.g. ratelimiting? I'm really curious how such a large operation is legally pulled off, especially when services like Google+ intentionally try to make scraping difficult (and presumably for large offenders, will attempt a C&D).
- stordoff 8y agoThey use the ArchiveTeam Warrrior VM, so they have access to many public IPs: https://www.archiveteam.org/index.php?title=ArchiveTeam_Warrior https://www.archiveteam.org/index.php?title=ArchiveTeam_Warr...
- ihuman 8y agoSometimes the operators of a service will contact the ArchiveTeam and work with them to speed up the archival process. The file hosting service pomf.se did this when they were shutting down. https://www.archiveteam.org/index.php?title=Pomf.se#Archiving https://www.archiveteam.org/index.php?title=Pomf.se#Archivin...
- creato 8y agoI've been surprised that there hasn't been much pushback/questioning (from what I've seen) on this archive effort. Of course people voluntarily published the information here. But they also had at least a little control over how and what was presented to the world. Now they don't. Is archive team going to respond if someone finds something in the archive they want to take down? Is there a listing somewhere of exactly what is archived (does it include pictures from google plus?)? I couldn't find it, and I don't want to go rooting around in the actual archive itself (if that is even possible).
- flatline 8y agoI find it curious and saddening that this is becoming an expectation. Someone voluntarily publishes something for the world to see - they are certainly within their rights to issue a retraction, but they have long since relinquished control over others’ actions with the data they published. This is pretty fundamental to how the internet works. The fact that there is now an expectation of continued control over said data just shows how far companies like Facebook have gone toward fundamentally changing the nature of the internet.
- sebthedev 8y agoUnder many data privacy laws, data subjects have a legal "right to correct" any inaccurate personal data concerning him or her. GDPR: https://gdpr-info.eu/art-16-gdpr/ https://gdpr-info.eu/art-16-gdpr/
- deleted 8y ago[deleted]
- bunderbunder 8y agoThe internet was created in a different age. The kinds of abuses people are worrying about right now were barely even possible when the Internet was created. Now they're quite feasible using relatively inexpensive and well-known technology, and people who are eager to share the knowledge of how to do it have created all sorts of MOOCs and bootcamps and even accredited degree programs on the subject.
- mbell 8y agoThis is rather creepy if I'm honest - why are you scraping user's profile data? Is there a way I can request that you delete any of my data that may be in there?
- toomuchtodo 8y agoOnly what was public was scrapped. You can make a request to the Internet Archive to remove your content when they have ingested it and made it available.