7 ms·
Help preserve the internet with Archiveteam's warrior
- iforgotpassword 5y agoHow likely is it you end up downloading child porn on behalf of them? In other words, how well curated or specific is the list of download jobs your node gets assigned? If it's something like "just grab everything from this blog platform" I guess chances are not zero.
- mhitza 5y agoI think you would be more likely to win the lottery without playing. That type of content has long moved from clearnet to the darknet. I would be inexplicably surprised if that type of content can be found on the clearnet. But I still can be wrong. However if you're in the US loli hentai is going to be a risk and legal headache for sure https://www.shouselaw.com/ca/blog/is-loli-illegal-in-the-united-states/ https://www.shouselaw.com/ca/blog/is-loli-illegal-in-the-uni... As far as I'm aware, maybe excepting Australia (?) as well, in the rest of the world that type of content is not something they'll classify as child pornography, you'll just get a few sketchy looks.
- charcircuit 5y ago>That type of content has long moved from clearnet to the darknet. A fraction of it. >I would be inexplicably surprised if that type of content can be found on the clearnet That kind of content is a single internet search away.
- smarx007 5y agoMy experience has shown that list to be extremely well-curated. See https://wiki.archiveteam.org/#Warrior-based_projects https://wiki.archiveteam.org/#Warrior-based_projects for the current list. Though if you join the Reddit archival project, all bets may be off but that's not AT team's fault, I guess.
- uniqueuid 5y agoYes! Archiving is important, we have already seen so much online history gone down the drain or just accidentally saved. Large institutions like the internet archive are doing an admirable job, but there is a lot of content that they cannot and will not cover. So we will definitely (also) need volunteer-based archival for the foreseeable future. 18TB drives are ~$300 a piece right now, go buy one and help our collective memory!
- prox 5y agoI kind of wonder how we can make it searchable again. Is this included in this archiving effort? In any case wonderful work.
- uniqueuid 5y agoThere is a standard set of tooling for indexing archives: CDX files. [1] They index WARC archives and can be used to quickly find records. You can build on top of this (and some systems do) to make a proper search front-end. But in general, these archives are NOT geared towards full-blown search because it would be pretty expensive to keep the indexes in hot cache. Plus you would need to deal with historical versions of records, which is not normally done in search UX. [1] https://wiki.archiveteam.org/index.php/The_WARC_Ecosystem#CDX_File_Format https://wiki.archiveteam.org/index.php/The_WARC_Ecosystem#CD...
- camtarn 5y agoAh, is the WARC format the reason it's called 'Warrior'? It seems like a very strange name for an archival program.
- myself248 5y agoArchiveTeam seems very guerrilla in their operations. I always imagined the Warrior as a camo-faced archivist operating under cover of darkness, preserving data even in the most hostile Yahoo-occupied territory.
- 5y ago
- Thorentis 5y agoThe Internet Archive has a huge noise to signal ratio, very much in favour of noise. I admire the effort and regularly make use of the quality archives. However, I wonder if much like Bitcoin, tremendous energy and amounts of resources are being put towards very little of value.
- sandgiant 5y agoCan you provide some details on this? I'm curious how noise and signal are defined and measured in this case.
- azeirah 5y agoI disagree with the op. This is historical data and includes all kinds of interesting content. Even if severely uninteresting today it may still be really valuable 40 years from now as part of research into colloquial language, design, trends, influence of events etc. Same reason why notes taken by random people 250 years ago are really valuable to historians today, even if it's just a todo list
- prox 5y agoValue is really hard to predict, but as someone who researches a lot in archives, there is no such thing as too little information. Especially if you want the views of several parties or organizations. In anthropology and history research this work (archiving) can be of tremendous value. Usually it’s hard to say if it’s valuable now , only time can tell.
- DoingIsLearning 5y agoI disagree, the unfiltered high noise is what makes it valuable. Curation is a bias. If someone wants to dive into any topic in the archive 30 years from now they will have access to everything, not access to what some of us deem 'worthy' of curating. I agree that it makes it harder to find things but I also see the value of IA as a time capsule.
- londons_explore 5y ago
- RNAlfons 5y agoMake it an easy installable/runable Windows application and it will spread like wildfire.
- capableweb 5y agoIf it was only that easy. To make distributed archiving as high quality as possible, you need reproducible environments as much as possible, which is why the "official" way of participating is to run virtual machines, instead of directly on the host. Not sure why this 3rd party is the submission site rather than the official page, which is this: https://wiki.archiveteam.org/index.php/ArchiveTeam_Warrior https://wiki.archiveteam.org/index.php/ArchiveTeam_Warrior Has a couple of different installation methods as well.
- jrwr 5y agoYep, Using Virtual box is rather easy to get the warrior running!
- cxr 5y agoWhy even require that? If the data in question is available over HTTP, it should be as easy as opening a page from the relevant origin in a browser tab, optionally opening a second tab for a "Warrior Dashboard", then invoking a bookmarklet on the former to slurp up data by XHR &tc. (If it's necessary to cross origins as the thing roves around, the dashboard can alert you to this while it continues doing what it can with the first origin. Just have the human return to the dashboard from time to time and repeat the second step to run as many in parallel as they want.)
- TheTechRobo 5y agoSimilar: github.com/InternetArchive/warcprox
- myself248 5y agoThat would be awesome, do you think you could write that?
- causi 5y agoWarrior is great for the community effort, but I wish someone would put some work into a modern local site archiver. HTTRACK just doesn't cut it anymore.
- myself248 5y agoOh jeez yeah. I've been going through https://github.com/ArchiveBox/ArchiveBox/wiki/Web-Archiving-Community https://github.com/ArchiveBox/ArchiveBox/wiki/Web-Archiving-... the last few days and I've concluded that none of 'em are appropriate for someone with my level of software ineptitude.
- nix23 5y agowget --recursive --page-requisites --adjust-extension --convert-links --no-parent https://YOURWEBPAGEHEREX.com https://YOURWEBPAGEHEREX.com NO "--convert-links" if you want a "pure" non local browsable copy.
- myself248 5y agoYes yes fine, and then I get throttled to 2 bytes/sec by the server. So I did some user-agent hijinks and set my delay to like 5000msec and that helped for a while, but my machine crashed and when I went to resume the task I was throttled again.
- nix23 5y ago>but my machine crashed Maybe it's not the servers who throttle you then ;)
- traverseda 5y agoWget will exhaust all available ram on a long enough crawl.
- 5y ago