9 ms·
Inside The Internet Archive's Infrastructure
https://github.com/internetarchive/heritrix3 https://github.com/internetarchive/heritrix3
- BryantD 9mo agoThey have come a very long way since the late 1990s when I was working there as a sysadmin and the data center was a couple of racks plus a tape robot in a back room of the Presidio office with an alarmingly slanted floor. The tape robot vendor had to come out and recalibrate the tape drives more often than I might have wanted.
- textfiles 9mo agoThere is a fundamental resistance to tape technology that exists to this day as a result of all those troubles.
- EvanAnderson 9mo agoThat's sad, but it mirrors my experience with commercial customers. Tape is so fiddly but the cost efficiency for large amounts of data and at-rest stability is so good. Tape is caught in a spiral of decreasing market share so industry has no incentive to optimize it. Edit: Then again, I recently heard a podcast that talked about the relatively good at-rest stability of SATA hard disk drives stored outdoors. >smile<
- duskwuff 9mo agoTape is also an extraordinarily poor option for a service like Internet Archive which intends to provide interactive, on-demand access to its holdings.
- stonogo 9mo agoThis is a common use for tape, which can via tools like HPSS have a couple petabytes of disk in front of it, and present the whole archive in a single POSIX filesystem namespace, handling data migration transparently and making sure hot data is kept on low-latency storage.
- BryantD 9mo agoYeah, it was like this (except not petabytes).
- EvanAnderson 9mo agoI presume backing-up the archive is a desirable thing. That's a place where I would see tape fitting well for them.
- duskwuff 9mo agoPerhaps? But unless tape, and the infrastructure to support it, is dramatically cheaper than disk, they might still be better served by more disk - having two or more copies of data on disk means that both of them can service load, whereas a tape backup is only passively useful as a backup.
- deleted 9mo ago[deleted]
- stonogo 9mo agounless tape, and the infrastructure to support it, is dramatically cheaper than disk, This turns out to be the case, with the cost difference growing as the archive size scales. Once you hit petascale, it's not even close. However, most large-scale tape deployments also have disk involved, so it's usually not one or the other.
- xk3 9mo agoYou might squirm at using refurbished or used media but those 3TB SAS ex-enterprise disks are often the same price or cheaper than tapes themselves (excluding tape drive costs!). Will magnetic storage last 30 years? Probably not but they don't instantly demagnetize either. Both tape and offline magnetic platters benefit from ideal storage conditions.
- EvanAnderson 9mo agoIt's not just cost / media, though. Automated handling is a big advantage, too. At the scale where tape makes sense (north of 400TB in retention) I think the inconvenience of handling disks with similar aggregate capacity would be significant. I guess slotting disks into a storage shelf is similar to loading a tape changer robot. I can't imagine the backplane slots on a disk array being rated at a significant lifetime number of insertions / removals.
- BryantD 9mo agoBack in the day, if you loaded a page from the web archive that wasn’t in cache, it’d tell you to come back in a couple of minutes. If it was in cache, it was reasonably speedy. Cache in this case was the hard drives. If I recall correctly, we were using SAM-FS, which worked fairly well for the purpose even though it was slow as dirt —- we could effectively mount the tape drive on Solaris servers, and access the file system transparently. Things have gotten better. I’m not sure if there were better affordable options in the late 1990s, though. I went from Alexa/IA to AltaVista, which solved the problem of storing web crawl data by being owned by DEC and installing dozens of refrigerator sized Alpha servers. Not an option open to Alexa/IA.
- Melatonic 9mo agoTape is almost always used for cold storage backups that are offline in case of ransomware attacks. Using it for on demand access would be insanely slow
- hinkley 9mo agoWe had a little server room where the AC was mounted directly over the rack. I don't think we ever put an umbrella in there but it sure made everyone nervous the drain pipe would clog. Much more recently, I worked at a medium-large SaaS company but if you listened to my coworkers you'd think we were Google (there is a point where optimism starts being delusion, and a couple of my coworkers were past it.) Then one day I found the telemetry pages for Wikipedia. I am hoping some of those charts were per hour not per second, otherwise they are dealing with mind numbing amounts of traffic.
- brcmthrowaway 9mo ago[flagged]
- brcmthrowaway 9mo agoDoes IA do deduplication?
- textfiles 9mo agoNot in the way I think you're talking about. The archive has always tried to maintain a situation where the racks could be pushed out of the door or picked up after being somewhere and the individual drives will contain complete versions of the items. We have definitely reached out to people who seem to be doing redundant work and ask them to stop or for permission to remove the redundant item. But that's a pretty curatorial process.
- HumanOstrich 9mo ago[flagged]
- sltkr 9mo agoI don't think the article mentions anything about deduplication. Can you be less snarky and actually quote the relevant sentence?
- zxcvasd 9mo agoheres the second paragraph in full: "Here, amidst the repurposed neoclassical columns and wooden pews of a building constructed to worship a different kind of permanence, lies the physical manifestation of the "virtual" world. We tend to think of the internet as an ethereal cloud, a place without geography or mass. But in this building, the internet has weight. It has heat. It requires electricity, maintenance, and a constant battle against the second law of thermodynamics. As of late 2025, this machine—collectively known as the Wayback Machine—has archived over one trillion web pages.1 It holds 99 petabytes of unique data, a number that expands to over 212 petabytes when accounting for backups and redundancy.3" can you help my small brain by pointing out where in this paragraph they talk about deduplication?
- hedora 9mo agoIt's frustrating that there's no way for people to (selectively) mirror the Internet Archive. $25-30M per year is a lot for a non-profit, but it's nothing for government agencies, or private corporations building Gen AI models. I suspect having a few different teams competing (for funding) to provide mirrors would rapidly reduce the hardware cost too. The density + power dissipation numbers quoted are extremely poor compared to enterprise storage. Hardware costs for the enterprise systems are also well below AWS (even assuming a short 5 year depreciation cycle on the enterprise boxes). Neither this article nor the vendors publish enough pricing information to do a thorough total cost of ownership analysis, but I can imagine someone the size of IA would not be paying normal margins to their vendors.
- toomuchtodo 9mo agoPick the items you want to mirror and seed them via their torrent file. https://help.archive.org/help/archive-bittorrents/ https://help.archive.org/help/archive-bittorrents/ https://github.com/jjjake/internetarchive https://github.com/jjjake/internetarchive https://archive.org/services/docs/api/internetarchive/cli.html https://archive.org/services/docs/api/internetarchive/cli.ht... u/stavros wrote a design doc for a system (codename "Elephant") that would scale this up: https://news.ycombinator.com/item?id=45559219 https://news.ycombinator.com/item?id=45559219 (no affiliation, I am just a rando; if you are a library, museum, or similar institution, ask IA to drop some racks at your colo for replication, and as always, don't forget to donate to IA when able to and be kind to their infrastructure)
- billyhoffman 9mo agoThere are real problems with the Torrent files for collections. They are automatically created when a collection is first created and uploaded, and so they only include the files of the initial upload. For very large collections (100+ GB) it is common for a creator to add/upload files into a collection in batches, but the torrent file is never regenerated, so download with the torrent results in just a small subset of the entire collection. https://www.reddit.com/r/torrents/comments/vc0v08/question_about_archiveorg_torrents_being/ https://www.reddit.com/r/torrents/comments/vc0v08/question_a... The solution is to use one of the several IA downloader script on GitHub, which download content via the collection's file list. I don't like directly downloading since I know that is most cost to IA, but torrents really are an option for some collections. Turns out, there are a lot of 500BG-2TB collections for ROMs/ISOs for video game consoles through the 7th and 8th generation, available on the IA...
- cowhax 9mo ago>And the rising popularity of generative AI adds yet another unpredictable dimension to the future survival of the public domain archive. I'd say the nonprofit has found itself a profitable reason for its existence
- schmuckonwheels 9mo agoDisappointed with the lack of pictures.
- parttimelarry 9mo agoProbably because this looks more like a Deep Research agent "delving" into the infrastructure -- with a giant list of sources at the end. The Archive is not just a library; it is a service provider.
- schmuckonwheels 9mo agoI wasn't expecting to read a podcast when clicking.
- textfiles 9mo agoWhat do you want some pictures of?
- schmuckonwheels 9mo agoAn article about "infrastructure" that opens up with a dramatic description of a datacenter stuffed into an old church, I would expect more than just generic clipart you'd see in the back half of Wired magazine.
- textfiles 9mo agoHere's some photos I took a long time ago. https://www.flickr.com/photos/textfiles/albums/72157633722203885/ https://www.flickr.com/photos/textfiles/albums/7215763372220...
- Tempest1981 9mo agoThanks! The church attendees (employees?) have a Severence Kier vibe... although I'm guessing the TV show came much later.
- 9mo ago
- mcpar-land 9mo agoIs this some kind of copypasted AI output? There are unformatted footnote numbers at the end of many sentences.
- NetOpWibby 9mo agoI was thinking the same thing. No proofreading is a sure sign to me. I also feel like I've read parts of this before.
- sltkr 9mo agoSome of the images are AI generated (see the Gemini watermark in the bottom right), and the final paragraph also reads extremely AI-generated.
- dvrp 9mo agoMaybe, but I was trying to find the original source of this article and couldn’t, at least not cursorily.
- ramon156 9mo agoI already stopped when I saw the AI-gen image
- eiiot 9mo agoThe table also seems like the kind of thing that Gemini seems to generate a lot. "Here's a table that communicates almost no information! One of the rows is constant for each item."
- lysace 9mo agoThe IA needs perhaps not just more money, but also more talented people, IMO. I worry that it has stagnated, from a tech pov.
- deleted 9mo ago[deleted]
- mixologic 9mo agoThey can offer a perk that literally no other tech job can offer: Someday have a statue of your likeness preserved in ceramic: https://www.atlasobscura.com/places/internet-archive-headquarters https://www.atlasobscura.com/places/internet-archive-headqua... "Inside the church's main room, with its still-intact pews, there are more than 120 ceramic sculptures of the Internet Archive's current and former employees, created by artist Nuala Creed and inspired by the statues of the Xian warriors in China."
- textfiles 9mo agoWe've hired a few dozen people over the past couple of years. We think they're pretty talented.
- lysace 9mo agoIs retreival from the wayback machine intentionally made slow?
- textfiles 9mo agoShow me the faster wayback machine we are competing against.
- brokensegue 9mo agoi'm a big fan of IA and wayback machine. i donate. but i do wish it were faster. i understand that would cost a lot more though. i wonder if maybe donors above a certain level could get priority on archiving pages or something.
- krunck 9mo ago[flagged]
- mjmas 9mo agoWas this reply meant for this story instead? https://news.ycombinator.com/item?id=46637127 https://news.ycombinator.com/item?id=46637127
- rarisma 9mo agoI think this was writen wholly by deep research. It just reads like a clunky low quality article
- astrange 9mo agoIt's clearly AI writing ("hum", "delve") but oddly I don't think deep research models use those words.
- joemi 9mo agoI think relying on the vocabulary to indicate AI is pointless (unless they're actually using words that AI made up). There's a reason they use words such as those you've pointed out: because they're words, and their training material (a.k.a. output by humans) use them.
- astrange 9mo agoNo American used "delve" before ChatGPT 3.5, and nobody outside fanfiction uses the metaphors it does (which are always about "secrets" "quiet" "humming" "whispers" etc). It's really very noticeable. https://www.nytimes.com/2025/12/03/magazine/chatbot-writing-style.html https://www.nytimes.com/2025/12/03/magazine/chatbot-writing-...
- pests 9mo agoBut now Americans do use "delve" since 3.5. So what? No Americans used "cromulent" as a word either until Simpsons invented it. Is it not a real word? Does using it mean the Simpsons wrote it?
- ashtonshears 9mo agoI bet the llm is biased towards the mtg card delver of secrets
- joemi 9mo agoThe link you posted doesn't back up the statement that "No American used "delve" before ChatGPT 3.5". Instead it states that _few_ people used it in _biomedical papers_. I've seen it (and metaphors using the other words you noted) used in fiction for my entire life, and I sure as hell predate chatgpt. This is why it's a bad idea to consider every use of particular words to be AI generated. There are always some people who have larger vocabularies than others and use more words, including words some people have deemed giveaways of AI use. That said, their use may raise suspicion of AI, but they are _not_ proof of AI. I don't want to live in a world where people with large vocabularies are not taken seriously. Such an anti-intellectual stance is extremely dangerous.
- bpiche 9mo agoIA is hosting a couple more of Rick Prelinger’s shows this month. Looking forward to visiting
- ghm2199 9mo agoDoes any one know how the size of this compares to archive.today?
- textfiles 9mo agoWe absolutely lap them with many, many more petabytes of material. But archive.today is also not doing speculative or multiple scheduled captures of the amount of sites that archive.org is.
- vladiim 9mo agoHow long will it take for them to send the PetaBox to space?
- textfiles 9mo agoThat project gets discussed every once in a while.
- semiquaver 9mo agoThis article is way too LLMey for my taste.
- alfgrimur 9mo agoI love to imagine this is all a cover and the Internet Archive is located in a remote cave in northern Sweden and consists of a series of endlessly self replicating flash drives powered by the sun.
- segalord 9mo agothis is every data hoarders dream setup haha
- jarboot 9mo agoHate to be the guy in the comments complaining about the css, but the sides of the text of this article are cut off. It looks like I'm zoomed in, and there's no way I can see the first few columns of the text without going to Reader view. I'm on a modern iPhone using safari, accessibility settings font larger than usual.
- shmeeed 9mo agoFWIW, it's the same for me on FF Android.
- nandomrumber 9mo agoSame for me, Safari iOS 18.7.1 no accessibility font size set, no browsers font size set.
- textfiles 9mo agoIt's an AI-generated article. It's going to be pretty terrible.
- initialg 9mo agoIs it still year 2006 and websites haven’t figured out responsive design?
- fedeb95 9mo agoThanks for this, I've always wondered how the Archive operates but always ended up not searching.
- ThinkBeat 9mo agoWow that piece of real-estate has to cost a bundle.
- bilater 9mo agoI have always wondered how archives manage to capture screenshots of paywalled pages like the New York Times or the Wall Street Journal. Do they have agreements with publishers, do their crawlers have special privileges to bypass detection, or do they use technology so advanced that companies cannot detect them?
- deleted 9mo ago[deleted]