15 ms·
Download the Entire Wikimedia Database
- nayuki 6y agoBack in 2014 I computed the PageRanks within English Wikipedia, thanks to their database dump. https://www.nayuki.io/page/computing-wikipedias-internal-pageranks https://www.nayuki.io/page/computing-wikipedias-internal-pag...
- crazygringo 6y agoThat's intriguing. Curious if you ever compared how PageRanks correlate to traffic? (They make their per-page traffic available too.) It would be interesting to see the largest disparities -- super-popular pages in visits but which don't have nearly as many internal Wikipedia links to them, versus unpopular pages but that have tons of internal Wikipedia links to them.
- tomaszs 6y agoWhat page had the highest PageRank?
- vinger 6y agoThe homepage.
- nathcd 6y agoHere is a plaintext version of Nayuki's results, via the link they posted above: https://www.nayuki.io/res/computing-wikipedias-internal-pageranks/wikipedia-top-pageranks.txt https://www.nayuki.io/res/computing-wikipedias-internal-page... Geographic coordinate system is first.
- bawolff 6y agoMaybe would be better to separate content links from more technically links that come from standardized templates which aren't really part of the article content in a certain sense
- orblivion 6y agoYou can also get it in a user-friendly format with the application Kiwix (https://www.kiwix.org/ https://www.kiwix.org/) if that's your use case. PC, phone, or server. You get subsets of the data, and images are smaller to save space.
- MeinBlutIstBlau 6y agoKiwix use is still somewhat hit or miss when browsing. Im not sure how it handles text parsing but it either takes forever or doesn't return results making it sort of unusable. But it's still a fantastic and incredible piece of software. When it gets to the point where I can portably keep the full 60gb zim file seemlessly, it will change simple computer for low broadband areas. Imagine the uses as a portable and versatile database that could accept json, html/css, data to make your own offline encyclopedias!
- kregasaurusrex 6y agoIt's useful to browse on a phone if you have limited mobile data, and the text-only English Wikipedia fits onto a modern micro SD card.
- m-p-3 6y agoThe complete archive with image is 82GB, well inside the capacity of an affordable modern microSD card.
- deleted 6y ago[deleted]
- puddingnomeat 6y agokiwix includes a relation to Qt5Core.dll with an invalid signature, not sure if relevant to anybody
- puddingnomeat 6y agoand why is it related to mbam? https://answers.microsoft.com/en-us/windows/forum/windows_10-files-winpc/qt5coredll-missing-from-computer/76e17b8e-e3b6-4f61-b351-04d72acac3b0 https://answers.microsoft.com/en-us/windows/forum/windows_10...
- dwheeler 6y agoI'm so glad the download-entire-wikipedia function continues to exist. That will help counter the "lost the entire library problem" from the city of Alexandria. To be fair, Wikipedia only has summaries, not the detailed material, but it's still important.
- porphyra 6y agoIt is pretty awesome that there are people like /r/datahoarder that are obsessed with backing up the collective knowledge of humanity.
- capableweb 6y agoI'm not familiar with r/datahoarder, but if the name bears any significance, it seems they are mostly centered on hoarding data, which means just digital I guess? If so, I much rather would want to promote efforts like Internet Archive that back up all kind of things, not just digital data.
- deleted 6y ago[deleted]
- KMnO4 6y agoWhat does the Internet Archive back up that isn’t represented by 1s and 0s?
- cguess 6y agoA ton of 35mm and 16mm film reels, vinyl, and even wax recording, physical books and a lot more. They make digital copies of them, but they also archive the physical versions as well. Here's a selection of the movies: https://archive.org/details/moviesandfilms?tab=about https://archive.org/details/moviesandfilms?tab=about
- LeoPanthera 6y agoThe Physical Archive. https://en.wikipedia.org/wiki/Internet_Archive#Physical_media https://en.wikipedia.org/wiki/Internet_Archive#Physical_medi...
- bawolff 6y agoWell not the entire db, just the public parts. User passwords are not included ;)
- dudus 6y agoWikipedia is always bugging me about donations, and yet here it is a feature they could charge for or at least hint to donate. It would be perfectly acceptable to charge here since abuse of this can rack up quite a bill. Maybe they don't pay as much as I do for outbound traffic on aws, but still
- capableweb 6y ago> would be perfectly acceptable to charge here since abuse of this can rack up quite a bill Not according to Wikipedia. Wikipedia much rather beg people from all corners of the world to donate, than restricting access to their data. That's what a good, honest and well-meaning foundation does. And yes, no sane person shuffling a lot of data around is using AWS because of their awful bandwidth pricing, Wikipedia included.
- morsch 6y agoI guess a hint would be fine, but charging for access, even bulk access, feels quite contrary to the spirit of the project. It excludes huge ranges of people who cannot afford it or don't have access to Internet payment methods. I suspect the traffic caused by this is minuscule compared to the overall traffic, anyway. But that's just a guess.
- bawolff 6y ago> Maybe they don't pay as much as I do for outbound traffic on aws, but still Almost certainly not. I doubt this is even a rounding error on their budget. I'm not sure how this stuff works at a high level (IANA network engineer), but i think they have peering presence at internet exchange points - https://www.peeringdb.com/asn/14907 https://www.peeringdb.com/asn/14907 so im not sure they actually pay for bandwidth at all, at least on some routes
- karlicoss 6y agoWouldn't it be cool if Steam supported distributing offline Wikipedia database? It's just a few gigs (depending on languages/images/etc, but it fits the DLC model perfectly), and it already uses bittorrent.
- m-p-3 6y agoAn IPFS cluster for Wikipedia like this https://collab.ipfscluster.io/ https://collab.ipfscluster.io/ would be nice.
- libraryofalex 6y agoOne of the only things you can do to ensure lasting democracy today is to download the pages, with complete history, put it on a usb drive or microsd card properly labelled for you to keep offline, and just forget about it. You can do this as a consumer, it's easy. There's no harm in it, it's not some kind of private data such as personal photos or documents. If you end up forgetting or losing track of it, it really is no big deal. You just decided to download it when you saw it on hacker news back in 2021, right? My reason for saying this is one of the only things you can do to ensure lasting democracy is that it is in the realm of what is possible in a physical sense that at some point through some mechanism the online version simply does not inform the public on some important public issue, whereas the history as you can download it today does. Though, I wouldn't speculate about what the mechanism might be or what kinds of subject. At that point in a physical sense you could consult your offline copy on an airgapped PC or future equivalent and I think it would be impossible for any group of any kind to even know you were doing that let alone stop it. How you might get the word out is another question but having this personal capability is easy for the people here, as technical users and simple consumers. Indeed the whole entire Internet was set up as a distributed network in case of nuclear attack, so the entire topology of the Internet is set up for you to do this easily today. It's a click and a cheap flash drive or slightly more expensive microsd card away. You can take this step in less than 20 active minutes of your time and for less than $50 if you go with an external spinning disk drive (such as 1 terabyte) or $200 or so if you go with a microsd card. It doesn't really matter if the file ultimately fails, this is not a critical backup for you to have just a nice to have. You could write the file's checksum onto the drive in marker so you can tell whether it's still correct later (as opposed to having bit errors). Maybe there is some file type that has a bit of redundancy (checksums) for long-term storage, since due to the large amount (several hundreds gigabytes) I wouldn't be all that surprised if a few bits flipped over the course of several years in cold storage. But I don't know what kind of file type has any sort of redundancy or parity built into it that is supposed to protect against this. (Does anyone know?) Most likely the hash just wouldn't match what you wrote in pen on it but it would still be useable. Regarding choice of spinning disk or microsd card: I guess it's in the realm of what's possible in a physical sense that at some point people would have their personal property rummaged through by some group and a hard drive is pretty obvious and could be stolen or removed for that reason. (In a physical sense, not speculating about social or political developments that might lead to that.) So for this reason perhaps best would be to put it on a microsd card even though it is quite a bit more expensive. I guess written once, bit rot causes microsd cards to decay within a few years if not used at all.[1] I don't know for spinning media but I guess it's also about 5-10 years at least.[2] You could put the microsd card under a postage stamp for example and put an important unrelated document into the envelope, which you would expect to keep for many years. Of course you could always end up accidentally discarding your envelope (while retaining its contents) but that risk shouldn't matter too much. In a physical sense it is possible for groups to x-ray all paperwork (such as envelopes as I just suggested) and a microsd card's electrical contacts are pretty obvious in an x-ray. (It looks like this [3]). I don't have any suggestion that works against this attack, which is within the realm of what's possible according to the laws of physics. I'm not speculating on what social or political developments might possibly make anything like this necessary at some point in the future, but we still live in a world governed by the laws of physics so as technical professionals you have a huge leg up on most of the world. Spending $50 doing this today might save democracy tomorrow. You could also leave it as a time capsule however the storage longevity is not that long (between 5 and 20 years I guess), and in a physical sense, a time capsule is not particularly secure and would require instructions for someone else to figure out so it's not great in that sense. So in terms of what you can do today, I would suggest just getting an external 1 terabyte usb drive ($50), downloading the dump together with history (20 active minutes), writing the checksum onto it in marker and just putting it somewhere. Obviously this small $50 investment is one you would hope never to have to use, but who knows, you might go down in history as the one who saved some small part of the world. Though, obviously, not in Wikipedia history. [1] https://www.quora.com/What-is-the-longevity-of-a-sd-memory-card https://www.quora.com/What-is-the-longevity-of-a-sd-memory-c... [2] https://serverfault.com/questions/986911/how-long-will-unused-hard-drive-last https://serverfault.com/questions/986911/how-long-will-unuse... [3] https://www.reddit.com/r/pics/comments/3b6bjw/i_xrayed_an_sd_card/ https://www.reddit.com/r/pics/comments/3b6bjw/i_xrayed_an_sd...
- guerby 6y agoIs the media content (images, videos) downloadable? When I follow the links I find 2012-2013 data but may be I missed something? https://meta.wikimedia.org/wiki/Mirroring_Wikimedia_project_XML_dumps https://meta.wikimedia.org/wiki/Mirroring_Wikimedia_project_...
- bawolff 6y agoNo. It starts to get prohibitive in terms of file size. There are 288 TB of uploaded media https://commons.wikimedia.org/wiki/Special:MediaStatistics https://commons.wikimedia.org/wiki/Special:MediaStatistics + https://en.wikipedia.org/wiki/Special:MediaStatistics https://en.wikipedia.org/wiki/Special:MediaStatistics . Much of that isnt used on english wikipedia, but nonetheless. I suspect if you had some genuinely good reason why you wanted it, and asked real nicely, and provided some way to transport the data, you might be able to make an arrangement of some sort to get it.
- ris 6y agoGenuine question: why is bittorrent not being used for this?
- bawolff 6y agoI imagine because wikimedia has lots of bandwidth (once your website gets to be a certain size, bandwidth gets sold very differently), new versions come out regularly, and there is a relatively small number of people who want every single one of these and keep them around to seed.
- 91aintprime 6y agoIt would be interesting intermittent releases supplemented with diffs
- bradleybuda 6y agoBecause they wanted it to be easy for people to get?
- Darthkoax 6y agoIt is available as a torrent
- sleavey 6y agoFurther down the page: > Backup dumps of wikis which no longer exist [...] This includes, in particular, the Sept. 11 wiki. There was a Sept. 11 wiki hosted by Wikimedia?
- duskwuff 6y agoA while ago, yes. https://meta.wikimedia.org/wiki/Sep11wiki https://meta.wikimedia.org/wiki/Sep11wiki
- Santosh83 6y agoAre images/media from Wikimedia Commons included in these dumps?
- bawolff 6y agoNo. Its just the wikitext (markup language) source of all pages.
- nromiun 6y agoI wish Wikipedia would offer incremental downloads (e.g. rsync). That would make it much easier to host your own Wikipedia.
- deleted 6y ago[deleted]
- jokoon 6y agoThere should be some form of compilation of quality articles per domain, like history, sciences, etc. In a way all articles should belong to a category...
- boramalper 6y agoKiwix provides that. :)
- jokoon 6y agoNo article are properly labelled by domain of science. There are collections of article, but for example there are no collections of all articles related to computer science or some part of Swedish history.
- boramalper 6y agoI see your point: what you are asking for is labelling the articles rather than grouping them in collections, right? I think labelling is useful when you can download individual articles (by label) efficiently. I don’t think it’s practical to create torrents for each nor efficient to ask users to scrape. IPFS could have helped if it wasn’t so slow in practise. Also see https://wiki.kiwix.org/wiki/Content_in_all_languages https://wiki.kiwix.org/wiki/Content_in_all_languages
- jokoon 6y agoI guess wikipedia should encourage all article editor to properly tag articles. Dewey categories are a good way to label things, I guess.