4 ms·
Wikipedia data dumps and stats
- fauigerzigerk 14y agoSadly, they don't publish up-to-date HTML dumps and there is no reliable way of reproducing them short of installing the entire wikipedia system locally, including the database. I know there are quite a few projects that claim to do it but they're all abandoned, incomplete or unsuitable in various other ways (as far as I know).
- atdt 14y ago"The entire Wikipedia system, including the database" is slightly overblown: it's just MediaWiki and MySQL.
- yuvipanda 14y agoMediaWiki (+ a finely tuned PHP system, because mediawiki is unusably slow without it), MySQL, and plenty of time / resources for doing the large import of enwiki (since that is what most people are interested in).
- fauigerzigerk 14y agoExactly, and at that point you still don't have the static HTML files. You have to crawl the entire local site, which takes ages. Then you have to repeat all of this according to your desired update frequency.
- alexkus 14y agoEC2 + S3 Minimal charge to cover the bandwidth and hosting costs. Any profit donate to Wikipedia Foundation. Solves that same problem for lots of people.
- linguaz 14y agoNot sure if this helps, but there's the Kiwix offline Wikipedia reader and its associated Zim files: http://www.kiwix.org/ http://www.kiwix.org/ http://download.kiwix.org/zim/0.9/ http://download.kiwix.org/zim/0.9/ http://www.kiwix.org/wiki/Tools/en#Generating_ZIM_Files_From_Wikis http://www.kiwix.org/wiki/Tools/en#Generating_ZIM_Files_From... http://www.openzim.org/ http://www.openzim.org/
- wikiburner 14y agoHey everybody, fauigerzigerk sort of gets into this, but I just downloaded the dump yesterday expecting there to be a relatively straightforward way to parse and search it with Python and extract and process articles of interest w/ NLTK. I'm not sure what I was expecting exactly, but it sure wasn't a single 40gb XML file that I can't even open in Notepad++. Is my only real option (for parsing and data mining this thing) to basically set up a clone of wikipedia's system, and then screen scrape localhost?
- fauigerzigerk 14y agoIt's not your only option. You can open the XML dump with a streaming XML parser (not a DOM parser) and use one of the existing wiki syntax parsers to extract what you need. If you just need a few specific items (for instance just the links to reconstruct the page graph or just the info boxes) that's a perfectly workable solution. There is a large number of small tools and scripts that extract various bits and pieces from the XML dump. You may well find a tool that suits your needs. But there are two issues: The available parsers are not very robust and not very complete, because the wiki syntax is extremely convoluted and there is no formal spec. Second, the wiki syntax includes a kind of macro system. Without actually executing those macros you don't get the complete page as you see it online. The only way to get the complete and correct page content, to my knowledge, is to install the mediawiki site and import the data. If you just want to look at the XML dump quickly you can use less or tail.
- wikiburner 14y agoHi fauigerzigerk, thanks for your response. Unfortunately, I'm going to need pretty much the entire page structure, content, links/graph of each article, so it looks like I might have to go the MediaWiki route. I'm noticing now that the dump is available as SQL as well, so maybe I'll check that out as well and see if a more streamlined approach is possible that way. Also, regarding less or tail, unfortunately I'm on Windows. Anyway, thanks again for your help, I appreciate it.
- deleted 14y ago[deleted]