3 ms·
I also didn't find much information about how long it would take to import into a db, so I used the xml dumps directly [1]. I only needed the wiki content (not
by thomas536 8y ago
I also didn't find much information about how long it would take to import into a db, so I used the xml dumps directly [1]. I only needed the wiki content (not the history), so the article xml files worked well for me. And then I used mwparserfromhell [2] to parse and extract from the wiki markup.
[1] https://dumps.wikimedia.org/enwiki/20190301/ https://dumps.wikimedia.org/enwiki/20190301/
[2] https://mwparserfromhell.readthedocs.io/en/latest/ https://mwparserfromhell.readthedocs.io/en/latest/
- paulmolloy 8y agoI've been working on some research for a recommender system using xml wiki article dumps the last few months. I've been using mwparserfromhell as well to get plain text and some other metadata I needed from articles to create a dataset. It seems to work pretty well for that use case anyway.