5 ms·
"That gives you Wikitext encapsulated in XML." avar: "The goal of Wikipedia should be to spread the content as far & wide as possible, the way OpenStreetMap op
by 10165 9y ago
"That gives you Wikitext encapsulated in XML."
avar: "The goal of Wikipedia should be to spread the content as far & wide as possible, the way OpenStreetMap operates is a better model."
I am confused.
Doesn't OSM data come encapsulated in XML or some binary format?
As for dispersion of content, I could have sworn I have seen Wikipedia content on non-Wikipedia websites. Is there some restriction that prohibits this?
I have seen Wikipedia data offered in DNS TXT records as well.
- 3131s 9y agoFor each article there is some metadata, but the entire text of an article is just a blob inside one XML element. For anyone who has not worked with the Wikipedia data dumps extensively before, trust us that it is not easily machine-readable and that even solutions like DBPedia / Wikidata are not yet suitable for many purposes.
- WikipediasBad 9y agoAs someone who contributes to many knowledge projects, including Wikipedia and Wikidata frequently, I'm curious about what you mean that Wikidata is not yet suitable for any purposes. Am I wasting my time contributing to it? I thought that it was helping a lot of machines understand data. Can you please explain further?
- 3131s 9y agoPlease reread, for many purposes! I love Wikipedia. The Wiki markup is extremely complicated and being user created, it is also inconsistent and error prone. I believe the MediaWiki parser itself is something like a single 5000 line PHP function! All of the alternate parsers I've tried are not perfect. There is a ton of information encoded in the semi-structured markup, but it's still not easy to turn that into actual structured data. That's where the problem lies.
- lacksconfidence 9y ago> believe the MediaWiki parser itself is something like a single 5000 line PHP function! It's not. I'm on mobile so not easiest to link, but the PHP versio of the parser is nothing like a single function. There is also a nodejs version of the parser under active development with the goal of replacing the php parser.
- 3131s 9y agoThanks, I had heard that somewhere but stand corrected.
- 10165 9y ago"... into actual structured data." Would there be some particular structure that everyone would agree on? Alternatively, what is the desired structure you want? Because the current format is so messy, I just focus on what I believe is most important: titles and externallinks. IMO, often the most interesting content in an article is lifted from content found via the external links. I also would like to capture the talk pages. Maybe just the contributing usernames and IP addresses. Opinions or explanations that have no supporting reference are inexpensive. One can always these for free on the web. No problem recruiting "contributors" for that sort of "content". Back to the question: I am curious what structure would you envision would be best for Wikipedia data? Assume hypothetically that a "perfect" parser has been written for you to do the transformation.
- rspeer 9y agoThe structure I need for my particular project (ConceptNet) is: * The definitions from each Wiktionary entry. * The links between those definitions, whether they are explicit templated links in a Translation or Etymology section, or vaguer links such as words in double-brackets in the definition. (These links carry a lot of information, and they're why I started my own parser instead of using DBnary.) * The relations being conveyed by those links. (Synonyms? Hypernyms? Part of the definition of the word?) * The links should clarify the language of the word they are linking to. (This takes some heuristics and some guessing so far, because Wiktionary pages define every word in any language that uses the same string of letters, and often the link target doesn't specify the language.) * The languages involved should be identified by BCP 47 language codes, not by their names, because names are ambiguous. (Every Wiktionary but the English one is good at this.) There are probably analogous relations to be extracted from Wikipedia, but it seems like an even bigger task whose fruit is higher-hanging. Don't get me wrong: Wiktionary is an amazing, world-changing source of multilingual knowledge. Wiktionary plus Games With A Purpose are most of the reason why ConceptNet works so well and is mopping the floor with word2vec. And that's why I'm so desperate to get at what the knowledge is.
- rspeer 9y agoThe GP said Wikidata isn't suitable for many purposes, different from any. It's a nice agreed-upon vocabulary for linked data. But you still need the data that the vocabulary refers to. The information you can get without ever leaving the Wikidata representation is still too sparse.
- StavrosK 9y agoHe's saying that Wikipedia doesn't give you clean, usable data, it gives you data with weird markup everywhere.