3 ms·
Wikidata statements (which roughly correspond to the edges in the Knowledge Graph) have quite a bit of Metadata associated with them: they can have refer to sou
by mmarx 5y ago
Wikidata statements (which roughly correspond to the edges in the Knowledge Graph) have quite a bit of Metadata associated with them: they can have refer to sources that state this particular bit of knowledge, they have a so-called rank that allows distinguishing preferred and deprecated statements, and the can be qualified by another statement in the graph. Temporal validity is encoded using a combination of rank and qualifiers, as for, e.g., Pluto[0], where the instance-of statement saying that “Pluto is a planet” is deprecated and has an “end time” qualifier, and the preferred statement says “Pluto is a dwarf planet,” with a corresponding “start time” qualifier.
In principle, all of this information is available through the SPARQL endpoint or as an RDF export (there is also the simplified export that contains only “simple” statements lacking all of that metadata), so reasoning over this data is not entirely out of reach, but the sheer size (the full RDF dump is a few hundred GBs) is also not particularly practical to deal with.
[0] https://www.wikidata.org/wiki/Q339#P31 https://www.wikidata.org/wiki/Q339#P31
- pcrh 5y ago>Wikidata https://en.wikipedia.org/wiki/Wikidata https://en.wikipedia.org/wiki/Wikidata Thanks for that! TIL. It seems a fascinating project in epistemiology!
- ablekh 5y agoThe size of Wikidata knowledge base / relevant graph (as well as Linked Open Data Cloud KBs and other large KBs) certainly presents some challenges. However, I think that the largest challenge and, in fact, the main obstacle, for practical programmatic solutions is the use of essentially meaningless alphanumeric identifiers assigned to entities and properties. All corresponding identifiers need to be discovered first manually in order to construct relevant SPARQL queries. Needless to say that these queries are not particularly human-readable (or, rather, human-interpretable) as well.
- tuukkah 5y agoWhy manually, when you have APIs to find Wikidata items and properties based on their labels, descriptions, aliases, data, metadata and use? If you mean autocomplete UI or tooltips, look no further than the query editor and its Ctrl+Space at https://query.wikidata.org/ https://query.wikidata.org/
- ablekh 5y agoI meant APIs, not UI or tooltips. And while Wikidata entities and properties could be accessed using MediaWiki API, arguably, there are, at least, two issues with this: 1) you have to know exact names of all the relevant metadata, which is quite overwhelming* (and the SPARQL query editor's autocomplete feature does not seem to help with this, except for top-level attributes); 2) entity disambiguation - yes, it can be implemented programmatically, however, it has to rely on knowing exact names (values), which brings us back to the point #1. *) Here is an example of the number of attributes for a single entity: https://www.wikidata.org/w/api.php?action=wbgetentities&ids=Q42&languages=en https://www.wikidata.org/w/api.php?action=wbgetentities&ids=....