5 ms·
It’s so sad that almost nobody knows or uses SPARQL…
by brodo 5y ago
It’s so sad that almost nobody knows or uses SPARQL…
- devbas 5y agoBecause the syntax is relatively complex and it is difficult to judge which endpoints and definitions to use.
- ZeroGravitas 5y agoI learned SPARQL recently, and would agrre its complicated to get info out of Wikidata. However, having read the article, they didnt have an easy time with scraping Wikipedia either. So I'd probably still recommend people look into wikidata and SPARQL if they want to do this kind of thing. Theres a few tools that generate queries for you, and some cli tools as well: https://github.com/maxlath/wikibase-cli#readme https://github.com/maxlath/wikibase-cli#readme It makes Wikipedia better too, in a virtuous cycle, with some infoboxes like those that he scraped being converted to be automatically populated from wikidata.
- matkoniecz 5y agoIn my experience SPARQL is really hard to use, and Wikidata data quality is really low. To the point that one of my larger project is trying to filter data to make it usable for my usecase. Yes, I made some improvements ( https://www.wikidata.org/wiki/Special:Contributions/Mateusz_Konieczny https://www.wikidata.org/wiki/Special:Contributions/Mateusz_... ). But overall I would not encourage using it, if I would know how much work it takes to get usable data I would not bother with it. Queries as simple as "is this entry describing event, bridge or neither" are requiring extreme effort to get right in a reliable way, including maintaining private list of patches and exemptions. And bots creating millions of known duplicated entries and expecting people to resolve this manually is quite discouraging. Creating Wikidata entries for Cebuano Wikipedia 'articles' was accepted, despite that Cebuano botpedia is nearly completely bot-generated. And that is without unclear legal status. Yes, they can legally import databases covered by database rights - but they should either make clear that Wikidata is a legal quagmire in EU or forbid such imports. But Wikidata community did neither.
- namedgraph 5y agoWho knew that a global machine-readable knowledge base would involve some complexity?
- matkoniecz 5y ago"is this entry describing (a) bridge (b) event" should have some reasonable way to answer. So far I have not found way to achieve this without laboriously maintaining my own database of errata, and new exceptions keep appearing.
- pphysch 5y agoWhy should more people know SPARQL?
- namedgraph 5y agoE.g. to get a job in FAANG, finance or pharma where SPARQL is used extensively on enterprise Knowledge Graphs? Check the jobs here: http://sparql.club/ http://sparql.club/
- pphysch 5y agoThe few openings I flipped through all mention SPARQL in an offhand manner, in the sense of "familiarity with query languages and data ontology".
- guidovranken 5y agoAt least the last times I checked, the WikiData SPARQL server was extremely slow, frequently timing out.
- lacksconfidence 5y agoseems to depend on the query. I can issue straight forward queries that visit a few hundred thousand triples easily. But when i write a query that visits tens of millions of triples it times out.
- epaulson 5y agoThere's some mix between "it's slow" and "it sets its timeout threshold too low" - a lot of queries would be OK if they just had a bit more time to run. And unfortunately, the time wasted on the badput of the killed queries just slows down everyone else. (They really need a batch queue) The Wikidata folks are well aware of the limits on their SPARQL service. They just posted an update the other day: https://lists.wikimedia.org/hyperkitty/list/wikidata@lists.wikimedia.org/thread/MSMKYTTWRZDD52JQLZCWPN4RSUCLFFMZ/ https://lists.wikimedia.org/hyperkitty/list/wikidata@lists.w...