3 ms·
SPARQL is nicer than SQL in a few cases: * Federation is better than ETL+Datawarehousing * Data is to be integrated on demand by the user A practical example o
by jerven 6y ago
SPARQL is nicer than SQL in a few cases:
* Federation is better than ETL+Datawarehousing
* Data is to be integrated on demand by the user
A practical example of mine is the UniProt database in the life sciences and the European Patent Office SPARQL endpoints.
These two datasets have some intersection of data. Combining these two in a classical datawarehouse with ETL pipelines would cost a few million in start up costs (Full data fidelity, small team 1 year work, optimisitic). The same with non federated RDF/SPARQL is 3 days work. This with federated querying is 2 minutes work.
SQL has a richer ecosystem with many more people confident in it's usage. More deployment options. etc.
Which is why often you will see tools like StarDog Virtual Graphs which will turn existing SQL DB's into SPARQL ones (via translating SPARQL->SQL) for in organization federated knowledge graphs. i.e. no to minimal ETL pipelines, direct querying on (standby, copies) of primary datasources.
In some domains the "business" analysts know SQL, even rarer but possible they know SPARQL. Letting them ask any query they can think of not bounded by what is one specific database can be extremely valuable. For organizations that can extract that value the lower market penetration of SPARQL is sad but not a real issue. This works in practice for SPARQL but not for SQL as what can be done with a user stated SERVICE clause in SPARQL requires a DBA to setup a foreign table connection in the SQL world.
Another example is when a few (N) hospitals need to exchange data. They have a few more relational databases (N+2) for this data in house for their patient groups. Upon project commencement they notice that all are SQL, but non are similar from vendor differences, but more problematic modelling issues.
Transforming their data to RDF is as complicated as to standardize on one new schema. But the RDF gives immediate integration with SnoMedCT, ICD10 and LOINC which allows easy queries taking into account hierarchical knowledge in those medical standards. Those queries would be possible two write by hand in SQL but are easier in RDF/SPARQL when attaching a minimal OWL reasoner. Then integrating with GeoSpatial data is again easier because in this country that is provided in RDF as well.
NOTE: A full fidelity UniProt Schema (including all subdata sets e.g. UniParc) would be a 150-200 tables depending on some modelling choices. EPO I assume to be in the same order.
NOTE2: While federated querying is a standard SPARQL feature this can of course be limited/turned off depending on the security/legal context.