5 ms·
We built a new semantic database first in university and then commercial open source (TerminusDB). We use the web ontology language (OWL) as a schema language,
by LukeEF 6y ago
We built a new semantic database first in university and then commercial open source (TerminusDB). We use the web ontology language (OWL) as a schema language, but made two important - practical - modifications: 1) we dispense with the open world interpretation; and 2) insist on the unique name assumption. This provides us with a rich modelling language which delivers constraints on the shapes in the graph. Additionally, we don't use SPARQL, which we didn't find practical (composability is important to us) and use a Datalog in its place (like Dataomic and others).
Our feeling on interacting with the semantic web community is that innovation - especially when it conflicts with core ideology - is not welcome. We understand that 'open world' is crucial to the idea of a complete 'semantic web', but it is insanely impractical for data practitioners (we want to know what is in our DB!). Semantic web folk can treat alternative approaches as heresy and that is not a good basis for growth.
As we came from university, I agree with comments that the field is too academic and bends to the strange incentives of paper publishing. Lots of big ideas and everything else is mere 'implementation detail' - when, in truth, the innovation is in the implementation details.
There are great ideas in the semantic web, and they should be more widespread. Data engineers, data scientists, and everybody else can benefit, but we must extract the good and remove ideological barriers to participation.
- tannhaeuser 6y agoYou're right to emancipate from the grab that SemWeb has had on the field for so long and turn to Prolog/Datalog and practical approaches IMO. Open world semantics and sophisticated theories may have been a vision for the semantic web of heterogenous data, but in reality RDF and co are only used in certain closed-world niches IME. Pascal Hitzler is one of the more prolific authors (especially with the EU-funded identification of description logic fragments of OWL2 which are some of the better results in the field IMO), but beginning this whole discussion with W3C's RDF is wrong IMO when description logic as more or less variable-free fragments of first-order logic with desirable complexities was a thing in 1991 or earlier already. Nit: careful with datomic. It's clearly not Datalog, but an ad-hoc syntax whereas Datalog is a proper syntactic subset of Prolog. And while I don't like SPARQL, it still gives quite good compat for querying large graph databases.
- j-pb 6y agoNitNit: I think the term "Datalog" the prolog subset, has been pretty much replaced with "Datalog" the recursive consjunctive query fragment with recursion (and sometimes stratified negation) term. Most papers and textbooks I read these days use it as a complexity class for queries and not as a concrete syntax.
- LukeEF 6y agoThis is the sense in which I was using Datalog - and how others like Datomic, Grakn and Crux use it (there is a growing movement of databases with a 'Datalog' query language) - althou in our case, we can also use in the former sense as TerminusDB is implemented in prolog.
- JPKab 6y agoI met Pascal Hitzler on a few occasions, not long after he ended up relocating from Germany to Ohio to work at Wright State. He was a rare bright spot in always trying to bridge the gap between theory and application in the SemWeb community. He was kind enough to meet up with me and a colleague at a coffee shop in Dayton to discuss our project, all pro-bono. A real good dude.
- Communitivity 6y agoAlmost every pragmatic implementation of semantic reasoning I've done involved both of the same modifications (closed world and unique names). A couple efforts used SPARQLX, something I created that was a binary form of SPARQL+SPARQLUpdate+StoredProcedures+Macros encoded using Variable Message Format. This was about 18 years ago, before SPARQL and SPARQL update merged, and before FLWOR. One of these days I'll recreate it again. The original work is not available, and I was not allowed to publish. Oh, and I forgot two things, SPARQLX had triggers, was customized for OWL DLP, and had commands for custom import and export using N3 (I was a big fan of the cwm software).
- wuschel 6y ago> [...] but we must extract the good and remove ideological barriers to participation. Could you point to some resources that explain the tradeoff between the practical solutions and concepts and the ideologic cruft for an outsider?
- lou1306 6y agoNot the commenter, but I hope to add something to the discussion. Generally, expanding on the current state of the art is paramount in academia. In this case, I guess that defaulting on closed-world and unique names is frowned upon because academic people "know" that SemWeb concepts would be "easy" to implement under such conditions (for some interpretation of "know" and "easy"). A university lab would be reluctant to invest on such a project, because it would likely result in less publications than, say, a bleeding-edge POC. Of course, practical solutions based on well-understood assumptions are exactly what a commercial operation needs, so it's no wonder that TerminusDB chose that path. They might not publish a ton of papers, but they have something that works and could be used in production.
- wuschel 6y agoVery interesting. Thanks!
- jerf 6y ago"Our feeling on interacting with the semantic web community is that innovation - especially when it conflicts with core ideology - is not welcome." I wasn't a big fan of the "semantic web" community when it first came out, and the years have only deepened my disrespect, if not outright contempt. The entire argument was "Semantic web will do this and that and the other thing!" "OK, how exactly will it accomplish this?" "It would be really cool if it did! Think about what it would enable!" "OK, fine, but how will this actually work!" "Graph structures! RDF!" "Yes, that's a data format. What about the algorithms? How are you going to solve the core problem, which is that nobody can agree on what ontology to apply to data at global scale, and there isn't even a hint of how to solve this problem?" "So many questions. You must be a bad developer! It would be so cool if this worked, so it'll work!" There has always been this vacuousness in the claims, where they've got a somewhat clear idea of where they want to go, but if you ever try to poke down even one layer deeper into how it's going to be solved, you get either A: insulted B1: claims that it's already solved just go use this solution (even though it is clearly not already solved since the semantic web promises are still promises and not manifested reality) B2: claims it's already solved and the semantic web is already huge (even though the only examples some using this can cite are trivial compared to the grand promises and the "semantic web" components borderline irrelevant, most frequently citing "those google boxes that pop up for sites in search results" just like this article does despite the fact they're wafer-thin compared to the Semantic Web promises and barely use any "Semantic Web" tech at all) or C: a simple reiteration of the top-level promises, almost as if the person making this response simply doesn't fundamentally grasp that the ideals need to manifest in real code and real data to work. This article does nothing to dispel my beliefs about it. The second sentence says it all. For the rest, while just zooming in to the reality may be momentarily impressive, compared to the promises made it is nothing. The whole thing was structured backwards anyhow. I'd analogize the "semantic web" effort to creating a programming language syntax definition, but failing to create the compiler, the runtime, the standard library, or the community. Sure, it's non-trivial forward progress, but it wasn't really the hard part. The real problem for semantic web and their community is the shared ontology; solve that and the rest would mostly fall into place. The problem is... that's an unsolvable problem. Unsurprisingly, a community and tech all centered around an unsolvable problem haven't been that productive. A fun exercise (which I seriously recommend if you think this is solvable, let alone easy) is to just consider how to label a work with its author. Or its primary author and secondary authors... or the author, and the subsequent author of the second edition... or, what exactly is an authored work anyhow? And how exactly do we identify an author... consider two people with identical names/titles, for instance. If we have a "primary author" field, do we always have to declare a primary author? If it's optional, how often can you expect a non-expert bulk adding author information in to get it correct? (How would such a person necessarily even know how to pick the "primary author" out of four alphabetically-ordered citations on a paper?) (I am aware of the fact there are various official solutions to these problems in various domains... the fact that there are various solutions is exactly my point. Even this simple issue is not agreed upon, context-dependent, it's AI-complete to translate between the various schema, and if you speak to an expert using any of them you could get an earful about their deficiencies.)
- usrusr 6y ago> "dispense with the open world interpretation“ That can mean anything from "we have some conventional (e.g. plain old RDBMS) CWA systems but describe their schemas in an OWA DL to ease integration across independent systems" (in particular this means no CWA implications outside those built into the conventional systems with or without a semweb layer on top) to "we do a big bucket of RDF and run it all through a set of rules formulated in OWL syntax but applied in an entirely different way" (CWA everywhere). The former would be semweb as intended, or at least a subset thereof, but the latter could easily end up somewhere between simple brand abuse and almost comical cargo culting. Well, at least that's how I feel as someone who never had to face the realities of the vast unmapped territories between plain old database applications and fascinating yet entirely impractical academic mind games of DL (old school symbolic AI ivory tower that suddenly happened to find itself in the center of the hottest w3c spec right before w3c specs kind of stopped being a thing, with WHATWG usurping html and Crockford almost accidentally killing XML) (also, when has "assumption" turned into "interpretation"? Guess I missed a lot)
- nut-hatch 6y agoI completed my PhD in the scope of Semantic Web technologies and I can share the same experience that the semantic web community is extremely closed (coming across as feeling "elite"). Having myself no supervisor from the field, it was still possible to publish my ideas (ISWC, WWW etc), but it was impossible to connect to the people and be taken seriously. I moved on from that field now, and I don't expect to come in touch with any Semantic Web stuff in a open-world context any time soon. I couldn't agree more with you that the strong ideology that drives this community is one of the main reason that these technologies are not widely adopted. This, and the failure to convince people outside academia that solving the problems it tries to solve is necessary in the first place. Good luck with TerminusDB, I think I listened to you at KGC.
- JPKab 6y agoI was "stuck" working with a bunch of leading academics and researchers on a SemWeb project using OWL/RDF, in collaboration with DARPA and the US Department of Defense, around 2008-2009. You are absolutely correct that they are hostile to anything outside of their "ideology". The awful, horrific performance of the RDF/OWL databases compared to the impure, heretical evil Neo4j that they despised for its practicality..... that was always funny. Another interesting thing I encountered in the field was the side-effect of an academic field that sounds really good to people who have never built anything EVER is that it can often get a ton of funding and grant money from central government organizations, thereby creating legions of rather shitty companies (especially in the EU, where the grants were everywhere) that have the word "semantic" in the name even though they do nothing with actual semantic technology. These shitty companies are often just there to employ the academics. The craziest thing was these projects where they would employ a dozen or so "library scientists" who were all just masters and phd students who had, for some reason, decided to study to be librarians in the digital era to create the OWL ontologies. None of them knew anything about computer science or programming, and they would just sit there and read thousands of policy documents and use an Eclipse based GUI to create and edit giant graphs of knowledge and rules. All were being paid six figures, and didn't produce a single goddammed thing of value. So much taxpayer money in those rooms going to complete waste. Glad it wasn't just me that thought the community was a joke. The semantic web will arrive one day, but OWL and SPARQL won't be anywhere in it. And it won't be any of these academics delivering it.
- ta988 6y agoAre you referring to the BFO crowd? I would be interested in your thoughts about BFO itself if thats the case.
- JPKab 6y agoIf by BFO you are referring to "Basic Formal Ontologies", then yes, I'm referring to that "crowd". My thoughts on ontologies in general are that they can certainly be powerful, and I've seen them used in the past in rules engines that powered fraud detection applications. In the SemWeb community, in the late 2000s, they successfully convinced a bunch of CIOs of massive organizations, especially in the US Federal Government, that a key to being able to centralize and federate all of their data, and save money on duplicative systems, they could simply have semantic mappings on top of every IT system's databases, and query this semantic layer. Ideally, they could eliminate duplicated data, so that all systems would get data from the "Authoritative Data Source" system instead of duplicating it locally in the application's database. I'm sure you can immediately see why this is wildly stupid and unrealistic. Imagine what it would look like if every single piece of data that I can technically obtain from another source has to remain in that source, and that storing that data locally with my application specific data is forbidden..... Suddenly, there is a massive increase in I/O, drop in performance, etc. The whole project taught me a lesson about the politics of academia, and how there is a segment of the population that is highly educated, and has learned how to manufacture work for themselves outside of academia by pushing for high-level government officials to implement programs based on their theories..... MITRE was a big part of this particular project.