6 ms·
The reason why tools like Protégé have not been sufficiently developed is because of infighting in the academic ontology community in addition to the reasons li
by hyperion2010 6y ago
The reason why tools like Protégé have not been sufficiently developed is because of infighting in the academic ontology community in addition to the reasons listed by the author. It has set the whole community back at least 5 years.
- j-pb 6y agoI think that's a symptom, not the cause. The complexity of web standards in general smother it with it's own weight. The common web has enough raw financial and person backing to grind through that. The semantic web does not. CURIEs and the depending standards alone are well over 100 pages. Language tags alone has 90. RDF has like 100, Sparql has a combined of more than 300, and OWL has more than 500, even though it assumes that the reader is generally familiar with Description logics, so it's probably a couple thousand if you take the required academic literature into account. Nobody is going to read all of that, let alone build that. Especially not a bunch of academics who don't care about the implementation as long as it's good enough to get the next paper out the door. So everybody pools on these few projects, because they're the only thing that's kinda working. OWLAPI, Protege, ... uh that's it. Because everything else, is broken and unfinished. Here's a thought experiment, name one production ready RDF libray for every major programming language (C, Java, Python, Js), that doesn't have major, stale, unresolved issues in their issue tracker. It's all broken, and there is simply too much work required to fix things. It's only natural that people start to infight when there is only few hospitable oasis. What we need is a simpler ecosystem, where people can stake their claim on their niche, where they have the ability and power to experiment and explore.
- syats 6y agoI agree with this. It is common to hear "Partial SPARQL 1.1 support"... or "Partial OWL compatibility" or "A variant of SKOS is supported". While it is true that full ECMA6/HTTP2/IPv6/SQL is also rarely provided by implementations, this doesn't hinder their use in productive environments. I think it is rare to reach the parts of ECMAscript that aren't implemented, or the corners of SQL that Postgres/MariaDB don't support. In many of the "Semantic Web Stack", however, one quickly reaches a "not implemented" portion of the 500 page owl standard.
- cheph 6y ago> CURIEs and the depending standards alone are well over 100 pages. The curie standard is 10 pages long, and those "dependent standards" includes things like RFC 3986 (Uniform Resource Identifiers (URI): Generic Syntax) and RFC 3987 (Internationalized Resource Identifiers (IRI)) - which are well established technologies that most people should be familiar with. And you really don't need to read all of the referenced standards to be able to understand and use CURIE quite proficiently. > RDF has like 100 Normative specifications of RDF is contained in two documents: - RDF 1.1 Concepts and Abstract Syntax ( https://www.w3.org/TR/rdf11-concepts/ https://www.w3.org/TR/rdf11-concepts/ ) = 20 pages - RDF 1.1 Semantics ( https://www.w3.org/TR/rdf11-mt/ https://www.w3.org/TR/rdf11-mt/ ) = 29 pages These page counts includes TOC, reference sections, appendices and large swathes of non-normative content also. And really the RDF 1.1 primer (https://www.w3.org/TR/rdf11-primer/ https://www.w3.org/TR/rdf11-primer/) should be quite sufficient for most people who want to use it, and that is only 14 pages. RDF and CURIE is simple as dirt really, maybe too simple, but I think I can explain it quite well to someone with some basic background in IT in about 30 minutes. And while the other aspects (e.g. SPARQL, OWL) are not that simple, there is inherent complexity they are trying to address that you cannot just ignore. And not everybody needs to know OWL, and SPARQL is really not that complicated either and again most people can become quite proficient with this rather quickly if they understand the basics. > What we need is a simpler ecosystem, where people can stake their claim on their niche, where they have the ability and power to experiment and explore. What are the alternatives? Proliferation of JSON schemas which is yet to be ratified as a standard and does not address most of the same problems as Semantic Web Technology? I think there are some validity to your concerns, but semantic web technologies are being used widely in production, maybe not all of them, but to suggest it is not usable is not true. I have used RDF in Java (rdf4j and jena), Python (rdflib) and JS (rdflib.js) without serious problems.
- j-pb 6y agoFamiliarity isn't nearly enough if you want to implement something. Talking about RDF is absolutely meaningless without talking about Serialisation (and that includes ...URGH.. XML serialisation), XML Schema data-types, localisations, skolemisation, and the ongoing blank-node war. The semantic web ecosystem is the prime example of "the devils in the detail". Of course you can explain to somebody who knows what a graph is, the general idea of RDF: "It's like a graph, but the edges are also reified as nodes." But that omits basically everything. It doesn't matter if SparQL is learnable or not, it matters if its implementable, let alone in a performant way. And thats really really questionable. Jena is okay-ish, but it's neither pleasant to use, nor bug free, although java has the best RDF libs generally (I think thats got something to do with academic selection bias). RDF4J has 300 open issues, but they also contain a lot of refactoring noise, which isn't a bad thing. C'mon, rdflib is a joke. It has a ridiculous 200 issues / 1 commit a month ratio, buggy as hell, and is for all intents and purposes abandonware. rdflib.js is in memory only, so nothing you could use in production for anything beyond simple stuff. Also there's essentially ZERO documentation. And none of those except for Jena even step into the realm of OWL. > What are the alternatives? Good question. SIMPLICITY! We have an RDF replacement running in production that's twice as fast, and 100 times simpler. Our implementation clocks in at 2.5kloc, and that includes everything from storage to queries, with zero dependencies. By having something that's so simple to implement, it's super easy to port it to various programming languages, experiment with implementations, and exterminate bugs. We don't have triples, we have tribles (binary triples, get it, nudge nudge, wink wink). 64 Byte in total, fits into exactly one cache line on the majority of Architectures. 16byte subject/entity | 16 byte predicate/attribute | 32 byte object/value These tribles are stored in knowledge bases with grow-set semantics, so you can only ever append (on a meta level knowledge bases do support non-monotonic set operations), which is the only way you can get consistency with open world-semantics, which is something that the OWL people apparently forgot to tell pretty much everybody who wrote RDF stores, as they all have some form of non-mononic delete operation. Even SparQL is non-monotonic with it's optional operator... Having a fixed size binary representation makes this compatible with most existing databases, and almost trivial to implement covering indices and multiway joins for. By choosing UUIDs (or ULIDs, or TimeFlakes, or whatever, the 16byte don't care) for subject and predicate we completely circumnavigate the issues of naming, and schema evolution. I've seen so many hours wasted by ontologists arguing about what something should be called. In our case, it doesn't matter, both consumers of the schema can choose their own name in their code. And if you want to upgrade your schema, simply create a new attribute id, and change the name in your code to point to it instead. If a value is larger than 32 byte, we store a 256bit hash in the trible, and store the data itself in a a separate blob store (in our production case S3, but for tests it's the file stystem, we're eyeing a IPFS adapter but that's only useful if we open-sourced it). Which means that it's also working nicely with binary data, which RDF never managed to do well. (We use it to mix machine learning models with symbolic knowledge). We stole the context approach from jsonLD, so that you can define your own serialisers and deserialisers depending on the context they are used in. So you might have a "legacyTimestamp" attribute which returns a util.datetime, and a "timestamp" which returns a JodaTime Object. However unlinke jsonLD these are not static transformations on the graph, but done just in time through the interface that exposes the graph. We have two interfaces. One based on conjunctive queries which looks like this (JS as an example): ``` // define a schema const knightsCtx = ctx({ ns: { [id]: { ...types.uuid }, name: { id: nameId, ...types.shortstring }, loves: { id: lovesId }, lovedBy: { id: lovesId, isInverse: true }, titles: { id: titlesId, ...types.shortstring }, }, ids: { [nameId]: { isUnique: true }, [lovesId]: { isLink: true, isUnique: true }, [titlesId]: {}, }, }); // add some data const knightskb = memkb.with( knightsCtx, ( [romeo, juliet], ) => [ { [id]: romeo, name: "Romeo", titles: ["fool", "prince"], loves: juliet, }, { [id]: juliet, name: "Juliet", titles: ["the lady", "princess"], loves: romeo, }, ], ); // Query some data. const results = [ ...knightskb.find(knightsCtx, ( { name, title }, ) => [{ name: name.at(0).ascend().walk(), titles: [title] }]), ]; ``` and the other based on tree walking, where you get a proxy object that you can treat as any other object graph in your programming language, and you can just navigate it by traversing it's properties, lazily creating a tree unfolding. Our schema description is also heavily simplified. We only have property restrictions and no classes. For classes there's ALWAYS a counter example of something that intuitively is in that class, but which is excluded by the class definition. At the same time, classes are the source of pretty much all computational complexity. (Can't count if you don't have fingers.) We do have cardinality restrictions, but restrict the range of attributes to be limited to one type. That way you can statically type check queries and walks in statically typed languages. And remember, attributes are UUIDs and thus essentially free, simply create one attribute per type. In the above example you'll notice that queries are tree queries with variables. They're what's most common, and also what's compatible with the data-structures and tools available in most programming languages (except for maybe prolog). However we do support full conjunctive queries over triples, and it's what these queries get compiled to. We just don't want to step into the same impedance mismatch trap datalog steps into. Our query "engine" (much simpler, no optimiser for example), performs a lazy depth first walk over the variables and performs a multiway set intersection for each, which generalises the join of conjunctive queries, to arbitrary constraints (like, I want only attributes that also occur in this list). Because it's lazy you get limit queries for free. And because no intermediary query results are materialised, you can implement aggregates with a simple reduction of the result sequence. The "generic constraint resolution" approach to joins also gives us queries that can span multiple knowledge bases (without federation, but we're working on something like that based on differential dataflow). Multi-kb queries are especially useful since our default in-memory knowledge base is actually an immutable persistent data-structure, so it's trivial and cheap to work with many different variants at the same time. They efficiently support all set operations, so you can do functional logic programming a la "out of the tar pit", in pretty much any programming language. Another cool thing is that our on-disk storage format is really resilient through it's simplicity. Because the semantics are append only, we can store everything in a log file. Each transaction is prefixed with a hash of the transaction and followed by the tribles of the transaction, and because of their constant size, framing is trivial. We can loose arbitrary chunks of our database and still retain the data that was unaffected. Try that with your RDMBS, you will loose everything. It also makes merging multiple databases super easy (remember UUIDs to prevent naming collisions, monotonic open world semantics keep consistency, fixed size tribles make framing trivial), you simply `cat db1 db2 > outdb` them. Again, all of this in 2.5kloc with zero dependencies (we do have one on S3 in the S3 blob store adapter). Is this the way to go? I don't know, it serves us well. But the great thing about it is that there could be dozens of equally simple systems and standards, and we could actually see which approaches are best, from usage. The semantic web community is currently sitting on a pile of ivory, contemplating on how to best steer the titanics that are protege, and OWLAPI through the waters of computational complexity. Without anybody every stopping to ask if that's REALLY been the big problem all along. "I'd really love to use OWL and RDF, if only the algorithms were in a different complexity class!"
- namedgraph 6y agoOWLAPI, Protege - that's it? RDF libraries broken? Dude what rock are you living under? What about Jena, RDF4J, rdflib, redland, dotNetRDF etc? Most of these libraries have been developed and tested for 20+ years and are active. See for yourself: https://github.com/semantalytics/awesome-semantic-web#programming https://github.com/semantalytics/awesome-semantic-web#progra... Why are you spreading FUD?
- j-pb 6y agoJena is everything but user friendly, it has a lot of weird edge cases, bugs, and a horrible API. RDF4J is okay for RDF, but completely ignores OWL. RDFLib is a bug ridden mess, have you ever used it, or checked their issue tracker, and commit history? With that amount of production breaking bugs that haven't been resolved for years, it might as well be unmaintained. Redland has last been updated years ago. Sure there's software that's finished, but with the complexity of RDF, OWL and friends, my hypothesis would be "it's dead Jim". I haven't used dotNetRDF, but it looks okay at first glance. So at least you can do semantic web on windows... YES! THEY'VE BEEN TESTED FOR 20+ YEARS AND ARE STILL HALVE BAKED, BUG RIDDEN, REASEARCH PROJECTS. MY POINT EXACTLY. These are all smart, dedicated people. The semantic web ecosystem is too complex to get right, no matter how many hours and $$$ you throw at it. I have nothing to gain from talking about the shortcomings of RDF and it's related standards, except maybe inspire people to come up with something better, and to save themselves some pain and suffering. The pot calling the cattle black. Your motives seem more questionable, with the whole "username matches the topic talked about", and has a Semantic Web consultancy business, shtick.
- namedgraph 6y agoResearch projects? Research is something coming out of academia, as you know. These are open-source projects with active developer communities. Jena has long been under Apache, RDF4J is now under Eclipse Foundation. Can you for once answer why large companies in the industry are using RDF/SPARQL as of today if it's so "dead"? Here's a list: http://sparql.club http://sparql.club
- 6y ago
- ta988 6y agoCould you give more details on that? How can it slow down an editor development?