6 ms·
> CURIEs and the depending standards alone are well over 100 pages. The curie standard is 10 pages long, and those "dependent standards" includes things like R
by cheph 6y ago
> CURIEs and the depending standards alone are well over 100 pages.
The curie standard is 10 pages long, and those "dependent standards" includes things like RFC 3986 (Uniform Resource Identifiers (URI): Generic Syntax) and RFC 3987 (Internationalized Resource Identifiers (IRI)) - which are well established technologies that most people should be familiar with. And you really don't need to read all of the referenced standards to be able to understand and use CURIE quite proficiently.
> RDF has like 100
Normative specifications of RDF is contained in two documents:
- RDF 1.1 Concepts and Abstract Syntax ( https://www.w3.org/TR/rdf11-concepts/ https://www.w3.org/TR/rdf11-concepts/ ) = 20 pages
- RDF 1.1 Semantics ( https://www.w3.org/TR/rdf11-mt/ https://www.w3.org/TR/rdf11-mt/ ) = 29 pages
These page counts includes TOC, reference sections, appendices and large swathes of non-normative content also.
And really the RDF 1.1 primer (https://www.w3.org/TR/rdf11-primer/ https://www.w3.org/TR/rdf11-primer/) should be quite sufficient for most people who want to use it, and that is only 14 pages.
RDF and CURIE is simple as dirt really, maybe too simple, but I think I can explain it quite well to someone with some basic background in IT in about 30 minutes.
And while the other aspects (e.g. SPARQL, OWL) are not that simple, there is inherent complexity they are trying to address that you cannot just ignore. And not everybody needs to know OWL, and SPARQL is really not that complicated either and again most people can become quite proficient with this rather quickly if they understand the basics.
> What we need is a simpler ecosystem, where people can stake their claim on their niche, where they have the ability and power to experiment and explore.
What are the alternatives? Proliferation of JSON schemas which is yet to be ratified as a standard and does not address most of the same problems as Semantic Web Technology? I think there are some validity to your concerns, but semantic web technologies are being used widely in production, maybe not all of them, but to suggest it is not usable is not true.
I have used RDF in Java (rdf4j and jena), Python (rdflib) and JS (rdflib.js) without serious problems.
- j-pb 6y agoFamiliarity isn't nearly enough if you want to implement something. Talking about RDF is absolutely meaningless without talking about Serialisation (and that includes ...URGH.. XML serialisation), XML Schema data-types, localisations, skolemisation, and the ongoing blank-node war. The semantic web ecosystem is the prime example of "the devils in the detail". Of course you can explain to somebody who knows what a graph is, the general idea of RDF: "It's like a graph, but the edges are also reified as nodes." But that omits basically everything. It doesn't matter if SparQL is learnable or not, it matters if its implementable, let alone in a performant way. And thats really really questionable. Jena is okay-ish, but it's neither pleasant to use, nor bug free, although java has the best RDF libs generally (I think thats got something to do with academic selection bias). RDF4J has 300 open issues, but they also contain a lot of refactoring noise, which isn't a bad thing. C'mon, rdflib is a joke. It has a ridiculous 200 issues / 1 commit a month ratio, buggy as hell, and is for all intents and purposes abandonware. rdflib.js is in memory only, so nothing you could use in production for anything beyond simple stuff. Also there's essentially ZERO documentation. And none of those except for Jena even step into the realm of OWL. > What are the alternatives? Good question. SIMPLICITY! We have an RDF replacement running in production that's twice as fast, and 100 times simpler. Our implementation clocks in at 2.5kloc, and that includes everything from storage to queries, with zero dependencies. By having something that's so simple to implement, it's super easy to port it to various programming languages, experiment with implementations, and exterminate bugs. We don't have triples, we have tribles (binary triples, get it, nudge nudge, wink wink). 64 Byte in total, fits into exactly one cache line on the majority of Architectures. 16byte subject/entity | 16 byte predicate/attribute | 32 byte object/value These tribles are stored in knowledge bases with grow-set semantics, so you can only ever append (on a meta level knowledge bases do support non-monotonic set operations), which is the only way you can get consistency with open world-semantics, which is something that the OWL people apparently forgot to tell pretty much everybody who wrote RDF stores, as they all have some form of non-mononic delete operation. Even SparQL is non-monotonic with it's optional operator... Having a fixed size binary representation makes this compatible with most existing databases, and almost trivial to implement covering indices and multiway joins for. By choosing UUIDs (or ULIDs, or TimeFlakes, or whatever, the 16byte don't care) for subject and predicate we completely circumnavigate the issues of naming, and schema evolution. I've seen so many hours wasted by ontologists arguing about what something should be called. In our case, it doesn't matter, both consumers of the schema can choose their own name in their code. And if you want to upgrade your schema, simply create a new attribute id, and change the name in your code to point to it instead. If a value is larger than 32 byte, we store a 256bit hash in the trible, and store the data itself in a a separate blob store (in our production case S3, but for tests it's the file stystem, we're eyeing a IPFS adapter but that's only useful if we open-sourced it). Which means that it's also working nicely with binary data, which RDF never managed to do well. (We use it to mix machine learning models with symbolic knowledge). We stole the context approach from jsonLD, so that you can define your own serialisers and deserialisers depending on the context they are used in. So you might have a "legacyTimestamp" attribute which returns a util.datetime, and a "timestamp" which returns a JodaTime Object. However unlinke jsonLD these are not static transformations on the graph, but done just in time through the interface that exposes the graph. We have two interfaces. One based on conjunctive queries which looks like this (JS as an example): ``` // define a schema const knightsCtx = ctx({ ns: { [id]: { ...types.uuid }, name: { id: nameId, ...types.shortstring }, loves: { id: lovesId }, lovedBy: { id: lovesId, isInverse: true }, titles: { id: titlesId, ...types.shortstring }, }, ids: { [nameId]: { isUnique: true }, [lovesId]: { isLink: true, isUnique: true }, [titlesId]: {}, }, }); // add some data const knightskb = memkb.with( knightsCtx, ( [romeo, juliet], ) => [ { [id]: romeo, name: "Romeo", titles: ["fool", "prince"], loves: juliet, }, { [id]: juliet, name: "Juliet", titles: ["the lady", "princess"], loves: romeo, }, ], ); // Query some data. const results = [ ...knightskb.find(knightsCtx, ( { name, title }, ) => [{ name: name.at(0).ascend().walk(), titles: [title] }]), ]; ``` and the other based on tree walking, where you get a proxy object that you can treat as any other object graph in your programming language, and you can just navigate it by traversing it's properties, lazily creating a tree unfolding. Our schema description is also heavily simplified. We only have property restrictions and no classes. For classes there's ALWAYS a counter example of something that intuitively is in that class, but which is excluded by the class definition. At the same time, classes are the source of pretty much all computational complexity. (Can't count if you don't have fingers.) We do have cardinality restrictions, but restrict the range of attributes to be limited to one type. That way you can statically type check queries and walks in statically typed languages. And remember, attributes are UUIDs and thus essentially free, simply create one attribute per type. In the above example you'll notice that queries are tree queries with variables. They're what's most common, and also what's compatible with the data-structures and tools available in most programming languages (except for maybe prolog). However we do support full conjunctive queries over triples, and it's what these queries get compiled to. We just don't want to step into the same impedance mismatch trap datalog steps into. Our query "engine" (much simpler, no optimiser for example), performs a lazy depth first walk over the variables and performs a multiway set intersection for each, which generalises the join of conjunctive queries, to arbitrary constraints (like, I want only attributes that also occur in this list). Because it's lazy you get limit queries for free. And because no intermediary query results are materialised, you can implement aggregates with a simple reduction of the result sequence. The "generic constraint resolution" approach to joins also gives us queries that can span multiple knowledge bases (without federation, but we're working on something like that based on differential dataflow). Multi-kb queries are especially useful since our default in-memory knowledge base is actually an immutable persistent data-structure, so it's trivial and cheap to work with many different variants at the same time. They efficiently support all set operations, so you can do functional logic programming a la "out of the tar pit", in pretty much any programming language. Another cool thing is that our on-disk storage format is really resilient through it's simplicity. Because the semantics are append only, we can store everything in a log file. Each transaction is prefixed with a hash of the transaction and followed by the tribles of the transaction, and because of their constant size, framing is trivial. We can loose arbitrary chunks of our database and still retain the data that was unaffected. Try that with your RDMBS, you will loose everything. It also makes merging multiple databases super easy (remember UUIDs to prevent naming collisions, monotonic open world semantics keep consistency, fixed size tribles make framing trivial), you simply `cat db1 db2 > outdb` them. Again, all of this in 2.5kloc with zero dependencies (we do have one on S3 in the S3 blob store adapter). Is this the way to go? I don't know, it serves us well. But the great thing about it is that there could be dozens of equally simple systems and standards, and we could actually see which approaches are best, from usage. The semantic web community is currently sitting on a pile of ivory, contemplating on how to best steer the titanics that are protege, and OWLAPI through the waters of computational complexity. Without anybody every stopping to ask if that's REALLY been the big problem all along. "I'd really love to use OWL and RDF, if only the algorithms were in a different complexity class!"
- cheph 6y ago> Talking about RDF is absolutely meaningless without talking about Serialisation (and that includes ...URGH.. XML serialisation), XML Schema data-types, localisations, skolemisation, and the ongoing blank-node war. Don't implement XML serialization. The simplest and most widely supported serialization is n-quads (https://www.w3.org/TR/n-quads/ https://www.w3.org/TR/n-quads/). 10 pages, again with exaples, toc, and lots of non-normative content. You don't need to handle every data type, and you can't even if you wanted to because data types are also not a fixed set. And whatever you need to know about skolemisation, localization, and blank-nodes is in the standards AFAIK. > C'mon, rdflib is a joke. It has a ridiculous 200 issues / 1 commit a month ratio, buggy as hell, and is for all intents and purposes abandonware. It works, not all functionality works perfectly but like I said I have used it and it worked just fine. > rdflib.js is in memory only, so nothing you could use in production for anything beyond simple stuff. Also there's essentially ZERO documentation. For processing RDF in browser it works pretty well, not sure what you expect but to me RDF support does not imply it should be a fully fledged tripple-store with disk backing. Also not really zero documentation: https://github.com/linkeddata/rdflib.js/#documentation https://github.com/linkeddata/rdflib.js/#documentation > > What are the alternatives? > SIMPLICITY! > But the great thing about it is that there could be dozens of equally simple systems and standards, and we could actually see which approaches are best, from usage. Okay, so you roll your own that fits your use case. Not much use to me and it is not a standard. Lets talk again when you standardize it. Otherwise do you mind giving an alternative that I can actually take off the shelf to at least the extent that I can with RDF? I am not going to roll my own standard, and if all the RDF data sets instead used their own standards instead of RDF it won't really improve anything. EDIT: If you compare support for RDF to JSON schema, things are really not that bad.
- j-pb 6y ago> Don't implement XML serialization. The simplest and most widely supported serialization is n-quads (https://www.w3.org/TR/n-quads/ https://www.w3.org/TR/n-quads/). 10 pages, again with exaples, toc, and lots of non-normative content. You omit the transitive hull that the n-quads standard drags along, as if implementing a deserializer somehow only involved a parser for the most top-level EBNF. Also, you're still tip-toeing around the wider ecosystem of OWL, SHACL, SPIN, SAIL and friends. The fact that RDF alone even allows for that much discussion is indicative of it's complexity. It's like a discussion about SVG and HTML that never goes beyond SGML. And you can't have your cake and eat it too. You either HAVE to implement XML-Syntax or you won't be able to load half of the worlds datasets, nor will you even be able to start working with OWL, because they do EVERYTHING with XML. You're still coming from a user perspective. RDF will go nowhere unless it finds a balance between usability and implementability. Currently I'd argue, it focuses on neither. JS is a bigger ecosystem than just the browser, if you want to import any real-world dataset (or persistence) you need disk backing. So anything that just goes poof on a power failure doesn't cut it. Sorry but "works pretty well", and 6 examples combined with an unannotated automatically extracted API, does not reach my bar for "production quality". It's that "works pretty well" state of the entire RDF ecosystem that I bemoan. It's enough to write a paper about it, it's not enough to trust the future of your company on. Or you know. Your life. Because the ONLY real world example of an OWL ontology ACTUALLY doing anything is ALWAYS Snowmed. Snowmed. Snowmed. Snowmed. [A joke we always told about theoreticians finding a new lower bound and inference engines winning competitions: "Can snowmed be used to diagnose a patient?" "Well it depends. It might not be able to tell you what you have, but it can tell you that your 'toe bone is connected to the foot bone' 5 million times a second!"] Imagine making the same argument for SQL, it'd be trivial to just point to a different library/db. And so far we've only talked about complexity inherent in the technology, and not about the complex and hostile tooling (a.k.a. protege) or even the absolut unmaintainable rats nests that big ontologies devolve to. Having a couple different competing standards would actually improve things quite a bit, because it would force them to remain simple enough that they can still somehow interoperate. It's a bit like YAGNI. If you have two simple standards it's trivial to make them compatible by writing a tool that translates one to the other, or even speaks both. If you have one humongous one, it's nigh impossible to have two compatible implementations, because they will diverge in some minute thing. See rich hickeys talk "simplicity matters", for an in-depth explanation on the difference between simple (few parts with potentially high overall complexity through intertwinement and parts taking multiple roles), and decomplected (consisting of independent parts with low overall system complexity). And regarding JSON Schema: I never advocated for JSON schema and the fact that you have to compare RDFs maturity to something that hasn't been released yet... You would expect a standard that work began on 25 YEARS ago to be a bit more mature in it's implementations. If it hasn't reached that after all this time, we have to ask the question, why is that? And my guess is that implementors see the standards _and_ their transitive hull and go TL;DR, and even if they try, they get overwhelmed by the sheer amount of stuff.