8 ms·
The semantic web is like the guy that tells everyone that he is an asshole, and then surprise it turns out he is an asshole. The best quote so far: How can we e
by wrnr 4y ago
The semantic web is like the guy that tells everyone that he is an asshole, and then surprise it turns out he is an asshole. The best quote so far: How can we explain this discrepancy from a
mathematical perspective (thereby patently ignoring strategic, economic, social,
and other aspects that play a role).
- j-pb 4y agoThe paper is pretty tone deaf and arrogant, but in line with the culture that I experienced in my short time at Vrije Universiteit Brussel. It feels like the semantic community has a fault line between the realists and the theorists. With technologies like JSON-LD and SHACL done by the realists under the linked data banner. And OWL and Description logics done by the theorists under the old semantic web label. Tim Berners-Lee is really hurting the progress of semantic data processing and representation by opposing any web technology that could be an alternative to RDF unless it's folded under the RDF banner.
- zcw100 4y agoThat's hilarious that they're so out of touch that the JSON-LD and SHACL people are the realists. JSON-LD is ridiculous and a desperate attempt to attach semantic web technology to something that is actually used. It was announced with the misleading post titled "JSON-LD and Why I hate the Semantic Web". Only it's completely about the semantic web and is just a serialization format for RDF expressed in JSON. So you take JSON a format that can be described in a single page and layer this monstrosity over it. For what? The best thing that can be said for it is you can ignore it (maybe) and just treat it as JSON. There was zero need for JSON-LD. They had a perfectly good serialization format in TURTLE. It was similarly easy to describe and understand as JSON, but nooooo. They're always riding the coat tails of some other popular technology trying to get a free ride. It's like that obnoxious kid who shows up to the party and tells you how much smarter they are than you and then complains that no one will talk to them. SHACL is almost dumber that JSON-LD, if you can believe that. At least JSON-LD is just a serialization format. If you can manage to get it to parse you can just reserialize it to a sane format like TURTLE and get rid of the stupid. With SHACL you're stuck with it. So you go and create the worlds slowest database by basically normalizing the hell out of it because if a little normalization is good a lot of normalization is better. Screw knowing anything about the world. Let's throw that all out the window and allow people to express anything. So now you've got a database that can express anything and people say, "hey, I actually know some things about the world that seem to hold and my database is becoming a ridiculous mess. Can we make it so that people can't express something like I already can with my relational database?" Well the semantic web people went off in deep thought for a decade and finally came back with SHACL. It's a constraints language, expressed in RDF, of course, and if you express it in JSON-LD that means you've got SHACL expressed in RDF expressed in JSON, joy. I guess you could implement is in a number of different ways but it ends up firing off a series of queries saying, "Is this ok? What about that? How about this?" and if they're all successful then it will allow you to run the query you actually wanted to run. So now after all that you've got the world's slowest database that is now orders of magnitude even slower so that it operates a bit like MySQL. Semantic web databases allow you to express just about any query you'd like but it allows you to express queries that you're never going to, and is extremely slow for the ones you are.
- PaulHoule 4y agoThat’s a backwards way to think about it. JSON-LD lets you paint some extra structure onto a JSON document by adding a little metadata. That is, add some semantic smarts to existing JSON document. What people miss is that the semantic web was never an attempt to fit everybody into a straightjacket but rather an attempt to mash together all the data in the world into a huge Katamari ball.
- zcw100 4y agoThat so beautifully captures the attitude of the semantic web community, "What people miss...". Didn't miss anything. It's this attitude, that the problem isn't with what's been done, it's that other people have failed to recognize how brilliant it all is.
- PaulHoule 4y agoIt hasn't been communicated well. The work hasn't been planned well. SHACL should have been introduced before OWL. RDFS doesn't really work as a data integration language because you can't write a rule like ?x :lengthInInches ?t -> ?x :lengthInCentimeters ?t*2.54 Since RDF isn't fit for the purpose it was designed for, people get really confused about it. The strange thing is that semweb work has been massively overfunded in zones that face severe multilingual problems (Europe), but it seems to be almost banned in academia in the United States. (except for a few enclaves I can enumerate on the fingers of a few hands.) I know quite a few Africans who are semweb believers because they are looking at markets fragmented by languages, but they don't have the overfunding that Europeans have. (Funny enough I am working on a standards doc for ISO 20022 which is badly wanted by Chinese authorities who are interested in semweb tech because they want to feel included in financial messaging despite language barriers.)
- eterevsky 4y agoPardon my ignorance, but is any of this used for anything practical?
- 4y ago
- zozbot234 4y agoYou may be right about the pre-existing "wedge" in the semantic/LD community, but this paper is an attempt to address it and provide a unifying perspective. I'm not quite sure what you're seeing as "arrogant" or tone deaf in it.
- PaulHoule 4y agoPeople get seduced by specifications that don't really specify anything. Conversely there is a lot of pushback against W3C standards because they are specific and unfortunately people don't see that at freedom (freedom to choose tools that interoperate) and don't see the slavery in being stuck with poorly specified "standards" that are controlled by one entity. GraphQL is improved (we almost know what the algebra is) but originally it was an asymmetric specification meant to keep power in the hands of Facebook. That is, they didn't want to specify what the exact rules for traversing the graph are because they have commercial reasons for controlling what information you can get plus some responsibility to protect user's privacy. Schema.org was another asymmetric standard, as it wasn't particularly good for exposing semantic metadata that people could consume as such (it took me a few years to really figure out how) but it was great for companies like Google to make a training set that would ultimately let them extract entities from documents that aren't marked up. It achieved some popularity because there is no limit to the hoops people will jump through if it improves their SERP rating from 78 to 55.
- j-pb 4y ago> People get seduced by specifications that don't really specify anything. Like the RDF 1.1 Spec?[https://www.w3.org/TR/rdf11-concepts/ https://www.w3.org/TR/rdf11-concepts/] The whole "abstract syntax" shenanigans that RDF pulls is one of its biggest flaws. It makes the entire ecosystem huge and unwieldy, and has little upsides besides giving everybody their favourite serialisation flavour. It makes things like canonical representations for content addressable hashing and singing pretty much impossible, which is a huge detriment to proper authentication and provenance tracking. It also pulls in all of these other open ended standards, where everything and anything is a valid subject identifier, so long as it's URI resolvable by http (which is pretty vague and random). Subjects and predicates should have always just been 16byte random UIDs, which at the same time would have delivered us from the bane of blank nodes, endless discussions on predicate names, and broken links. The object part should also have been limited in the amount of data it can hold, just hash anything bigger and store it in some form of content addressable blob store.