9 ms·
A Review of the Semantic Web Field
- stareatgoats 6y agoThis seems to me to be an insightful and comprehensive overview of the Semantic Web, both current status and how we got here. People like me, who have long been wanting to better understand the (obviously sprawling) concepts involved will be able to use the article as a good entry point. That said, the expressed hope of consolidation in the field is likely still some way off. AI has taken over a lot of the promise that the Semantic Web originally held. But AFAICS there are two drivers (also mentioned in the article) that potentially could provide the required impetus for a reignited interest in the Semantic Web: Firstly the need for explainable AI, and secondly the probable(?) coming breakthrough in natural language processing and automatic knowledge graph or ontologies from text. All in all, it seems way too early to write off the Semantic Web field at this point.
- bryanrasmussen 6y agoAs always should look at metacrap (http://www.well.com/~doctorow/metacrap.htm http://www.well.com/~doctorow/metacrap.htm) when discussing the semantic web - Certain kinds of implicit metadata is awfully useful, in fact. Google exploits metadata about the structure of the World Wide Web: by examining the number of links pointing at a page (and the number of links pointing at each linker), Google can derive statistics about the number of Web-authors who believe that that page is important enough to link to, and hence make extremely reliable guesses about how reputable the information on that page is. This sort of observational metadata is far more reliable than the stuff that human beings create for the purposes of having their documents found. It cuts through the marketing bullshit, the self-delusion, and the vocabulary collisions. in short, engineering triumphs over data entry.
- sammorrowdrums 6y agoI found that Job Postings are an exception. Google picks up on them, has a special API to submit them direct (due to slow crawling) and close them. So long as you're a good actor that will get you far. If your data is low quality, wrong, error prone or otherwise you'll not get shown and will likely receive manual actions and end up in the Google proverbial sin bin. I have found that incentives align for job postings. That obviously doesn't prove that metadata is not flawed, just that there are areas where it seems to work well.
- LukeEF 6y agoWe built a new semantic database first in university and then commercial open source (TerminusDB). We use the web ontology language (OWL) as a schema language, but made two important - practical - modifications: 1) we dispense with the open world interpretation; and 2) insist on the unique name assumption. This provides us with a rich modelling language which delivers constraints on the shapes in the graph. Additionally, we don't use SPARQL, which we didn't find practical (composability is important to us) and use a Datalog in its place (like Dataomic and others). Our feeling on interacting with the semantic web community is that innovation - especially when it conflicts with core ideology - is not welcome. We understand that 'open world' is crucial to the idea of a complete 'semantic web', but it is insanely impractical for data practitioners (we want to know what is in our DB!). Semantic web folk can treat alternative approaches as heresy and that is not a good basis for growth. As we came from university, I agree with comments that the field is too academic and bends to the strange incentives of paper publishing. Lots of big ideas and everything else is mere 'implementation detail' - when, in truth, the innovation is in the implementation details. There are great ideas in the semantic web, and they should be more widespread. Data engineers, data scientists, and everybody else can benefit, but we must extract the good and remove ideological barriers to participation.
- tannhaeuser 6y agoYou're right to emancipate from the grab that SemWeb has had on the field for so long and turn to Prolog/Datalog and practical approaches IMO. Open world semantics and sophisticated theories may have been a vision for the semantic web of heterogenous data, but in reality RDF and co are only used in certain closed-world niches IME. Pascal Hitzler is one of the more prolific authors (especially with the EU-funded identification of description logic fragments of OWL2 which are some of the better results in the field IMO), but beginning this whole discussion with W3C's RDF is wrong IMO when description logic as more or less variable-free fragments of first-order logic with desirable complexities was a thing in 1991 or earlier already. Nit: careful with datomic. It's clearly not Datalog, but an ad-hoc syntax whereas Datalog is a proper syntactic subset of Prolog. And while I don't like SPARQL, it still gives quite good compat for querying large graph databases.
- hyperion2010 6y agoThe reason why tools like Protégé have not been sufficiently developed is because of infighting in the academic ontology community in addition to the reasons listed by the author. It has set the whole community back at least 5 years.
- j-pb 6y agoI think that's a symptom, not the cause. The complexity of web standards in general smother it with it's own weight. The common web has enough raw financial and person backing to grind through that. The semantic web does not. CURIEs and the depending standards alone are well over 100 pages. Language tags alone has 90. RDF has like 100, Sparql has a combined of more than 300, and OWL has more than 500, even though it assumes that the reader is generally familiar with Description logics, so it's probably a couple thousand if you take the required academic literature into account. Nobody is going to read all of that, let alone build that. Especially not a bunch of academics who don't care about the implementation as long as it's good enough to get the next paper out the door. So everybody pools on these few projects, because they're the only thing that's kinda working. OWLAPI, Protege, ... uh that's it. Because everything else, is broken and unfinished. Here's a thought experiment, name one production ready RDF libray for every major programming language (C, Java, Python, Js), that doesn't have major, stale, unresolved issues in their issue tracker. It's all broken, and there is simply too much work required to fix things. It's only natural that people start to infight when there is only few hospitable oasis. What we need is a simpler ecosystem, where people can stake their claim on their niche, where they have the ability and power to experiment and explore.
- syats 6y agoI agree with this. It is common to hear "Partial SPARQL 1.1 support"... or "Partial OWL compatibility" or "A variant of SKOS is supported". While it is true that full ECMA6/HTTP2/IPv6/SQL is also rarely provided by implementations, this doesn't hinder their use in productive environments. I think it is rare to reach the parts of ECMAscript that aren't implemented, or the corners of SQL that Postgres/MariaDB don't support. In many of the "Semantic Web Stack", however, one quickly reaches a "not implemented" portion of the 500 page owl standard.
- tammet 6y agoThe whole field has been dominated by research, i.e. the wish to make simple things complicated (in order to publish papers) as opposed to engineering, i.e. making complicated things simple (in order to produce usable software efficiently). As a result the standards are horrendously - and needlessly - complicated. The few major practical outcomes like the schema.org, json-ld and the google annotation system, are results of engineering, not research. Alas, json-ld has also taken a turn towards hypercomplexities.
- huskyr 6y agoYeah, this is an unfortunate consequence of having the whole ecosystem mostly within academia, including the lack of tutorials and proper documentation (e.g. not a 500 page standard). IMO the most interesting place right now for semantic web development is Wikidata. It's still pretty difficult for newcomers to contribute (as is the case for all Wikimedia projects) but at least it has many eyeballs and a very active community / ecosystem.
- ivan_ah 6y ago+1 for WIKIDATA There are lots of useful WIKIDATA links and demos on this page: https://www.wikidata.org/wiki/User:Daniel_Mietchen/FSCI_2017#Introductions_to_Wikidata https://www.wikidata.org/wiki/User:Daniel_Mietchen/FSCI_2017...
- krallistic 6y agoMaybe a good indicator that there is only minor (industry) need/benefit. The "biggest" Knowledge Graph is Google, but it is unclear, how much there is actually Semantic Web and how much search, ML, NLP etc.. They are all nice ideas, but the practical usecases are rare. I am skeptical of the often touted usecase in Medicine/Drug Interactions. The only time i saw it in the industry, it was not really used by the lab technicians. Because all questions the system could answer, were trivial. The promise of "the system can inference new combinations/interactions" was never fulfilled.
- cheph 6y ago
- xkvhs 6y agoMaybe for its time it seemed like a good idea.. Like SOAP or manual features for image classification. Today, it's clear that languages and knowledge don't really work like that, and it's not practical to approach them this way. I've learned about the OWL and SPARQL 12 years ago, and it already felt like a very dated idea. But then who knows... Everybody have given up on NNs once too.
- krallistic 6y agoThe comparisons to NLP presents a good view on the problems. Its "easy" to write some logic rules to parse input text for a 50% demo. But then you want to improve & scale, and suddenly all the nuances, bites you. The rules get bigger, nested and complicated. Traditional NLP tried that avenue for a while, with decent success in small usecases, but for larger problems without success. (Compared to stuff like BERT & GPT, which still have a lot of problems) Similar with Knowledge Graphs, you can show some nice properties on inferring knowledge on small problems, but the real world is much more approximate and unclear than some (binary) relationships. Personally i think we Humans lack the mental capacity to build large models with complex interactions.
- ta988 6y agoThat's true as individuals I do not believe we can. Only as a group and with the help of tools, which is what semweb tried to achieve. We found out the tools weren't the most practical and learned a lot. Now we need the Tensorflow of those approaches, something easy to use not platform centric and with a low barrier of entry.
- cheph 6y ago> Today, it's clear that languages and knowledge don't really work like that, and it's not practical to approach them this way. There are many applications of Semantic Web that has little to do with natural languages. If you have a better option for all the existing RDF data sets (https://lod-cloud.net/ https://lod-cloud.net/, https://www.wikidata.org/ https://www.wikidata.org/) and ontologies (http://www.ontobee.org/ http://www.ontobee.org/, https://schema.org/ https://schema.org/) it would be good to be explicit about it. I would prefer to have more data (e.g. data from US federal reserve data, world bank data) as RDF and accessible via SPARQL endpoints than less, because it is much more useful as RDF than as CSV, in my opinion.
- actsofthecla 6y agoI'm only a hobbyist in this area, but I wonder why the review wouldn't mention some of the graph databases as, at least, semantic web adjacent. Their relative success seems to lend credence to the overall vision of the semantic web and its supporting technologies. For example, are there really more than surface syntactical differences between SPARQL and Cypher? Even though it was over-hyped, I like the semantic web because it supports a conception for the future that includes something other than neural network black-boxes. However, whether the ideas deliver remains to be seen. If anyone is looking for an introduction, then I think the Linked Data book from Manning is worth mentioning--it might be a little dated at this point. The author provides a coherent introduction and helps, especially, in cutting through the confusing proliferation of acronyms that characterizes this field. As others have mentioned, reliable software is a major stumbling block. It's especially unfortunate that there isn't better browser support, of RDFa for example.
- namedgraph 6y agoCheck our SPARQL-driven Knowledge Graph management system :) https://atomgraph.github.io/LinkedDataHub/ https://atomgraph.github.io/LinkedDataHub/
- tragomaskhalos 6y agoMy 10,000 ft layperson's view, to which I invite corrections, is broadly: - The semantic web set off with extraordinarily ambitious goals, which were largely impractical - The entire field was trumped by Deep Learning, which takes as its premise that you can infer relationships from the exabytes of human rambling on the internet, rather than having to laboriously encode them explicitly - Deep Learning is not after all a panacea, but more like a very clever parlour trick; put otherwise, intelligence is more than linear algebra, and "real" intelligences aren't completely fooled by one pixel changing colour in an image, etc. - Hence, we have come back round to point 1 again ?
- breck 6y agoI think you are spot on. I think what we'll see is Deep Learning/Human Editor "Teams". DL will do the bulk of the relationship encoding, but human domain experts will do "code reviews" on the commits made by DL agents. Over time fewer and fewer commits will need to be reviewed, because each one trains the agent a bit more.
- cheph 6y agoDeep Learning does not even operate in the same space as where most of Semantic Web is being used today, some examples: - https://schema.org/ https://schema.org/ - https://www.wikidata.org/ https://www.wikidata.org/ - https://lod-cloud.net/ https://lod-cloud.net/ - http://www.ontobee.org/ http://www.ontobee.org/ - https://catalog.data.gov/dataset?res_format=RDF&_res_format_limit=0 https://catalog.data.gov/dataset?res_format=RDF&_res_format_... - https://ukparliament.github.io/ontologies/ https://ukparliament.github.io/ontologies/ - https://ckan.publishing.service.gov.uk/dataset?res_format=SPARQL&_res_format_limit=0 https://ckan.publishing.service.gov.uk/dataset?res_format=SP... - https://ckan.publishing.service.gov.uk/dataset?res_format=RDF&_res_format_limit=0 https://ckan.publishing.service.gov.uk/dataset?res_format=RD... - https://data.nasa.gov/ontologies/atmonto/index.html https://data.nasa.gov/ontologies/atmonto/index.html - https://data.europa.eu/euodp/linked-data https://data.europa.eu/euodp/linked-data
- fauigerzigerk 6y ago>The entire field was trumped by Deep Learning, which takes as its premise that you can infer relationships from the exabytes of human rambling on the internet, rather than having to laboriously encode them explicitly I don't think machine learning can ever replace data modeling, because data modeling is often creative and/or normative. If we want to express what data must look like and which relationships there should be, then machine learning doesn't help and we have no other choice than to laboriously encode or designs. And as long as we model data we will have a need for data exchange formats. You could categorise data exchange formats as follows: a) Ad-hoc formats with ill defined syntax and ill defined semantics. That would be something like the CSV family of formats or the many ad-hoc mini formats you find in database text fields. b) Well defined syntax with externally defined often informal semantics. XML and JSON are examples of that. c) Well defined syntax with some well defined formal semantics. That's where I see Semantic Web standards such as RDF (in its various notations), RDFS and OWL. So if the task is to reliably merge, cleanse and interpret data from different sources then we can achieve that with less code on the basis of (c) type data exchange formats. But it seems we're stuck with (b). I understand some of the reasons. The Semantic Web standards are rather complex and at the same time not powerful enough to express all the things we need. But that is a different issue than what you are talking about.
- mxmilkb 6y agoNice though nothing about Turtle or LV2 https://www.w3.org/2007/02/turtle/primer/ https://www.w3.org/2007/02/turtle/primer/ https://github.com/lv2/lv2/wiki https://github.com/lv2/lv2/wiki Also, #swig (semantic web interest group) exists on freenode.
- smarx007 6y agoI do research in this field but I am a programmer by training before I entered this research field. I have talked to many academics and they agree that industry needs something simpler, more approachable and something that solves their problems in a more direct way, so it's definitely not an "academic exercise" for many researchers. However, I failed to convince people that we need to implement the 2001 SciAm use case (https://www-sop.inria.fr/acacia/cours/essi2006/Scientific%20American_%20Feature%20Article_%20The%20Semantic%20Web_%20May%202001.pdf https://www-sop.inria.fr/acacia/cours/essi2006/Scientific%20..., see the intro before the first section) using 2021 technologies (smartphones are here, assistants are here, shared calendars are easy, companies have APIs, the only thing missing is a proper glue using semantic web tech). This goes to the core thesis of this paper that semantic web is awesome as the set of ideas and approaches but the Semantic Web as the result of all this work may look underwhelming or irrelevant today. I like to point everyone who disagrees with me to the 1994 TimBL presentation at CERN (https://videos.cern.ch/record/2671957 https://videos.cern.ch/record/2671957) where he talks about the early vision of semantic web (https://imgur.com/aS2dbf6 https://imgur.com/aS2dbf6 or around 05:00 in the video), which looks awfully like IoT (many years before the term even existed). We simply cannot fault someone who envisioned communication technologies for IoT in 1994 for getting the technology a bit wrong. Today's technologies simply cannot handle the use-cases for which SemWeb was designed for properly: 1) The web is still not suitable for machines. Yes, we have IoT devices that use APIs but nobody will say it's truly M2M communication at its best. When APIs go down devices get bricked, there is no way to get those devices to talk to any other APIs. There is no way for two devices in a house to talk to each other unless they were explicitly programmed to do so. 2) We don't have common definitions for the simplest of terms. Schema.org made a progress but it's very limited because it serves search engine interest, not the IoT community. There is no reason something like XML NS or RDF NS should not be used across every microservice in a company. Using a key (we call them predicates, but not important here) "email:mbox" (defined in https://www.w3.org/2000/10/swap/ https://www.w3.org/2000/10/swap/ very long time ago) you can globally denote the value is an email. 3) Correctness of data and endpoint definition still matters. We threw away XML and WSDL but came back to develop JSON Schema and Swagger. We are trying to get there. JSON Schema, Swagger etc. all make efforts in the direction of the problems SemWeb tried to address. One of the most "semantic" efforts I see done recently is GraphQL federation, which has been a semantic web dream for a long while: being able to get the information you need by querying more than one API. This only indicates the problems that semantic web tried to address are still viable. If anyone has attempted an OSS reimplementation of the 2001 "Pete and Lucy" semantic web use case (ie as an Android app and a bunch of microservices), please point me in the right direction. Otherwise, if anyone is interested in doing it, I am all ears (https://gitter.im/linkeddata/chat https://gitter.im/linkeddata/chat is an active place for the LOD/EKG/SW discussion).
- mark_l_watson 6y agoEven though I have been working off and with SW and linked data tech for twenty years, I share some of the skeptical sentiments in comments here. I am keenly interested in fusion of knowledge representation with SW tech and deep learning. I wrote a short and effective NLP interface to DBPedia two weekends ago that you can experiment with on Google Colab https://colab.research.google.com/drive/1FX-0eizj2vayXsqfSB2ONuJYG8BaYpGO https://colab.research.google.com/drive/1FX-0eizj2vayXsqfSB2... that leverages Hugging Face’s transformer model for question answering. You can quickly see example use in my blog https://markwatson.com/blog/2021-01-18-dbpedia-qa-transformer/ https://markwatson.com/blog/2021-01-18-dbpedia-qa-transforme...
- ivan_ah 6y agoWow, what a great summary with lots of realism and nuances. I agree with the author's conclusions that what is missing is consolidation and interoperability between standards (e.g. make Protégé easier to use and ensure libraries for RDF parsing and serializations exist for all languages). No technology will be adopted if it requires PhD-level ability to handle jargon and complexity... but if there were tutorials and HOWTOs, we could see big progress. Personally, I'm not a big fan of the "fancy" layers of the Semantic Web Stack like OWL (see https://en.wikipedia.org/wiki/Semantic_Web_Stack https://en.wikipedia.org/wiki/Semantic_Web_Stack ), but the basic layers of RDF + SPARQL as a means for structured exchange of data seem like a solid foundation to build upon. It's really simple in the end: we've got databases and identifiers. INTERNALLY to any company or organization, you can setup a DB of your choosing and ensure data follows a given schema, with data linked through internal identifiers. When you want to publish data EXTERNALLY, you need to have "external identifiers" for each resource, and URIs are a logical choice for this (this is also a core idea of REST APIs of hyperlinked resources). Similarly, communicating data using the a generic schema capable of expressing arbitrary entities and relations like RDF and JSON-LD is also a logical next step, rather than each API using it's own bespoke data schema... As for making web data machine-readable, the key there is KISS: efforts like schema.org with opt-in, progressive enhancements annotations are very promising. For anyone wanting to know more about this domain, there is an online course here: https://www.youtube.com/playlist?list=PLoOmvuyo5UAeihlKcWpzVzB51rr014TwD https://www.youtube.com/playlist?list=PLoOmvuyo5UAeihlKcWpzV... The whole course is pretty deep (would take a month to go through it all), but you can skip ahead to lectures of specific interest.
- kmerroll 6y agoHonestly disconcerting to see mostly negative responses in this thread: awful community, overly complicated, research focused, academic nitwits gone wild, etc. Pretty sure there's some truth here, but would suggest the deeper argument is against semantic web as evolution of the world-wide-web. Agree this isn't likely to happen in my lifetime. Right up there with, mostly hated, Javascript, I happen to think there are good parts of the semantic web technologies and that the pivot towards industry adoption of the graph data models related to knowledge graphs, ontologies, and SPARQL shows there are benefits outside of academic paper mills. I don't have a dog in this fight (TerminusDB), but applying some reasonable expectations and accepting the limitations of the semantic web tools has been very successful on many projects. Even more so, innovation and improvements in graph data repositories are making triple-stores and graph-based models compelling for some use cases. Not going back to CSV hell if there are better alternatives.