4 ms·
With all potential advantages of semantic technologies, I'm wondering about whether their adoption is slowed down by performance issues of inference engines (re
by ablekh 7y ago
With all potential advantages of semantic technologies, I'm wondering about whether their adoption is slowed down by performance issues of inference engines (reasoners) on very large (> 100B of nodes) datasets (e.g., AWS has decided to exclude semantic inference functionality from their Neptune graph database, citing performance issues - though relevant product leads have expressed interest in including inference, based on use cases etc.). Are there any recent achievements (preferably, open source) on the front of dramatically speeding up relevant engines?
- __afk__ 7y agoAs a semantic architect, this is not my experience. In fact, I see very few large graphs in the wild. The problem is, unsurprisingly, that describing data is difficult. Relating your own conceptualization of a domain to anothers is frustrating and time consuming. It will always be easier to create a bespoke model. So, people just don't do it. As for OBO, there are many interesting comments here. The OBO ontologies all utilize BFO as an upper-level and in this regard they are united. But otherwise, their quality and utility varies tremendously. I still believe in this work and hope that one day everyone will think about their data as being longer-lived and more important than the software that generated it.
- ablekh 7y agoThank you for sharing your thoughts. Just curious: If you were tasked with architecting and implementing a semantic layer for a complex SaaS platform in a large domain from scratch, what would be your approach and what technology stack would you prefer to use and why? What best practices would you adopt, if any?
- chrismungall 7y agoMany large OBO ontologies use EL++ reasoning (e.g. Elk), performing DL reasoning on smaller chunks (e.g. relations). Having said that newer reasoners like Konklude apparently do well with DL reasoning over combinations of large ontologies. For "data" / ABoxes we have had a lot of success with the RL subset. We use this a lot https://github.com/balhoff/arachne https://github.com/balhoff/arachne But ultimately it depends on what you want to do. In the life sciences subsets of FOL only buy you so much, and some kind of statistical or probabilistic inference is required. Mostly this is combined with logical inference in crude ad-hoc ways...
- ablekh 7y agoI appreciate you sharing your insights and references. Will definitely research on relevant aspects and linked projects.