7 ms·
Ontology for the life sciences: genes, proteins, diseases, all expressed in OWL
- jakeogh 7y agoIs https://bitbucket.org/jibalamy/owlready2/src/default/ https://bitbucket.org/jibalamy/owlready2/src/default/ the standard py lib for accessing these? Is there a recent tutorial on querying the ontology from the terminal? I found a few on yt, this one is 5y old: https://www.youtube.com/watch?v=5DCS9LE-8rE https://www.youtube.com/watch?v=5DCS9LE-8rE
- haddr 7y agoJena API was quite common to use with Java, and RDFLib for python.
- rzzzt 7y agoOWL API is also a contender in the Java libraries category: https://github.com/owlcs/owlapi https://github.com/owlcs/owlapi
- chrismungall 7y agoFor a general purpose OWL python library there is OWLReady2 and funowl: https://github.com/hsolbrig/funowl https://github.com/hsolbrig/funowl For a Python library that intends to provide a level of abstraction more appropriate to bioinformatics use cases see https://github.com/biolink/ontobio/ https://github.com/biolink/ontobio/ The ontology data model for this is based more on obograph json than on OWL. See https://douroucouli.wordpress.com/2016/10/04/a-developer-friendly-json-exchange-format-for-ontologies/ https://douroucouli.wordpress.com/2016/10/04/a-developer-fri... and https://douroucouli.wordpress.com/2016/10/04/a-developer-friendly-json-exchange-format-for-ontologies/ https://douroucouli.wordpress.com/2016/10/04/a-developer-fri... Of course, it's always possible to use an RDF level library such as rdflib, but this can be low-level for OWL. Even simple bio-ontologies often make frequent use of existential restrictions and axiom annotations. And of course it's possible to use a python-jvm bridge to access the fully featured java OWL API. (I am involved in both ontobio and funowl)
- gibsonf1 7y agoThis is fantastic work and will most definitely speed up innovation in the life sciences.
- bordercases 7y agoI've tried to work with these in both health science and agricultural industries. The bottlenecks are the reasoning they support and the quality of the data being annotated with the ontologies. The long and the short of it is that every organization is going to partially disagree with the subtleties of certain categorizations which proliferates standards and ad-hoc modifications. There needs to be means of adjusting mappings at runtime and to store changes. It also needs to be so stupid/simple that an old-guard biologist can use it and immediately comprehend the value. Querying ontologies is easy, working with annotations-qua-annotations is more difficult than it has to be, and as such organizations typically will want to roll their own.
- deleted 7y ago[deleted]
- killjoywashere 7y agoWhy does this not include the clinically relevant ontologies like ICD9, ICD10, SNOMED, LOINC, and CPT? Also, folks may find this article interesting: Classification, Ontology, and Precision Medicine (1) (1) https://www.nejm.org/doi/full/10.1056/NEJMra1615014 https://www.nejm.org/doi/full/10.1056/NEJMra1615014
- dan_fornika 7y agoThe OBOFoundry has a set of principles[0] that not all ontologies conform to. An ontology project can register with the OBOFoundry through their GitHub site [1] if they follow the principles. [0] http://obofoundry.org/principles/fp-001-open.html http://obofoundry.org/principles/fp-001-open.html [1] https://github.com/OBOFoundry/OBOFoundry.github.io https://github.com/OBOFoundry/OBOFoundry.github.io
- jerven 7y agoIt is a combination of licensing. Snomed-CT is there in part via ULMS. But these days there is an official Snomed and Loinc via FHIR maintained by their hosting organizations.
- mnemonicsloth 7y ago1. Good link 2. I wish I could answer your question about the other projects. All I can say is that I saw this on Friday in a bioinformatics lecture. I'm not a bioinformatician. I don't know all the various projects relate to each other, although there should be some kind of liaison effort (right?). I do think I remember something about SNOMED, but the rest are blanks.
- i_am_nomad 7y agoOne thought about this, ICD9/10 are medical coding standards. They have only a very shallow semantic depth to them, they’re more about completeness and specificity. They don’t map well to what OWL tries to do.
- jerven 7y ago
- glofish 7y agoit is one of these misguided efforts to bring some order to life sciences but they go about it the wrong way. It is so heavy-handed, the website so obtuse and confusing that no life scientist I know (and I have worked with hundreds of them) is even aware of let alone understand what an ontology is or how to use it. Might be hard to believe for the uninitiated but they even got some of the namings is wrong from the start, for example, they have a GO (Gene Ontology) but that ontology is not actually describing genes! It describes gene products (like proteins) a huge big difference! Not to mention grossly misleading the very life scientists it is meant to help.
- xipho 7y agoCurious, coming from a user named 'glofish' that they are unaware of GOs pervasiveness in the zebrafish model organsim world. GO is everywhere in the genomics/model organism world, and only growing in use. Also funny that the user is looking at labels, when concepts are the core of an ontology. Take a look at how many terms in GO are deprecated, it has evolved over time, I know very few scientists who get things right the first time. Sure there are issues, but many FOSS principles apply throughout. Also, NIH would be to differ that OWL/OBO isn't important- https://monarchinitiative.org/ https://monarchinitiative.org/.
- netfl0 7y agoExcellent reply. The GO is very impressive and I don’t think most folks have seen it.
- glofish 7y agoPeople do not understand and misuse GO. They think it is Gene-Ontology - it is not. It is gene product ontology. In addition, I guarantee you that the vast majority of people using GO understand neither: - what the ontology actually is - how it works - who decides what gets into the data - why is something labeled a certain way - what evidence is there - how the terms interconnect - what the hierarchy all means All that because the concepts are not explained properly, nor the site is of any use to help you figure these out. It is mostly an illusion - and I am saying that as someone that uses GO a lot. I am intimately familiar with all of its pitfalls. At best some people know is that a label is attached to a gene. Finally GO is also perhaps the odd one out, the only ontology that is known somewhat because it is misused a lot. I invite you to go to the link on the top post and note how many other ontologies are there ... hundreds? Ask a life scientist how many they have heard of. Is my original statement all that wrong really? I don't think so. These ontologies are dead-end.
- i_am_nomad 7y agoThis has been around for a while. It’s excellent, and I’ve tried to get it adopted at several life science organizations. The problem is culture. Principal investigators are notoriously stubborn about how they do things, and that includes what they call things. Manufacturers and supply companies also don’t see much benefit to standardizing around this, and of course realize the large downside (alienating the aforementioned PIs). What would help this would be the introduction of OWL into major software suites, primarily LIMS and ELN packages, and include tools for aligning terminology and concepts with OWL.
- dan_fornika 7y agoI think the various OBOFoundry ontologies are at different stages of maturity. One of their design principles is interoperability, but I'm not sure how often one is able to reason effectively across ontologies unless great care is taken to ensure that logical axioms are sound. I'm involved in a project[0] that has an aim to provide better data integration by using OBOFoundry ontologies, but it's been a challenge in practice to merge the software with the ontologies in a coherent way. [0] https://www.irida.ca/data-integration/ https://www.irida.ca/data-integration/
- jerven 7y agoThere are a few which are really great. Basically my shortcut is Chris Mungall involved then it is logically and biologically ;) sound. While we at SIB are more towards the RDF/SPARQL part of the spectrum (https://edu.sib.swiss/course/view.php?id=440 https://edu.sib.swiss/course/view.php?id=440). We do use OWL and obofoundry projects like Uberon and GO. For UniProt I took a lot of pain to make sure it is really compatible. We did the same for Rhea and ChEBI. This has paid of handsomely in new query capabilities.
- sndean 7y ago> Principal investigators are notoriously stubborn about how they do things I've come across things a simple as "I wrote a Perl script that does that, but it uses KEGG not GO terms" for their reasoning. That, or – if they're okay with using gene ontology – we have to also use the other system in parallel. And then the results becomes a discussion of how the two different annotation systems display the results. Life science could use a benevolent dictator.
- dan_fornika 7y agoThere's an interesting clojure library called 'Tawny-OWL'[0] that is designed for building ontologies like this. It allows one to define a set of entities that follow patterns and logical axioms. [0] https://github.com/phillord/tawny-owl https://github.com/phillord/tawny-owl
- ablekh 7y agoWith all potential advantages of semantic technologies, I'm wondering about whether their adoption is slowed down by performance issues of inference engines (reasoners) on very large (> 100B of nodes) datasets (e.g., AWS has decided to exclude semantic inference functionality from their Neptune graph database, citing performance issues - though relevant product leads have expressed interest in including inference, based on use cases etc.). Are there any recent achievements (preferably, open source) on the front of dramatically speeding up relevant engines?
- __afk__ 7y agoAs a semantic architect, this is not my experience. In fact, I see very few large graphs in the wild. The problem is, unsurprisingly, that describing data is difficult. Relating your own conceptualization of a domain to anothers is frustrating and time consuming. It will always be easier to create a bespoke model. So, people just don't do it. As for OBO, there are many interesting comments here. The OBO ontologies all utilize BFO as an upper-level and in this regard they are united. But otherwise, their quality and utility varies tremendously. I still believe in this work and hope that one day everyone will think about their data as being longer-lived and more important than the software that generated it.
- ablekh 7y agoThank you for sharing your thoughts. Just curious: If you were tasked with architecting and implementing a semantic layer for a complex SaaS platform in a large domain from scratch, what would be your approach and what technology stack would you prefer to use and why? What best practices would you adopt, if any?
- chrismungall 7y agoMany large OBO ontologies use EL++ reasoning (e.g. Elk), performing DL reasoning on smaller chunks (e.g. relations). Having said that newer reasoners like Konklude apparently do well with DL reasoning over combinations of large ontologies. For "data" / ABoxes we have had a lot of success with the RL subset. We use this a lot https://github.com/balhoff/arachne https://github.com/balhoff/arachne But ultimately it depends on what you want to do. In the life sciences subsets of FOL only buy you so much, and some kind of statistical or probabilistic inference is required. Mostly this is combined with logical inference in crude ad-hoc ways...
- mark_l_watson 7y agoVery nice! I have a keen interest in the semantic web and linked data (I have written two books on the subject, and the current book I am writing on the Hy programming language also has linked data and knowledge representation examples). Much of the most interesting work has been done in medicine and biology. Off topic, but I am really split by wanting to use full graph databases vs. RDF/RDFS/OWL. Different but overlapping use cases.
- smadge 7y agoI suspect a lot of interesting work on semantic knowledge graphs has been done internally at FAANG and others, but the results are trade secrets.
- jerven 7y agoIf you are in Bio the SPARQL usecase is just so nice. I work on the sparql.uniprot.org and others at SIB. Being able to federate without a hassle is such a powerful thing for research.
- chrismungall 7y ago> I am really split by wanting to use full graph databases vs. RDF/RDFS/OWL I think having an alternative OWL->graph mapping that utilized the features of property graphs would help a lot with using non-RDF graph databases. The RDF layering is sub-optimal. I wrote about this here: https://douroucouli.wordpress.com/2019/07/11/proposed-strategy-for-semantics-in-rdf-and-property-graphs/ https://douroucouli.wordpress.com/2019/07/11/proposed-strate...
- xvilka 7y agoSadly, there are not enough libraries to work with OWL from Julia and/or Rust.
- chrismungall 7y agoI don't have any recommendations for Julia, but for Rust have you tried fastobo? - https://github.com/fastobo/fastobo https://github.com/fastobo/fastobo - https://github.com/fastobo/fastobo-owl https://github.com/fastobo/fastobo-owl It has been tested heavily on all the ontologies on the obofoundry site. But I agree in general that working with OWL ontologies (particularly those that use nested OWL constructs) can be difficult to work with in non-JVM languages.
- throwawaysea 7y agoCan someone ELI5 this? What are some real life use cases this solves?
- jerven 7y agoPart of it is trivial. i.e. is "9606" used to identify a species or a pubmed paper. (While trivial it's a huge source of errors, a bit like manually managing memory trivial but still a huge source of CVEs) The other is for logical data quality inference. E.g. data is coded by super sub class, while query asks for middle class the right data is still retrieved. This goes beyond that into ever more complicated scenarios which are not double by connect by queries. Which get's into questions is the hind leg of an ant similar enough to a human leg to have an argument about the function of a protein transfer. (Uberon)
- TallGuyShort 7y agoCan someone help me understand what an ontology is supposed to be or do? I've seen the word come up in a number of HN articles lately and any definition I find online seems impossibly vague.
- steerpike 7y agoAn ontology is supposed to be a formal way of describing things so that the people who want to discuss their concepts have a very clear and disambiguated set of terms for conducting that discussion. There's more to it than that, but really that's the most practical explanation of ontologies within the most common context you'll likely encounter it. A good example is the BBC Program ontology[0] which you can see the practical usage of in the (still fantastic) talk 'Beyond the polar bear'[1] [0]https://www.bbc.co.uk/ontologies/po https://www.bbc.co.uk/ontologies/po [1]https://www.youtube.com/watch?v=8iaqcf9-riI https://www.youtube.com/watch?v=8iaqcf9-riI
- ImaCake 7y agoIt's a big graph, from graph theory. Think of a bunch of words connected to each other by lines ("edges") where the lines are some kind of causal link (e.g. both involved in processing sugars). It's basically organised like an XML file or JSON. The ontologies are actually quite small to download (46kb for the gene ontology). You can go look at one here [0] The idea of the ontologies is to use the same words for the same things, and have a kind of summary for all the research into the things listed in the ontology. Having this all standardised allows for a lot of interoperability. Big problems with the ontologies is they are years out of date and have poor uptake in many parts of life science research. Which is shame because they are a fantastic fundamental tool, it's just that the life scientists are underfunded and already work ridiculous hours. 0. http://geneontology.org/docs/download-ontology/#go_obo_and_owl http://geneontology.org/docs/download-ontology/#go_obo_and_o...
- zmmmmm 7y agoOne very simple thing it does is help you understand hierarchical semantic relationships. So, for example, you could have sibling => => brother => sister Then you know when eg: parsing text that if you observe the symbol "sister" in one place and "sibling" in another, they could be the same thing. This kind of thing is used in the medical space where one doctor might discribe a problem using a general term and the next in a more specific way and you have to be able to match them.
- DrScientist 7y agoI'm skeptical about these - as often it's missing the point. Two main challenges: - one world view of categorization does not meet all needs and is fragile over time - the challenge is still shared understanding between humans, not between computers. Let me explain the second: You have a large set of computer readable definitions - painstakingly built. You have some new data that needs mapping into the ontology. Typically a person has to do that. To map they need to understand the ontology in the same way as every other person doing mappings. The larger, more complex/'definitive' the ontology the harder this is. Finally, science moves on and your ontology/and or understanding is out of date. That's not to say shared definitions are not incredibly useful - just that writing them down in a computer friendly way doesn't guarantee that the definitions are actually shared between people.
- chrismungall 7y agoI am part of OBO (and GO, which has also been mentioned here a few times). It's nice to see the discussion here, lots of good comments, including the critical ones! I'm going to address a few themes that have emerged, hopefully this is informative and can generate useful discussion. Yes, the OBO site is not really very biologist friendly (or really anyone-friendly). It is more geared towards ontology developers, biocurators, and the kind of people who might build tools biologists would use. I would recommend portals such as the OLS (https://www.ebi.ac.uk/ols https://www.ebi.ac.uk/ols) for biologists -- but even this site tends to be used by bioinformatics-savvy folks. Domain scientists and users often use ontologies indirectly. For example, the Human Phenotype Ontology is used in many clinical settings for entering patient phenotypes, and subsequently in diagnosis (making use of the logical structure of the ontology). The cell ontology is central to many single-cell seq efforts. And of course the GO is ubiquitous in interpretation of experimental data via enrichment analyses. One thing we have tried hard to do with the OBO site is ensure we have up to date metadata for all registered ontologies -- including GitHub trackers. So if you have comments on any particular ontology then engaging the developers via their issue tracker is strongly encouraged. And OBO itself has a tracker, as well as a mail list anyone can join. One thing we have not done a great job of is in giving any kind of order to the many ontologies now listed. It can be overwhelming to someone not familiar with the field. Which ontology should I use ?(or avoid?) Which ontologies integrate well together, and which ones have duplication or incompatibilities? We have a number of new developments in the pipeline that should improve the situation here through the development of a new mid-upper level ontology called OBO Core (https://github.com/OBOFoundry/Experimental-OBO-Core/ https://github.com/OBOFoundry/Experimental-OBO-Core/). We're also encouraging all ontology developers to use standard tooling and best practice which will make interoperation easier (https://github.com/INCATools/ontology-development-kit https://github.com/INCATools/ontology-development-kit). A few people were asking for explainers on ontologies and bio-ontologies in particular. We have a few on our resources page, which is linked from the main site: https://github.com/OBOFoundry/OBOFoundry.github.io/blob/master/resources.md https://github.com/OBOFoundry/OBOFoundry.github.io/blob/mast... (pull requests welcome!) Finally, the OBO project itself receives only a tiny amount of short term funding, and most of the ontologies that comprise it have little or no direct funding (exceptions being widely used ones such as GO). A lot of work is community effort. Not saying that to deflect any criticism - constructive criticism is good! Just to provide an explanation of which some things are the way they are.