5 ms·
There are several databases that annotate metabolic pathways, although you are right that we lack a comprehensive integrated data source. Some of the larger one
by ravila4 5y ago
There are several databases that annotate metabolic pathways, although you are right that we lack a comprehensive integrated data source. Some of the larger ones:
Reactome: https://reactome.org/ https://reactome.org/
KEGG: https://www.kegg.jp/kegg/pathway.html https://www.kegg.jp/kegg/pathway.html
Wikipathways: https://www.wikipathways.org/index.php/WikiPathways https://www.wikipathways.org/index.php/WikiPathways
BioCarta: https://www.hsls.pitt.edu/obrc/index.php?page=URL1151008585 https://www.hsls.pitt.edu/obrc/index.php?page=URL1151008585
SMPDB: https://smpdb.ca/ https://smpdb.ca/
ConsensusPathDB: http://cpdb.molgen.mpg.de/ http://cpdb.molgen.mpg.de/
Regarding the issue of of parsing pathway figure data from published papers: pathway-figure-ocr (https://github.com/hiplot/pathway-figure-ocr/ https://github.com/hiplot/pathway-figure-ocr/) is a project by the developers of Wikipathways which is trying to solve that issue.
- joshuamcginnis 5y agoThat's a great list, but it's not just about annotated metabolic pathways; there also needs to be information about the gene(s) that have been characterized as being involved in those pathways, along with information about the source organisms and potential target organisms for production. Same with respect to substrates needed for intermediary pathways.
- alevskaya 5y agoKEGG is a very useful index into known genes for this, I used it all the time when mining for possible enzymes in the past.
- joshuamcginnis 5y agoAgreed! And if it were complemented with added information about the organisms transcription machinery (promoters, terminators, etc), it could empower folks to generate DNA to test a new or modified metabolic pathway in a manner that's never been tried before.
- sjg007 5y agoGene Ontology (GO) is what you are looking for. KEGG is also useful. Welcome to systems biology. Mapping transcriptional networks/pathways is more challenging and a major focus of current research across all model organisms.
- ravila4 5y agoAll that information is available through various databases and APIs, but the exact pipeline you are looking for might not be easily accessible via a nice user interface. It requires a lot of "glue" code to to connect the data from one resource to another.
- malux85 5y agoCan you please list them? Then I can write all the glue code and pull them together.
- joshuamcginnis 5y agoMy idea is to essentially be the glue and build the interface.
- eweitz 5y agoSome notes towards those ends: WikiPathways supports advanced queries via their SPARQL API and UI. See [1] and [2]. I find WikiPathways nice because it lets logged-in users create and edit pathways, with a low barrier to entry. I've been building a way to find related genes using biochemical pathways [3]. The source code linked there includes practical examples for fetching information on genes in those pathways, which you rightly note is needed for something compelling. That and other code there might help spark ideas for you on how to glue together various biochemistry and molecular biology APIs to achieve your vision. I'm currently working on a way to drastically expand the set of organisms and pathways covered by WikiPathways. Yeast has 66 pathways there, compared to 1319 for human. By doing fast ortholog detection at runtime (using another SPARQL API, provided by OrthoDB [4]) I'm hoping to be able to convert relevant annotated pathways across organisms, e.g. human to yeast, mouse to rat, Arabidopsis to rice -- and vice versa. [1] http://sparql.wikipathways.org http://sparql.wikipathways.org [2] https://www.wikipathways.org/index.php/Help:WikiPathways_Sparql_queries https://www.wikipathways.org/index.php/Help:WikiPathways_Spa... [3] https://eweitz.github.io/ideogram/related-genes?q=RAD51&org=homo-sapiens https://eweitz.github.io/ideogram/related-genes?q=RAD51&org=... [4] https://sparql.orthodb.org https://sparql.orthodb.org
- ravila4 5y ago
- cknoxrun 5y agoPathWhiz is pretty general (https://smpdb.ca/pathwhiz https://smpdb.ca/pathwhiz) and covers most of that. Give it a go! You can sign in as a guest and start creating pathways. All of the gene information is included and pulled in from Uniprot as well as a bunch of other sources. Small molecule data is integrated with HMDB, DrugBank, and a bunch of other sources as well. (disclaimer: although I didn't work on this tool, I worked on the first version of SMDPB).
- killjoywashere 5y agoFun historical note: Steve Sprang (1) added a T terminator to inkpad for ipad (2) developed originally under the company name Taptrix, which was in a YC'10 batch (3), after being his first app, Brushes, was an early smash success on the iphone (4), at my request the day he open-sourced it (5) because I wanted a T terminator for drawing biochemical pathway maps! (1) https://news.ycombinator.com/user?id=ssprang https://news.ycombinator.com/user?id=ssprang (2) http://inkpad.art/ http://inkpad.art/ (3) https://news.ycombinator.com/item?id=6665261 https://news.ycombinator.com/item?id=6665261 (4) https://www.macworld.com/article/198770/brushes.html https://www.macworld.com/article/198770/brushes.html (5) https://github.com/sprang/Inkpad/issues/22 https://github.com/sprang/Inkpad/issues/22