3 ms·
The global alliance for genomics and health is a similar idea but not designed to be completely open nor linked. They don't acknowledge the semantic web. Practi
by eggie 11y ago
The global alliance for genomics and health is a similar idea but not designed to be completely open nor linked. They don't acknowledge the semantic web. Practically, it is implemented as CORBA using JSON. Everything is an API and after that an implicit data model (in JSON) is being produced for each type of concept. AFAIK security is a huge limitation here. The idea of GA4GH is data silos can communicate some things with each other, but not personally identifiable information.
I work on the project but find it pretty uninspiring. It presents a dark image of the future in which a handful of large tech companies control all of our biomedical data and we have to beg them to allow us to share it. I guess that sounds pretty similar to the present. Just switch bio and social and here we are.
- dekhn 11y agoI don't think that GA4GH is literally using CORBA. The data model is not implicit, it's explicit (there is a schema). The "data silos can communicate some things ... but not personally identifiable information" is a constraint placed on the alliance by legal system. As for the semantic web, every bio project I've seen which adopted the semantic web ultimately failed -- the semantic web seems like a great idea, but attempts to fully implement to the point where it's useful for research always fail. So I think they're focusing on areas where they are likely to succeed (collection and processing of large amounts of raw and derived data using pretty conventional processes, but at a much larger scale, with a solid authentication and access mechanism).
- eggie 11y ago> I don't think that GA4GH is literally using CORBA. It's not literally CORBA, but people who spent time implementing literal CORBA in a bioinformatics context (for instance, the original author of https://github.com/bioperl/bioperl-corba-server https://github.com/bioperl/bioperl-corba-server) have noted that the design pattern and discussions followed by the GA4GH are pretty similar to those had in the EBI when they attempted to unify everything using CORBA. > The data model is not implicit, it's explicit (there is a schema). There is a schema, but the semantics of the data model are encoded in the comments of the schema. Without hooking into some kind of ontological basis it doesn't seem possible to avoid this. > As for the semantic web, every bio project I've seen which adopted the semantic web ultimately failed -- the semantic web seems like a great idea, but attempts to fully implement to the point where it's useful for research always fail. I'm aware of at least one group in the GA4GH that uses RDF internally, then converts into the custom schemas produced by the group in order to maintain compatibility with the top-down designs of the project. I believe this is the phenotype group. These are the people who are most interested what the author of the linked page is describing, they have decided to use the technology you believe is doomed to fail. But, they aren't failing. As far as I can tell from their presentations they are one of two or three groups in the project that have produced a functioning system. It's very easy so say that hard things are impossible. This tends to keep them that way. I doubt we have any other viable option for building large distributed knowledge systems. The fact that these don't exist does not mean they are impossible to construct, but simply that no one has managed to do so yet. People leveled the same kinds of arguments against neural networks up until a few years ago, saying that they were a nice idea but destined to fail because they were too hard. > So I think they're focusing on areas where they are likely to succeed (collection and processing of large amounts of raw and derived data using pretty conventional processes, but at a much larger scale, with a solid authentication and access mechanism). The scales we're talking about are not even an order of magnitude above that which existing techniques allow. So I agree that they will succeed insofar as they simply adopt these existing community-driven standards and slap access control on top. However, in terms of generating new data models for genomics, I'm not so convinced that the centralized design and API-based approach which they are taking will work. I guess we will have to meet back here in a few years and see what happened.
- x1k 11y agoFor NN, we have a clear target: for example, beat HMM on speech recognition. To achieve that, you write a tool on some standard test data sets. You don't need to interact with many parties. NN is only technically hard. For GA4GH, things are quite different. Technically, it is hard, but it is much simpler than NN in my view. What is hard is 1) we lack a target and 2) the communications between developers and users. People don't know what we need, don't know what is the right approach and don't know how to evaluate the success. They have changed course back and forth, and still have clashes between programmers and those with more biological background.
- heuermh 11y agoCurious of your affiliation, if you're willing to provide it. I recently joined the Big Data Genomics team at UC Berkeley AMPLab, a GA4GH contributor. I was initially quite impressed with the GA4GH effort, because it was transparent and producing useful things, in terms of schema and code in github repositories. I am afraid now it has all gone working groups and private email threads and very little is happening in the open any more. I wish I understood semantic web. I've built more than one system on it but have never found a use case that sells it for me.