5 ms·
I am very supportive of the idea in general. I think that the author underestimates some of the reasons for the diversity in the data sources schemas etc... The
by jerven 11y ago
I am very supportive of the idea in general. I think that the author underestimates some of the reasons for the diversity in the data sources schemas etc... They are different often because of the data being different.
The linked data/semantic web approaches are slowly eating away at the unneeded data diversity. With small shared standards for common data, and unique schemas for unique data. The single endpoint solution unfortunately does not scale; it would be the AOL of biomedical data or entrez as you prefer. To be truly open and free anyone should be able to contribute their data and tools. This means decentralized infrastructure which means confusion and difficult to find information. However, as acedemia and research is decentralized their IT infrastructure must match that reality. This leads to infrastructure that can integrate on demand, such as made possible by the SPARQL service keyword.
Open source has always been an important part of bio-IT and that is not going to change. But a single source is not the answer to our problems. We need to make it easier to find information, but most importantly need to make it easier to answer questions with the data that is available.
- crypt1d 11y agoI have no knowledge of the field, so apologizes in advance if this idea does not make sense. Would it be possible to make a compromise here? Perhaps by maintaining the decentralized structure, but at the same time introducing standards to categorize data on a global level and allow it to be 'mined' by a centralized entity that has API access to this data. Kind of like Google does indexing and searching, just with more cooperation of those being indexed.
- nikolamilosevic 11y agoYes, this is kinda what I wanted to propose. An umbrella organization that would make an infrastructure using which it would be possible to query, index and integrate data acros the web. And I am not running away from decentralized structure or infrastructure.
- crypt1d 11y agoGlad to see that I got the idea after all. As I've said I'm no expert, but I'd be interested to use a project like this to learn more about the field and to find ways to contribute. Atm, I can help setup the initial infrastructure for the project and cover the AWS costs for the first few months until you get some funding. Feel free to reach out if you think this could help, email is in my profile.
- jerven 11y agoHave you ever looked at the sadiframework? The HCLS W3C note on data set descriptions is also a good starting point.
- nikolamilosevic 11y agoLinked data and SPARQL are definitely very possible solution and infrastructure can be decentralized. There is a bit of resistance in one part of the community from these technologies, because people are not used to them and compared to some other data storages they tend to be a bit slower, but thats other discussing. I do not have anything against these technoogies. What I currently don't like is that there are a lot of resources that are technically open source and free, but they are burried somewhere on the internet and sometimes hard to find and it takes quite a lot of time to review all existing resources. What I wanted to recommend is one central umbrella organization that will be (1) platform for collaboration in biomedical field, (2) central endpoint to all major existing project, possibly with some maturity level of projects and internal review in order to arrange projects into maturity levels, so it can be relatively easy to review how much you can "trust" that project of data, (3) central repository for open source NLP, data curation and semantic web tools, (4) some relevant body that would be able to propose and work on standards for data curation that would take in account all field specific needs.
- x1k 11y agoYou have seriously underestimated 1) efforts needed to develop and to maintain such resources – your best hope is to work with government-funded institutes; 2) resistance from the convention of a particular research field – you can rarely bend how people in a field work on things; 3) culture differences between biologists/doctors and programmers – biologists/doctors think very differently, which is frequently overlooked by programmers; 4) bureaucracy – everyone thinks he/she is the best; when you work with top groups to make things happen, you will find how problematic it is; 5) technical challenges – as you care about pheonotype data: there are no good ways to integrate various pheonotypes from multiple sources. Everyone in biomedical research dreams about integrated resources. I have heard multiple people advocating SPARQL as well. If it had been that easy, this would have occurred years ago. In the real world, no one is even close. If you want to attract collaborators, learn Linus: say you have a working prototype and demonstrate how wonderful it is. Your ideas are cheap. The difficult part is a clear roadmap to make it happen.
- dekhn 11y ago