3 ms·
I worked for a company that tried to do exactly what this article is proposing, I was responsible for parsing the data from publications into exactly this sort
by fudgefactorfive 4y ago
I worked for a company that tried to do exactly what this article is proposing, I was responsible for parsing the data from publications into exactly this sort of table.
The primary reason is that it is very hard to come up with a schema that even 5% of papers would adhere to. The vast majority of this knowledge is phrased as natural language.
There are databases that track compounds and the publications related to them, but those papers again are natural language and cannot be readily converted to tabular data. Our first basic approach involved POS tagging and then trying to associate proper nouns with numeric values. Again the issue became how do you interpret a sentence like "may lead to sudden death" as a symptom? Something like "may lead to symptom X in Y% of respondents" is a nightmare to consistently parse without heavy ML running over huge datasets of just text.
In the end we wound up having to shut down concluding that until papers are released with not only arbitrary XML tables/results we were not equipped to handle the task. And even worse, what if our models didn't interpret things correctly and a consumer got {symptom:"sudden death", chance:0%} and insisted on that compound for their indication only to later realize the paper stated "in lab setting 0% of animals didn't experience sudden death after being administered X after diabetes diagnosis". Paying a hundred students to work around the clock couldn't get the volume processed accurately for months, let alone getting a second army of validators to confirm each entry.
- fudgefactorfive 4y agoIn the end the only functional parts that worked algorithmicly are exactly those featured in the GlacierMD demo. What trials are running related to this indication, what compounds are being tested for the indication and what other indications are related. That's the easy part, it's effectively a word association game, TF-IDF did this job admirably, scoring proper nouns by their uniqueness and then associating them with one another and searching for publications with similar words as the requested indication. Effectively a medical word cloud for each indication and compound. Parsing them into symptoms is the first nightmare, the second is numeric values associated with those symptoms and paper results. There is a very good reason the demo only has one indication and a handful of symptoms, it's being done manually and then at best showing publications related to the words encountered. It's not a matter of cost, although the author is all but doomed if they want to cover more than a few indications, it's a matter of not forcing publicly funded health publications to use an electronically parseable Format despite the simplicity of them being able to parse their paper by definition. See the standards XKCD, the issue is getting many different academics and departments to agree on a set of schema to include alongside their publications. PubMed at least tries with their XML dumps but even those are inconsistent at best and non-syntactically interpretable at worst. The Japanese compound tracker is great to learn about a specific compound and their indications but stops there.
- intrasight 4y agoJust make it a condition of funding and they'll probably get on board with machine-readable standards. But the first thing the feds will have to do is fund and do a big competition to define those standards.