3 ms·
In the end the only functional parts that worked algorithmicly are exactly those featured in the GlacierMD demo. What trials are running related to this indicat
by fudgefactorfive 4y ago
In the end the only functional parts that worked algorithmicly are exactly those featured in the GlacierMD demo. What trials are running related to this indication, what compounds are being tested for the indication and what other indications are related.
That's the easy part, it's effectively a word association game, TF-IDF did this job admirably, scoring proper nouns by their uniqueness and then associating them with one another and searching for publications with similar words as the requested indication. Effectively a medical word cloud for each indication and compound. Parsing them into symptoms is the first nightmare, the second is numeric values associated with those symptoms and paper results.
There is a very good reason the demo only has one indication and a handful of symptoms, it's being done manually and then at best showing publications related to the words encountered.
It's not a matter of cost, although the author is all but doomed if they want to cover more than a few indications, it's a matter of not forcing publicly funded health publications to use an electronically parseable Format despite the simplicity of them being able to parse their paper by definition.
See the standards XKCD, the issue is getting many different academics and departments to agree on a set of schema to include alongside their publications. PubMed at least tries with their XML dumps but even those are inconsistent at best and non-syntactically interpretable at worst. The Japanese compound tracker is great to learn about a specific compound and their indications but stops there.
- intrasight 4y agoJust make it a condition of funding and they'll probably get on board with machine-readable standards. But the first thing the feds will have to do is fund and do a big competition to define those standards.