4 ms·
The second layer is hard. I tried something in this space in mid-2018. Full text extraction and sentence segmentation tech was adequate, but extracting the disc
by stonerri 4y ago
The second layer is hard. I tried something in this space in mid-2018. Full text extraction and sentence segmentation tech was adequate, but extracting the discourse tree and building the graph was a bit of a struggle (trying to repurpose a collection of academic/open tools to get something useful). Never published or released the code.
If interested, a few rabbit holes to explore (no affiliations):
https://scite.ai https://scite.ai -> best option for citation mapping, but same issues you described above
https://www.semanticscholar.org https://www.semanticscholar.org and AI2 -> the best group working on tooling in this space
https://www.weave.bio https://www.weave.bio -> early startup trying to build this out
The hardest challenge in my view is solving the intermediate representation issue. You have to establish a DSL/nomenclature that provides the range required to represent a complete scholastic discourse while also being computable.
- 21eleven 4y ago> The hardest challenge in my view is solving the intermediate representation issue. Right, you'd basically be writing an interpreter for English
- nerpderp82 4y agoHave GPT translate down into a Controlled Natural Language. I tried having it translate to OWL, but it sucked. https://en.wikipedia.org/wiki/Controlled_natural_language https://en.wikipedia.org/wiki/Controlled_natural_language
- uticus 4y ago“OWL”?
- xyzzy3000 4y agoI assumed OWL='Web Ontology Language', which seems to make sense given that the topic is semantics - but OWL isn't a controlled vocabulary of a natural language, so I may be wrong.
- nerpderp82 4y agoYou are correct, I just reached for the nearest machine readable knowledge format.