3 ms·
This is a really cool and very smart looking mathematical way of looking at something that I do with lower-tech taxon labels using LTREE in Postgres to store th
by chartpath 6y ago
This is a really cool and very smart looking mathematical way of looking at something that I do with lower-tech taxon labels using LTREE in Postgres to store the annotations.
In PG you can do a kind of KNN for things that are X LTREE degrees away as well as ordering by distance from trigrams of text. With an ensemble of online learning models for different blobs of taxa, it is a way to prevent concept drift and give weights to the different subdomains of training data to fight bias.
This kind of topological approach is great. One challenge in SQL is that class membership can only be so expressive and I have a nagging desire for slightly more RDF-like edges of predicates/verbs, but that still feels like overkill to me since I can already do rule engines with facts based on the labels. Therefore I end up with class names like "things that verb" or "things that are verbable" like would be done for naming interfaces.
Scanning the paper I don't understand the math well enough (am more of a code plumber). How would different "types" of edges be encoded to get past the perceived limitations of class membership?
- Topolomancer 6y agoThat's an awesome application that I was totally unaware of! To encode more edge types, you would have to be able to devise a way to 'hash' them into a unique edge label. There's a variant of WL that collects edge labels instead of node labels; the rest of the algorithm remains unchanged. Is this roughly what you were looking for?
- chartpath 6y agoYeah totally! I have done a couple different LTREE label columns on the same table before but that was for separating named entity recognition from other things, so still just different kinds of class membership (e.g. people, places, things). It is possible to join these across tables, and therefore treat the matching nodes as having the same meaning for classification. However, it doesn't make sense to concatenate trees with different roots because it would just look like regular nesting. Even if I make it clear at the column level like resource_table.node_label::ltree and predicate_table.edge_label::ltree, there would still be no obvious way to parse where in the forest we traverse over nodes and edges, since they all look and function the same. I feel like this calls for a well thought-out naming convention using the only thing LTREE allows, alphanumeric and underscore characters. Something like root0.__node__name0.__edge__name0.__node__name1. This "merged" forest could be a "projection" within queries/views. Thanks again, happy topologizing :)