4 ms·
Using Topology to Classify Labelled Graphs
- chartpath 6y agoThis is a really cool and very smart looking mathematical way of looking at something that I do with lower-tech taxon labels using LTREE in Postgres to store the annotations. In PG you can do a kind of KNN for things that are X LTREE degrees away as well as ordering by distance from trigrams of text. With an ensemble of online learning models for different blobs of taxa, it is a way to prevent concept drift and give weights to the different subdomains of training data to fight bias. This kind of topological approach is great. One challenge in SQL is that class membership can only be so expressive and I have a nagging desire for slightly more RDF-like edges of predicates/verbs, but that still feels like overkill to me since I can already do rule engines with facts based on the labels. Therefore I end up with class names like "things that verb" or "things that are verbable" like would be done for naming interfaces. Scanning the paper I don't understand the math well enough (am more of a code plumber). How would different "types" of edges be encoded to get past the perceived limitations of class membership?
- Topolomancer 6y agoThat's an awesome application that I was totally unaware of! To encode more edge types, you would have to be able to devise a way to 'hash' them into a unique edge label. There's a variant of WL that collects edge labels instead of node labels; the rest of the algorithm remains unchanged. Is this roughly what you were looking for?
- chartpath 6y agoYeah totally! I have done a couple different LTREE label columns on the same table before but that was for separating named entity recognition from other things, so still just different kinds of class membership (e.g. people, places, things). It is possible to join these across tables, and therefore treat the matching nodes as having the same meaning for classification. However, it doesn't make sense to concatenate trees with different roots because it would just look like regular nesting. Even if I make it clear at the column level like resource_table.node_label::ltree and predicate_table.edge_label::ltree, there would still be no obvious way to parse where in the forest we traverse over nodes and edges, since they all look and function the same. I feel like this calls for a well thought-out naming convention using the only thing LTREE allows, alphanumeric and underscore characters. Something like root0.__node__name0.__edge__name0.__node__name1. This "merged" forest could be a "projection" within queries/views. Thanks again, happy topologizing :)
- nerdponx 6y agoMachine learning (and more generally modeling) on graphs seems really powerful. Interesting to see continued progress in this area.
- TACIXAT 6y agoI use WL in my work. Is there a digestible intro to persistent homology? Programmer level explanation (or code!), or should I just put my reading pants on and dive into the papers?
- Topolomancer 6y agoThere's a neat blog post written by one of my colleagues that might get you started: https://christian.bock.ml/posts/persistent_homology/ https://christian.bock.ml/posts/persistent_homology/ There's also a great article, written with biologists in mind, which does an excellent job in explaining what is going on: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7383827/ https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7383827/ In terms of programming, I think giotto-tda is a nicely-documented library: https://giotto-ai.github.io/gtda-docs/latest/ https://giotto-ai.github.io/gtda-docs/latest/ Message me for more!
- ent0py 6y agoAlways great to see people applying TDA. For anyone interested in experimenting with tools like barcodes (persistent homology), I'd recommend checking out https://scikit-tda.org/ https://scikit-tda.org/.