4 ms·
Thanks. I skimmed the linked "Mapper" article they cite as their method, and it looks about as topological as t-SNE, i.e. in the sense of caring about local nea
by improbable22 8y ago
Thanks. I skimmed the linked "Mapper" article they cite as their method, and it looks about as topological as t-SNE, i.e. in the sense of caring about local nearness but not global distance.
But did I miss any heavier stuff? Is this all people mean when they talk about topological data analysis?
- ajudson 8y agoThe other famous TDA technique is persistent homology, this paper is a good intro https://www.math.upenn.edu/~ghrist/preprints/barcodes.pdf https://www.math.upenn.edu/~ghrist/preprints/barcodes.pdf
- soVeryTired 8y agoBut now: persistent homology. Same question as above.
- kaitai 8y agoProgress has been slower than expected, in some ways. But I think that persistent homology and topological methods like Mapper are slowly allowing interesting research. My favorite recent application: two-parameter persistent homology for drug discovery (link to pdf of talk: https://www.ima.umn.edu/materials/2017-2018/SW8.13-15.18/27458/Bryn-Keller_Two-Parameter-Persistence-for-Virtual-Ligand-Screening.pdf https://www.ima.umn.edu/materials/2017-2018/SW8.13-15.18/274...). Frankly I find two-parameter persistence very hard to interpret, though I hope to spend some time on it this summer. But it's undeniable that the work described in the link is an application that can't be reduced to clustering. I also feel that describing Mapper as 'just a clustering tool' is not really accurate; I'm getting good results from using Mapper for feature discovery in some specific domains where clustering has truly failed because traits lie on continua in a way that renders usual clustering muddy and useless. Edited to add: here's another interesting paper, about subgroups of type 2 diabetes patients: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4780757/ https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4780757/ They use Mapper to find 'clusters' and then go back to more traditional statistics to test significance. I don't have the data so I don't know if clustering alone would have worked.
- llamaz 8y agoI'm currently doing a thesis using TDA. Excuse my shameless opportunism, but the topic I've come up with is using figure h [1] as a channel selection mechanism to select EEG channels [2]. I plan to then use a Gaussian process (similar to an SVM) to detect the time when a seizure occurs. From your position of expertise, can you see any major problems off the top of your head? [1] https://media.springernature.com/lw785/springer-static/image/art%3A10.1186%2Fs13104-018-3482-7/MediaObjects/13104_2018_3482_Fig3_HTML.png https://media.springernature.com/lw785/springer-static/image... - from the open-access paper https://link.springer.com/article/10.1186/s13104-018-3482-7 https://link.springer.com/article/10.1186/s13104-018-3482-7 [2] quote: "Step III... we applied the Vietoris–Rips filtration to determine which particular sensors (thus, which areas of the brain) are more “involved”∖“significant” concerning the spreading of epileptic seizures"