3 ms·
Yeah, this is what I ended up doing. I maintain a little project[1] that parses the CSV files and graphs your spending habits as well. I also wanted fuzzy stri
by amboar 11y ago
Yeah, this is what I ended up doing. I maintain a little project[1] that parses the CSV files and graphs your spending habits as well.
I also wanted fuzzy string matching, but I ended up writing my own clusterer in C to get the speed I wanted, and then wrapped it up as a python module. Must admit I hadn't considered difflib. The c code now lives in ccan[2]. I've got a patch to add the string vector cosine measurement as a filter which gives quite a big performance boost. As a rough indicator, on my laptop it clusters 3500 transaction descriptions in about 600ms.
[1] https://github.com/amboar/fpos/ https://github.com/amboar/fpos/
[2] http://ccodearchive.net/info/strgrp.html http://ccodearchive.net/info/strgrp.html