9 ms·
I made an automatic document tagger and categorizer. It collects any docs or HTML pages saved to Dropbox, dropped into a Telegram channel, saved with zotero, Sl
by kusmi 9y ago
I made an automatic document tagger and categorizer. It collects any docs or HTML pages saved to Dropbox, dropped into a Telegram channel, saved with zotero, Slack, Mattermost, private webdav, etc, cleans the docs, pulls the text, performs topic modeling, along with a bunch of other NLP stuff, then renames all docs into something meaningful, sorts docs into a custom directory structure where folder names match the topics discovered, tags docs with relevant keywords, and visually maps the documents as an interactive graph. Full text search for each doc via solr. HTML docs are converted to clean text PDFs after ads are removed. This 'knowledge base' is contained in a single ECMS, external accounts for data input are configured from a single yaml file. There's also a web scraper that takes crawl templates as json files and uploads data into the CMS as files to be parsed with the rest of the docs. The idea is to be able to save whatever you are reading right now with one click whether you are on your mobile or desktop, or if you are collaborating in a group, and have a single repository where all the organizing is done actively 24/7 with ML.
Currently reconstructing the entire thing to production spec, as an AWS AMI, perhaps later polished into a personal knowledge base saas where the cleaned and sorted content is public accessible with REST/cmis api.
This project has single handedly eaten almost a third of my life.
- echion 9y agoThis sounds really interesting -- can you share anything else, or pieces of the pipeline...especially topic modeling?
- kusmi 9y agoI use LDA algorithm for topic modeling. It has been the standard go-to for a while now within NLP community. There are implementations of it in many languages. The tricky part is cleaning the text, domain specific stopword lists, and in general controlling how text is processed depending on the context to make useful topic assignments when the text corpus represents more than a single field of knowledge. There are also some interesting ways of combining recent advances in RNNs on top of the more old school LDA topic modeling. I think this will be where most substantial advances will be coming from.