4 ms·
For those interested in related/alternative approaches, one or more of the following established open-source libraries might appeal to you: - Snorkel (training
by jointpdf 6y ago
For those interested in related/alternative approaches, one or more of the following established open-source libraries might appeal to you:
- Snorkel (training data curation, weak supervision, heuristic labeling functions, uncertainty sampling, relation extraction): https://github.com/snorkel-team/snorkel https://github.com/snorkel-team/snorkel
- AllenNLP (many pretrained NLP research models for tasks beyond text classification, model training and serving, visualization/interpretability utilities): https://github.com/allenai/allennlp https://github.com/allenai/allennlp
- Spacy (tokenization, NER/POS + tagging visualizer, pretrained word vectors, integration with DL models): https://github.com/allenai/allennlp https://github.com/allenai/allennlp
-huggingface Transformers (latest and greatest pretrained models, e.g. BERT): https://github.com/huggingface/transformers https://github.com/huggingface/transformers
- ...or a barebones “from scratch” solution in less than an hour with a Colab notebook and scikit-learn (preprocess text into tf-idf vectors, LSA/NMF to generate “document embeddings”, visualize embeddings with t-SNE/UMAP [facilitates weak supervision/active learning], classify with LogReg/RF/SVM/whatever). You could also tack on pretrained gensim/TF/PyTorch models quite easily as a next step. But this basic flow quickly gives you a handle on your corpus.
By the way, the docs for DeepDive (the predecessor of Snorkel) are some amazingly detailed background reading: http://deepdive.stanford.edu/example-spouse http://deepdive.stanford.edu/example-spouse
- Der_Einzige 6y agoKill the LSA/NMF middle-man and use UMAP directly. It supports sparse (tf-idf) vectors.
- jointpdf 6y agoTrue, good point. That may be better for classification performance. But at least for visualization and interpretability purposes using NMF is extremely simple and versatile (e.g. you can induce sparsity in the representation, setting the rank to be artificially low can cause high-level structure to “pop out”). That is, it gives you a few more knobs to turn than UMAP alone.
- nxpnsv 6y agoSpacy link is for allennlp!
- jointpdf 6y agoOops good catch. I should really stop writing posts on my phone. That’s especially sad since their docs are so good. I can’t edit it now, but: https://spacy.io/usage/spacy-101 https://spacy.io/usage/spacy-101 Bonus—An excellent interactive Spacy course from Ines Montani (also includes a template to build similar courses!): https://github.com/ines/spacy-course https://github.com/ines/spacy-course
- nl 6y agoHas anyone outside the Snorkel team done anything with it? I've tried multiple times (although mostly with DeepDive) and it was pretty complicated to get to do anything outside the demos. The Spacy link is here BTW: https://spacy.io/ https://spacy.io/