4 ms·
I've been going down the academic NLP rabbit hole lately, and at least in my domain (unsupervised key phrase extraction), the problem isn't the papers, it's the
by staticautomatic 6y ago
I've been going down the academic NLP rabbit hole lately, and at least in my domain (unsupervised key phrase extraction), the problem isn't the papers, it's the code (surprise!).
Let's start with the fact that in applied NLP, everyone has a plan until they get punched in the face by any number of pre-processing issues. And let's set aside the fact that in the end it's all going to regress to supervision, without which you can't optimize. Let's also set aside the fact that performance against a "gold standard" SemEval dataset doesn't mean shit in a lot of real world applications.
So you try out the standard issue "top of the line" algo, like YAKE, which is so fucking slow in pure Python that it'll choke a Bayesian optimizer. You sit around for a while debating whether or not to port it to Cython, having little idea if the effort will pay off because you aren't sure how well YAKE is going to work to begin with and it might get bested by another algo anyway.
So you go looking through the literature and you're delighted to find that within just the last few months, there have been some really cool and promising algos coming out with solid benchmarks, and there's code available to boot. Yay!
So you download the "weakly supervised" statistical one and it turns out to be a fucked up polyglot of Bash, C++, and a stale version of OpenJDK, some of which you have to compile yourself with g++, and then you have to dump your corpus into text files even though you've already got it in memory, run it through a tokenizer you neither want nor need, and then read the results back out of other text files. Sure, there's a docker version. It's full of bloat and solves some of the more negligible problems at hand.
Then you download a graph-based algo and it's such an undocumented mess of spaghetti it might as well have been written by an Italian restaurant. So you spend a really unreasonable amount of time just trying to figure out which function even takes your text as an input, and you read through a bunch of other functions trying to figure out if it needs to be pre-tokenized or not and if it wants the input as sentences or not or whatever. It also wants your input as a text file.
Then you download a language model-based algo and you think you're going to run the BERT variant you have at hand, but you double check the paper and it happens to perform way better with ELMO and then if you're lucky you don't spend a whole day trying to get AllenNLP running because you're using WSL on a laptop without a GPU and the non-gpu Tensorflow dependency is shitting itself all over the stack trace. You finally get the environment going in all its bloated glory even though you just wanted the pre-trained ELMO model, which you finally get deployed to Cortex or whatever and breathe a sigh of relief. And then it turns out your corpus is so domain specific that your matrix is sparser than swiss cheese because it's chock full of unks.
What have you learned after all this? That building an ensemble model which plays nicely with spaCy or SparkNLP is going to be an order of magnitude harder. Have fun!
- Der_Einzige 6y agoI wrote a summarizer which (when using the right settings) performs unsupervised key phrase extraction using language models. It is available here: https://github.com/Hellisotherpeople/CX_DB8 https://github.com/Hellisotherpeople/CX_DB8 It seems that it would be very useful to you. Like most data science code, it's non-trivial to install (it used to be when it was still updated)mostly because some dependencies are out of date and I will not risk a lawsuit from my current employer due to the similarity between this work and my day-to-day work. There is a jupyter notebook available which will allow you to use it without an install