2 ms·
Currently, we're text-mining the english version of Wikipedia. There's a lot of room for improvement: optimizing for speed and pruning down the results are at
by eserorg 18y ago
Currently, we're text-mining the english version of Wikipedia.
There's a lot of room for improvement: optimizing for speed and pruning down the results are at the top of our "TODO" list.
Also, the UI is simplistic -- that's because we've been spending 99% of our time working on the algorithm in matlab.
But, we wanted to get something out -- warts and all -- to get some feedback on the general idea.
We'd value any feedback -- positive or negative.
- gaika 18y agoPretty cool, what is the kind of math behind it? PLSA?
- eserorg 18y agoYes. At this point we're really constrained by the number of cores we're running on. Once we can get a hold of some more servers, we should be able to drastically improve the performance and prune many of the results. We'd also like to run the algorithm on additional corpora. Specifically: (1) the US patent database (back to 1975); and (2), a collection of United States federal and state case law (the JURIS database).