3 ms·
I have a good amount of experience in natural language processing and machine learning, and I don't think offering an API that provides easy access to the algor
by law 15y ago
I have a good amount of experience in natural language processing and machine learning, and I don't think offering an API that provides easy access to the algorithms is the right solution. The major algorithms in text classification aren't that complex to implement, and can be done in a few hundred lines. Moreover, all of the most widely used, widely tested, and reliable algorithms have public implementations that are readily adaptable to your needs. And that's the problem: understanding your needs.
Understanding your needs (or your company's needs) is where people with PhDs make their money. Machine learning isn't a panacea, and we won't be seeing a one-size-fits-all approach for awhile. Even though data has become more accessible, it might be noisy, incomplete, streaming, partially labeled, etc. This is why understanding exactly what you're trying to model with these algorithms is crucial and why "just applying" them is impractical at best and misleading at worst.
- St-Clock 15y agoThe other big catch is randomness, which is often not understood by neophytes. If you try to find relations in sufficiently large data set, you're bound to find some that are "caused" by randomness. Tools like p-values are of little help when you fish for many relations (and not just one in particular).
- law 15y agoRandomness can be a curse, but can also be a blessing when introduced as in the random subspace methods. This again abstracts to understanding your business needs and whether the results encountered make sense given the features' [absence of] independence. An API giving you a wide choice of algorithms will still rely on you to run something like ICA as a pre-processing step to identify this statistically independent randomness.
- PaulHoule 15y agoWell. Here's my take. There are a number of text analysis SaaS offerings such as OpenCalais, AlchemyAPI, Zemanta, and OpenAmplify. They've all got impressive science under the hood, but none of them are accurate enough to be useful. I spend most of my time these days thinking about why that is and what to do about it. For systems to do better, they'll need to incorporate world knowledge; they'll need to test different interpretations of a text and select the ones that "make sense". This is likely to be a form of statistical inference rather than Cyc style logic. Based on some systems I've worked with, I'd estimate that a space optimized "background" knowledge base that can estimate satisfiability in the common sense domain is on the order of 10-100 GB. It will puff out to at least an order of magnitude beyond that in the process of creating it. Few users will have the ability to create a KB of that type, and it would be a serious thing to download and install. Hosting the services of that kind of system in a SaaS manner makes a lot of sense.
- zeratul 15y agoWithout looking under the hood I'd say there could be at least four reasons why they fail (based on what most of the NLP literature is lacking): - did not remove contradicting information from the training sets (two very similar vectors having contradicting labels) - did not try enough feature selection algorithms - did not estimate ALL learner parameters using the training sets with internal CV - did not include domain knowledge The last one refers to Paul Houle comment. Just, beside using tools like OpenCyc, WordNet, UMLS, there many other ways to embed domain expertise in an automated classification process. Injecting semantically related features into a vector representation of a document is extremely difficult. Forward feature selection doesn't work well for sparse and noisy data.
- PaulHoule 15y agothe curse of dimensionality is the worst problem that affects machine learning customers don't want to create training sets large enough to train text classifiers; often the number of documents they need to sort into a category is too small to fit in a category. As for semantic indexing, it was hard to do in 2005. In 2011 it's easy. DBpedia and Freebase are a chromosome map for the human memome. With large amounts of instance information, it's possible to do things that a big rulebox can't. These tools are aiming for the market segment that Cyc aimed for, but will use very different methodologies.