4 ms·
The gap that I see in current machine learning is that everyone is learning how to use the popular models, but no one knows how to construct a new model that so
by codingslave 7y ago
The gap that I see in current machine learning is that everyone is learning how to use the popular models, but no one knows how to construct a new model that solves a new problem. So everyone can download word vectors and use them for what they're good at, but the second you get off the beaten track, almost all machine learning practitioners fall flat. I really dont think this is due to how new the field is, rather that very few people have the mathematical maturity and experience to actually use optimization theory and linear algebra to construct new highly specific language models. There is simply no information available about doing this. You can learn about the underpinnings of matrix factorization and how that relates to word vectors, take that further and read about eigen vectors, but still, its too thin.
- M5x7wI3CmbEem10 7y agowouldn’t this just require a few textbooks in linear/non-linear optimization, mathematical modeling, and real analysis?
- codingslave 7y agoI think its way different than that, those would just be precursors, and in cases like real analysis, superfluous. Instead it would look something like, I have a Universal Sentence Encoder architecture, but its not performing well on my data, aside from tweaking the training set, how can I take this architecture and change it to work better with my individual problem? Assume here that the problem one is trying to solve is very similar to the models original use case. The number of people on the planet that can do this successfully, without wasting months of time messing around with tensorflow is extremely small. But this is where the value is. These massive catch all models only work for the people creating them, just jamming them into any NLP model will always produce sub par and probably unusable results
- rckoepke 7y agoI believe Google's AutoML is attempting to answer these types of questions. It's obviously internal-only so others can't fork the research...but it has helped them invent new specific networks like "EfficientNet for EdgeTPU" [0]. I think humans can still invent new macro structures like CNN's...but humans are inherently shit at analyzing "what if we removed one neuron in the 2nd hidden layer?". The subtle tweaking is really best left to an automated recursion process. Humans are better at seeing/inventing macro structures - such as adapting the unidirectional GPT to a bidirectional ELMO/BERT. After the invention, humans are generally pretty good at determining "whether" a network can be used to solve a particular task, although not infallible [1: Can BERT generate sentences from a prompt like GPT?] But computers are once again often better at quickly determining whether which (ELMO, BERT, or GPT) perform better on a particular task for which they are all at least feasibly suited. 0: http://ai.googleblog.com/2019/08/efficientnet-edgetpu-creating.html http://ai.googleblog.com/2019/08/efficientnet-edgetpu-creati... 1: https://ai.stackexchange.com/questions/9141/can-bert-be-used-for-sentence-generating-tasks https://ai.stackexchange.com/questions/9141/can-bert-be-used...
- codingslave 7y agoSure, but auto ML has not in most cases panned out. The ML name for auto ml is neural architecture search, which is mostly useless these days. NAS has shown to not be any better than standard random search across a neural network architecture. I do not say this to disparage googles results, only that they came up with the networks they did by expending huge amounts of computational power.
- Buttons840 7y agoWhat mathematical maturity went into designing the most popular models? What mathematical maturity lead to using ReLU instead of tanh for activation functions, for example? As far as I know, a lot (most?) advances in the field are just trying new ideas that happen to work. Is this correct?
- codingslave 7y agoSure, there's lots of trial and error. But consider something like Universal Sentence Encoders (USE) versus facebooks InferSent. USE is superior, mostly because its new, but still under performs infersent in a few areas. This is the kind of thing where actually building a more specialized model for question answering or textual similarity could see huge performance boosts for companies, but nobody is doing it. If they are, its under lock and key. Anyone looking to perform these tasks is mostly just pulling the code, tweaking the data, messing with the heads of the networks, and then calling it a day. EDIT: copying and pasting this from another answer: I think its way different than that, those would just be precursors, and in cases like real analysis, superfluous. Instead it would look something like, I have a Universal Sentence Encoder architecture, but its not performing well on my data, aside from tweaking the training set, how can I take this architecture and change it to work better with my individual problem? The number of people on the planet that can do this successfully, without wasting months of time messing around with tensorflow is extremely small. But this is where the value is. These massive catch all models only work for the people creating them, just jamming them into any NLP model will always produce sub par and probably unusable results
- denzil_correa 7y agoOne needs to form an experimental design with a ability to detect the challenge, understand the properties and come up with an appropriate computational solution to that challenge. This isn’t just for Machine Learning but for any kind of algorithm you develop.
- jononor 7y agoWhat kind of learning resources to use for this?