5 ms·
Except this doesn't work consistently because words in real language are ambiguous and not every combination results in something that can meaningfully map to r
by pluma 7y ago
Except this doesn't work consistently because words in real language are ambiguous and not every combination results in something that can meaningfully map to real-world language.
You need to start from a word list with 100% unambiguous and clearly defined words and even then you're no step closer to working with real language because while superficially similar your word list is actually a highly specialised DSL.
Of course in many cases this DSL approximation of the target language is good enough for certain tasks but the entire process is inherently flawed.
- brudgers 7y agoI agree: king - manliness is nonsensical in the context of chess.
- tonyarkles 7y agoThat's a really good point, and related to my sibling comment, one of the other things that makes me uncomfortable about ML models is the mysterious generalizability. Some generalize well, some start spewing nonsense when you move slightly outside of their training zone. Special-purpose models, I think, do quite a bit better than attempted larger generalized models. For your chess example, if you trained a word2vec model using only a large corpus of text about chess, you very likely wouldn't get the "king - manliness" vector being anything meaningful at all, but you would likely see word associations that are meaningful and also potentially unexpected.
- tonyarkles 7y ago> approximation of the target language is good enough for certain tasks but the entire process is inherently flawed That is pretty much the definition of a "model" :) I recently went through the "Tensorflow in Practice" specialization on Coursera and it was illuminating. The thing about ML models, whether CNNs for images, or word2vec+RNN, or whatever else, is that they really don't have any rigid scientific basis for why they work. You're doing, say, Stochastic Gradient Descent to optimize the neuron weights across your dataset. Out the other side of the training, you have a mostly meaningless set of coefficients that work well to classify other unseen data. I dual-majored in CS and EE, and I leaned towards the "science" side of things, where things get modelled mathematically and analyzed, accepting that the model is likely incomplete but still useful. The thing that drives me nuts with ML is that there's no explanation of what the terms in the ML model actually mean (because the process that produced them doesn't actually investigate meaning, it just optimizes the terms). But... I've accepted that even though the models are pretty much semantically meaningless, they work.
- amelius 7y ago> I've accepted that even though the models are pretty much semantically meaningless, they work. Until they don't. Which may happen easily if you deploy a model for the first time. My personal view is that (at this moment) ML is mostly correlation detection and pattern recognition, but has little to do with intelligence.
- bitL 7y agoThe point is that we don't have mental capacity to understand this stuff. Nobody has any clue how to interpret millions of dimensions, some non-linear manifold there and how to translate it to something humans are capable of understanding. These things might be done automatically by our brains on subconscious level in a similar fashion (or not), but on conscious level we are completely clueless and basically shoot darts to see which ones become somewhat useful. I think you object to the lack of "mathematical beauty", but my point is "who cares?". Not sure why should reality conform to some mental model we find "appealing" for whatever reason. Deep Learning is similar to experimental physics.
- woliveirajr 7y agoThis. Explainable AI is a emerging field, I hear about this necessity specially in NLP and Law. We expect to understand how some decision was reached, and we'll never accept some computer-generated decision if it wasn't explained how each logical step was done. And just giving millions of weights of each neuron won't give us that, because we won't be able to reach the same decision with just those parameters. We know that IA is a bunch of probabilities, weights and relations in n-dimensions. Our rational brain can know that too, but can't feel it.
- Der_Einzige 7y agoThat's why you use interpretability tools like LIME Example of this would be here: https://github.com/Hellisotherpeople/Active-Explainable-Classification https://github.com/Hellisotherpeople/Active-Explainable-Clas...
- fromthestart 7y agoI'm afraid you misunderstand the way embeddings work - at least for BERT based models, which are currently state of the art. BERT embeddings, after training change with context. In other words if you feed a paragraph about bank robbers and look at the encoding for bank, it will be meaningfully different from the encoding for the same word produced from a paragraph (or sentence) about river banks. We use BERT at the startup I work at, and one of our tests was the sentence "the bank robbers robbed the bank and then rested by the river bank". BERT was able to generate three different semantically meaningful encodings for the word bank in this sentence. The first two instances were much closer to each other in vector space (euclidean distance) than the last. This is huge, because it is arguably the first step in building an AI which can perform basic reasoning about information encoded in text. For example, if you average up the encodings of a paragraph of words, you can create an "encoding" which assigns a summary meaning or topic. Simple vector math becomes a powerful reasoning tool. The future is here.
- yorwba 7y agoThe contextualized word embeddings you get out of BERT are still generated from fixed per-word vectors. And while you get one output vector for each input vector, that doesn't mean they correspond to each other. The model could arbitrarily reshuffle information between outputs, so long as the output as a whole reflects the input sufficiently well. So BERT embeddings are not "word embeddings" in the usual sense.
- thelazydogsback 7y ago> This is huge, because it is arguably the first step in building an AI which can perform basic reasoning about information encoded in text. Well, except for the many many decades of previous work on NLP using symbolic methods that are quite capable. Although DNNs are en vogue and have some amazing properties, we shouldn't forget that symbolic AI/NLP using explicitly semantic representations is powerful and has a rich history, and complements DNNs quite well -- such as being easily explainable, for one.