4 ms·
The previous paper they mention explains the core insight that makes unsupervised translation possible: https://arxiv.org/abs/1710.04087 https://arxiv.org/abs/1
by heydenberk 8y ago
The previous paper they mention explains the core insight that makes unsupervised translation possible: https://arxiv.org/abs/1710.04087 https://arxiv.org/abs/1710.04087
The original paper didn't receive the attention I thought it would, but I continue to think this is a fascinating result which has deep implications for machine learning and for linguistics.
- schoen 8y agoThese word embeddings keep on yielding all kinds of amazing benefits. Is there any kind of explainability research to help people understand them better in terms of human psychology?
- yorwba 8y agoWord embeddings work because they reflect co-occurrences. I don't know whether that counts as an explanation in terms of psychology, but humans tend to put related things together. In a newspaper the articles aren't jumbled together, but there are sections on different topics, and in each section the articles are clearly delineated instead of mixing their sentences and each sentence represents a single unit instead of giving partial information on a dozen unrelated things. It might seem obvious that things should be done that way, but if you consider servers hosting lots of different websites on the same physical machine, or data structures spread out over several memory allocations held together by pointers, it's clear that there are other possibilities. So it does seem to be specific to the way humans use language. And because human language has this property of co-occurrences corresponding to relatedness in meaning, you can represent the meaning of a word by building a model that only predicts the probability that two words occur together.
- PeterisP 8y agoIn linguistics, there's a classic principle "You shall know a word by the company it keeps" (Firth, J. R. 1957) - colocations (a sequence of words or terms that co-occur more often than would be expected by chance) are very informative about what a word means.
- dmreedy 8y agoNot so much in psych, where language remains a pretty fundamental mystery, but from the philosophy of language side, word embeddings are, I think it's fair to say, fundamentally an implementation of Sassurian structuralism and semiotics. Words (signs) have no intrinsic, native meaning. And so, the only way to figure out which one is which is by measuring it in relation to all other signs in the lexicon; that is to say, words only mean something because they don't mean all the other things, and that the Structure as a whole is what provides meaning. There has been much in the way of discussion for, and concern about (for example, the Deconstructionist movement), these ideas, for the past 60 years or so. And a bit of practical exploration in the field of child development. Interestingly, the fact that embeddings between languages seem to share some common shapes (per the linked paper in this thread), would seem to suggest that A) Fundamentally, most languages have the same deep structure, whether through coincidence or common evolutionary root. or B) The brain has a hard-wired structure for language, evolved alongside the development of language itself. The Chomskian Language Acquisition Device. We're not born tabula rasa, we've got some hardcoding indicating how we're going to understand things Or a little of A, a little of B maybe, as it does end up being a boostrapping problem.
- shawn 8y agoThis is genius: https://imgur.com/a/1aRZ3sI https://imgur.com/a/1aRZ3sI Normally this technique wouldn't be useful, because it's overfitting a specific training set. (If you make space X as similar as possible to space Y, then this mapping from X to Y is only useful for X to Y – it can't generalize to other situations, which is often the goal of an ML model.) But since the task is "Translate from English to Italian," and since all languages have similar embedding structures (Zipf mystery), overfitting is exactly what we want: we want, for any given English phrase, to find the closest-fitting mapping to a corresponding Italian phrase. The more I learn about ML and data mining, the more I'm astounded by how clever many of the techniques are, and how much artistry is involved. You have to be clever to make a certain model perform well in a certain domain. If you want to make a stock trading bot, you can't randomly subdivide stock market data into e.g. 70% training data and 30% test data, because the data is ordered by time. You have to use the past 3 years of stock market data as a training set, and validate it against the subsequent 1 year of market data. I really like ML because the techniques applicable for training a stock market bot seem unrelated to the algorithms for doing unsupervised machine translation, which differ from how to model credit fraud, which are no doubt different from how to build a dota 2 bot. :)
- LittlePeter 8y agoSplitting the time series data into training and test datasets by making a single cut is not some artistry or epiphany following years of research. It is common sense if you work in time series domain.