3 ms·
i think they don't. given enough data those related forms will have similar colocation profiles (aka fingerprints). not a model of semantics.. at scale it might
by espe 3y ago
i think they don't. given enough data those related forms will have similar colocation profiles (aka fingerprints). not a model of semantics.. at scale it might become a proxy though.
- 19h 3y agoYes, we don't actively cover this. Our training dataset is composed of 12 different languages. The English corpus has a size of around 25TB -- ~18TB after a rather aggressive filtering process (char frequencies [i.e. https://gist.github.com/19h/b983769b7ec9c4a9528377d7819892a7 https://gist.github.com/19h/b983769b7ec9c4a9528377d7819892a7], too many unusual features in the sentences, too short sentences, etc.). Most common misspellings of words have nearly identical fingerprints, and we align the retinas across languages so that translated sentences result in similar fingerprints (via competitive learning).
- espe 3y ago"retina" describes the orientation of the vector space? first time i hear it in this context.
- 19h 3y agoApologies, that's a very semantic folding specific term -- with retina I refer to a word -> fingerprint database that gives you either fingerprints of single words or of blocks of text (by applying relevant heuristics to filter out irrelevant noise / heavily used words in a specific language).