3 ms·
very efficient but also brittle. that must be vast amounts of relatively clean data. you have to magically set the number of top n words to in- and exclude. for
by espe 3y ago
very efficient but also brittle. that must be vast amounts of relatively clean data. you have to magically set the number of top n words to in- and exclude. for most user generated content one would need to heavily normalize the text, e.g. by stemming (to keep in line with the computational austerity). 16384 is very little even if it is neatly seperated concepts. applied to that volume of data it should amount to keyword matching.. that only works if users are basically self-tagging their texts via constrained language use.
edit: short version: not semantics and not a fingerprint :)
- 19h 3y agoWe also trained on all of pushshift and have an average ”unknown” word rate of less than 0.007% — the Reddit corpus is rather amazing to capture pretty much all misspellings of a word. We may only be using 16k vector values but that doesn’t mean we only have a vocab of 16k —- our vocab is more around 1.9 million words each described by a sparse fingerprint of 16k.
- espe 3y agothanks for the clarification. if your base population is that large then it's frequencies and you get a fingerprint. well done.
- foolswisdom 3y agoI'm curious though, how do you handle related forms of a word (assuming you don't use stemming)? It doesn't seem to me that this process would automatically handle that.
- espe 3y agoi think they don't. given enough data those related forms will have similar colocation profiles (aka fingerprints). not a model of semantics.. at scale it might become a proxy though.
- 19h 3y agoYes, we don't actively cover this. Our training dataset is composed of 12 different languages. The English corpus has a size of around 25TB -- ~18TB after a rather aggressive filtering process (char frequencies [i.e. https://gist.github.com/19h/b983769b7ec9c4a9528377d7819892a7 https://gist.github.com/19h/b983769b7ec9c4a9528377d7819892a7], too many unusual features in the sentences, too short sentences, etc.). Most common misspellings of words have nearly identical fingerprints, and we align the retinas across languages so that translated sentences result in similar fingerprints (via competitive learning).
- espe 3y ago"retina" describes the orientation of the vector space? first time i hear it in this context.
- 19h 3y agoApologies, that's a very semantic folding specific term -- with retina I refer to a word -> fingerprint database that gives you either fingerprints of single words or of blocks of text (by applying relevant heuristics to filter out irrelevant noise / heavily used words in a specific language).