3 ms·
There's a good Google Books dataset for doing exactly that, and you can doing something in Hive to get you the lowest-entropy n-grams (with something like https
by lsb 13y ago
There's a good Google Books dataset for doing exactly that, and you can doing something in Hive to get you the lowest-entropy n-grams (with something like https://github.com/lsb/text-entropy/blob/master/passphrase-safety.hiveql https://github.com/lsb/text-entropy/blob/master/passphrase-s... ), and then you can the low-entropy n-grams into a Bloom filter (like http://www.leebutterman.com/passphrase-safety/how-it-works.html http://www.leebutterman.com/passphrase-safety/how-it-works.h... ) and your enormous corpus gets fitted into a few dozen megs of memory.
(That sounds like a cool idea, hit me up if you're game for hacking on something like that)
- sitkack 13y agoMight also help to get feedback from the user and to know what n-grams the user has already encountered in their life.
- mallamanis 13y agoI've been using n-grams heavily in my PhD. I like a lot what you are describing. Also a simple extension is to use "cache" n-gram models, that can naively simulate longer-term "memory". What's fascinating is that in the entropy/information-theoretic sense, instead of using wpm, one could aim for "constant" information rate of reading (as in bits/token). That's really cool as a concept. I don't know if it would work, obviously. I'd definitely be up to for hacking this up, but currently a bit too busy. Will ping you at some point though!
- pkghost 13y agoYou're welcome to fork it! http://www.github.com/cameron/squirt http://www.github.com/cameron/squirt. Would love to see such an implementation