4 ms·
found the JS one.. https://github.com/google-research/google-research/tree/master/jslm https://github.com/google-research/google-research/tree/mast...
by willwade 2y ago
found the JS one..
https://github.com/google-research/google-research/tree/master/jslm https://github.com/google-research/google-research/tree/mast...
- _akhe 2y agoVery nice, thanks for sharing! I noticed their token sequence length is fixed: Constructing 7-gram LM ... Created trie with 21,502,513 nodes. People on here will be happy to say that I do a similar thing, however my sequence length is dynamic because I also use a 2nd data structure - I'll use pretentious academic speak: I use a simple bigram LM (2-gram) for single next-word likeliness and separately a trie that models all words and phrases (so, n-gram). Not sure how many total nodes because sentence lengths vary in training data, but there are about 200,000 entry points (keys) so probably about 2-10 million total nodes in the default setup. "Constructing 7-gram LM": They likely started with bigrams (what I use) which only tells you the next word based on 1 word given, and thought to increase accuracy by modeling out more words in a sequence, and eventually let the user (developer) pass in any amount they want to model (https://github.com/google-research/google-research/blob/5c87ca575097f11f1763ebed80bd5066aaa98e9e/jslm/language_model_driver.js#L51 https://github.com/google-research/google-research/blob/5c87...). I thought of this too at first, but I actually got more accuracy (and speed) out of just keeping them as bigrams and making a totally separate structure that models out an n-gram of all phrases (e.g. could be a 24-token long sequence or 100+ tokens etc. I model it all) and if that phrase is found, then I just get the bigram assumption of the last token of the phrase. This works better when the training data is more diverse (for a very generic model), but theirs would probably outperform mine on accuracy when the training data has a lot of nearly identical sentences that only change wildly toward the end - I don't find this pattern in typical data though, maybe for certain coding and other tasks there are those patterns though. But because it's not dynamic and they make you provide that number, even a low number (any phrase longer than 2 words) - theirs will always have to do more lookup work than with simple bigrams and they're also limited by that fixed number as far as accuracy. I wonder how scalable that is - if I need to train on occasional ~100-word long sentences but also (and mostly) just ~3-word long sentences, I guess I set this to 100 and have a mostly "undefined" trie. I also thought of the name "LMJS", theirs is "jslm" :) but I went with simply "next-token-prediction" because that's what it ultimately does as a library. I don't know what theirs is really designed for other than proving a concept. Most of their code files are actually comments and hypothetical scenarios. I recently added a browser example showing simple autocomplete using my library: https://github.com/bennyschmidt/next-token-prediction/tree/master/examples/ui-autocomplete https://github.com/bennyschmidt/next-token-prediction/tree/m... (video) And next I'm implementing 8-dimensional embeddings that are converted to normalized vectors between 0-1 to see if doing math on them does anything useful beyond similarity, right now they look like this: [nextFrequency, prevalence, specificity, length, firstLetter, lastLetter, firstVowel, lastVowel] For example, here's the embedding for "dog": [0.6666666666666666, 0.6, 0.45714285714285713, 0.15, 0.12, 0.24, 0.56, 0.56] where it used to just be 1D: { dog: 23 } Up until now I've only effectively used nextFrequency to predict the next word, but with these other 7 metrics I'm able to store more info about each word and being vectors am able to store it efficiently, and do what all the smart people keep saying like dot product similarity index etc. The big boys are using like 1536D arrays etc. but with just 7 it still works, and it's weird that it works. Thanks again for sharing!
- willwade 2y agoTheir library was actually made for dasher.. http://www.inference.org.uk/dasher/ http://www.inference.org.uk/dasher/ - there was a web version being made (https://github.com/dasher-project/dasher-web https://github.com/dasher-project/dasher-web We hit a bottleneck with the graphics driving. Note in dasher pretty much the entire tree is in dynamic view). Now this may help to understand the use case. Dasher is for people with disabilities who cant speak. It needs to be a personalised LM that trains on the fly and and keeps track of new words/sentences. But in truth too, utterances are usually small. Don't get too knocked back by comments. A) If it works - it works. B) Your learning is as valuable as the outcome. Oh have a look at https://imagineville.org/software/ https://imagineville.org/software/ for some other things that may be of interest..
- _akhe 2y agoThe visual tool on their site is trippy! Very cool. I appreciate the words of wisdom/motivation :) Since my last comment the embedding vector is now 16-dimensions, but I'm not quite getting "King - Woman = Queen" from it yet. It actually does find similar words if I score word features, normalize the scores to a min/max spectrum between 0-1, then multiply 2 vectors (dot products) to get similiarity. But in the middle of the experiment I was like "what is this for..?" so I just kinda stopped working on the embedding refactor until I have a real world use case. I know for example text-embedding-ada-002 uses 1536 length arrays, and others use over 3k, so maybe at some point the usefulness of having very complex embeddings emerges. For now, the original approach still seems superior for next token prediction, bigram frequency (how often a word follows another) just need to have enough sentences modeled and scored.