4 ms·
Love stuff like this. Tangentially I'm working on useful language models without taking the LLM approach: Next-token prediction: https://github.com/bennyschmi
by bschmidt1 2y ago
Love stuff like this. Tangentially I'm working on useful language models without taking the LLM approach:
Next-token prediction:
https://github.com/bennyschmidt/next-token-prediction https://github.com/bennyschmidt/next-token-prediction
Good for auto-complete, spellcheck, etc.
AI chatbot:
https://github.com/bennyschmidt/llimo https://github.com/bennyschmidt/llimo
Good for domain-specific conversational chat with instant responses that doesn't hallucinate.
- p1esk 2y agoWhy do you call your language model “transformer”?
- bschmidt1 2y agoLanguage is the language model that extends Transformer. Transformer is a base model for any kind of token (words, pixels, etc.). However, currently there is some language-specific stuff in Transformer that should be moved to Language :) I'm focusing first on language models, and getting into image generation next.
- p1esk 2y agoNo, I mean, a transformer is a very specific model architecture, and your simple language model has nothing to do with that architecture. Unless I’m missing something.
- richrichie 2y agoFor a century, transformer meant a very different thing. Power systems people are justifiably amused.
- p1esk 2y agoAnd it means something else in Hollywood. But we are discussing language models here, aren’t we?
- bschmidt1 2y agoAnd it fits the definition doesn't it since it tokenizes inputs to compute them against pre-trained ones, rather than being based on rules/lookups or arbitrary logic/algorithms? Even in CSS a matrix "transform" is the same concept - the word "transform" is not unique to language models, more a reference to how 1 set of data becomes another by way of computation. Same with tile engines / game dev. Say I wanted to rotate a map, this could be a simple 2D tic-tac-toe board or a 3D MMO tile map, anything in between: Input [ [0, 0, 1], [0, 0, 0], [0, 0, 0] ] Output [ [0, 0, 0], [0, 0, 0], [0, 0, 1] ] The method that takes the input and gives that output is called a "transformer" because it is not looking up some rule that says where to put the new values, it's performing math on the data structure whose result determines the new values. It's not unique to language models. If anything vector word embeddings are much later to this concept than math and game dev. An example of use of word "Transformer" outside language models in JavaScript is Three.js' https://threejs.org/docs/#examples/en/controls/TransformControls https://threejs.org/docs/#examples/en/controls/TransformCont... I used Three.js to build https://www.playshadowvane.com/ https://www.playshadowvane.com/ - built the engine from scratch and recall working with vectors (e.g. THREE Vector3 for XYZ stuff) years before they were being popularized by LLMs.
- bschmidt1 2y agoI still call it a transformer because the inputs are tokenized and computed to produce completions, not from lookups or assembling based on rules. > Unless I'm missing something. Only that I said "without taking the LLM approach" meaning tokens aren't scored in high-dimensional vectors, just as far simpler JSON bigrams. I don't think that disqualifies using the term "transformer" - I didn't want to call it a "computer" or a "completer". Have a better word? > JSON instead of vectors I did experiment with a low-dimensional vector approach from scratch, you can paste this into your browser console: https://gist.github.com/bennyschmidt/ba79ba64faa5ba18334b4ae06c857641 https://gist.github.com/bennyschmidt/ba79ba64faa5ba18334b4ae... But the n-gram approach is better, I don't think vectors start to pull away on accuracy until they are capturing a lot more contextual information (where there is already a lot of context inferred from the structure of an n-gram).
- kgeist 2y agoCalling it a "transformer" is misleading when discussing language modelling because it now means a very specific ML architecture while your project seems to be about Markov chains + hardcoded rules using regexps https://github.com/bennyschmidt/llimo/blob/master/models/Chat/index.js https://github.com/bennyschmidt/llimo/blob/master/models/Cha... The idea of tokenizing words and producing completions is not unique to the original transformers, it's a basic idea from NLP. So I'm not sure why you think it should be called a transformer just because it uses tokenized inputs and produces completions as well. It's like saying your new programming language has a "Java-based architecture" simply because they both have classes (and nothing else in common otherwise). >I didn't want to call it a "computer" or a "completer". Have a better word? I've seen projects which also use Markov chains + additional rules ontop, for example there's quite a few projects called "Markov chains with POS tagging": https://github.com/26medias/context-aware-markov-chains https://github.com/26medias/context-aware-markov-chains >not from lookups or assembling based on rules. Not quite sure about "it's not based on rules" when your code has things like: const MATCH_FIRST_MODAL = new RegExp(/IS|AM|ARE|WAS|HAS|HAVE|HAD|MUST|MAY|MIGHT|WERE|WILL|SHALL|CAN|COULD|WOULD|SHOULD|OUGHT|DOES|DID/); or const properNoun = `${part.value} `; if (isPrevNNP) { result += prependArticle(query, properNoun); } Pretty sure your examples in the video are also cherry-picked. The very first example is you asking "where is Paris?" What really happens is, one of the hardcoded regexps transforms it to "Paris is" and then the bigram model repeats the second sentence in the Paris dataset verbatim.
- vunderba 2y agoI took a very cursory look at the code, and it looks like this is just a standard Markov chain. Is it doing something different?
- bschmidt1 2y agoI get this question only on Hacker News, and am baffled as to why (and also the question "isn't this just n-grams, nothing more?"). https://github.com/bennyschmidt/next-token-prediction https://github.com/bennyschmidt/next-token-prediction ^ If you look at this GitHub repo, should be obvious it's a token prediction library - the video of the browser demo shown there clearly shows it being used with an <input /> to autocomplete text based on your domain-specific data. Is THAT a Markov chain, nothing more? What a strange question, the answer is an obvious "No" - it's a front-end library for predicting text and pixels (AKA tokens). https://github.com/bennyschmidt/llimo https://github.com/bennyschmidt/llimo This project, which uses the aforementioned library is a chat bot. There's an added NLP layer that uses parts-of-speech analysis to transform your inputs into a cursor that is completed (AKA "answered"). See the video where I am chatting with the bot about Paris? Is that nothing more than a standard Markov chain? Nothing else going on? Again the answer is an obvious "No" it's a chat bot - what about the NLP work, or the chat interface, etc. makes you ask if it's nothing more than a standard [insert vague philosophical idea]? To me, your question is like when people were asking if jQuery "is just a monad"? I don't understand the significance of the question - jQuery is a library for web development. Maybe there are some similarities to this philosophical concept "monad"? See: https://stackoverflow.com/questions/10496932/is-jquery-a-monad https://stackoverflow.com/questions/10496932/is-jquery-a-mon... It's like saying "I looked at your website and have concluded it is nothing more than an Array."
- kgeist 2y ago>Simpler take on embeddings (just bigrams stored in JSON format) So Markov chains
- bschmidt1 2y agoSee https://news.ycombinator.com/item?id=41419329 https://news.ycombinator.com/item?id=41419329