3 ms·
I still call it a transformer because the inputs are tokenized and computed to produce completions, not from lookups or assembling based on rules. > Unless I'm
by bschmidt1 2y ago
I still call it a transformer because the inputs are tokenized and computed to produce completions, not from lookups or assembling based on rules.
> Unless I'm missing something.
Only that I said "without taking the LLM approach" meaning tokens aren't scored in high-dimensional vectors, just as far simpler JSON bigrams. I don't think that disqualifies using the term "transformer" - I didn't want to call it a "computer" or a "completer". Have a better word?
> JSON instead of vectors
I did experiment with a low-dimensional vector approach from scratch, you can paste this into your browser console: https://gist.github.com/bennyschmidt/ba79ba64faa5ba18334b4ae06c857641 https://gist.github.com/bennyschmidt/ba79ba64faa5ba18334b4ae...
But the n-gram approach is better, I don't think vectors start to pull away on accuracy until they are capturing a lot more contextual information (where there is already a lot of context inferred from the structure of an n-gram).
- kgeist 2y agoCalling it a "transformer" is misleading when discussing language modelling because it now means a very specific ML architecture while your project seems to be about Markov chains + hardcoded rules using regexps https://github.com/bennyschmidt/llimo/blob/master/models/Chat/index.js https://github.com/bennyschmidt/llimo/blob/master/models/Cha... The idea of tokenizing words and producing completions is not unique to the original transformers, it's a basic idea from NLP. So I'm not sure why you think it should be called a transformer just because it uses tokenized inputs and produces completions as well. It's like saying your new programming language has a "Java-based architecture" simply because they both have classes (and nothing else in common otherwise). >I didn't want to call it a "computer" or a "completer". Have a better word? I've seen projects which also use Markov chains + additional rules ontop, for example there's quite a few projects called "Markov chains with POS tagging": https://github.com/26medias/context-aware-markov-chains https://github.com/26medias/context-aware-markov-chains >not from lookups or assembling based on rules. Not quite sure about "it's not based on rules" when your code has things like: const MATCH_FIRST_MODAL = new RegExp(/IS|AM|ARE|WAS|HAS|HAVE|HAD|MUST|MAY|MIGHT|WERE|WILL|SHALL|CAN|COULD|WOULD|SHOULD|OUGHT|DOES|DID/); or const properNoun = `${part.value} `; if (isPrevNNP) { result += prependArticle(query, properNoun); } Pretty sure your examples in the video are also cherry-picked. The very first example is you asking "where is Paris?" What really happens is, one of the hardcoded regexps transforms it to "Paris is" and then the bigram model repeats the second sentence in the Paris dataset verbatim.
- bschmidt1 2y agoIt's literally what it is. You for some reason think transformers are unique to language models - boy are you late to the game https://en.wikipedia.org/wiki/Transformation_matrix https://en.wikipedia.org/wiki/Transformation_matrix A CSS matrix "transform" is the same concept. Same with tile engines & game dev. Say I wanted to rotate a map: Input [ [0, 0, 1], [0, 0, 0], [0, 0, 0] ] Output [ [0, 0, 0], [0, 0, 0], [0, 0, 1] ] The function is a "transformer" because it is not looking up some rule that says where to put the new values, it's performing math on the data structure whose result determines the new values. > Not quite sure about "it's not based on rules" when your code has things like: > > const MATCH_FIRST_MODAL Totally irrelevant to the topic. This is the chat interface itself which mostly just parses questions into cursors to be completed. You would be a fool to think ChatGPT has no NLP or parts-of-speech analysis. text-ada-embedding itself uses POS. > Pretty sure your examples in the video are also cherry-picked Fantastic detective work, you caught me. But just to confirm - why not just use it yourself? npm i next-token-prediction Here is an example you can run very easily in Chrome, so you don't have to rely solely on your amazing bullshit detector: https://github.com/bennyschmidt/next-token-prediction/tree/master/examples/ui-autocomplete https://github.com/bennyschmidt/next-token-prediction/tree/m... Don't forget to log the completions to prove that they aren't broken down by token, and instead just doing key/val lookups or text searches as you said. > What really happens is, one of the hardcoded regexps transforms it to "Paris is" The only thing you got right - that questions are transformed into sentences using conventional NLP in order to complete them. This functionality is what makes it a chat bot that you can ask questions.
- kgeist 2y ago>A CSS matrix "transform" is the same concept It's still misleading to call it a transformer in the context of NLP. It doesn't matter what it means in other, non-NLP areas (linear algebra, CSS or gamedev). It's like creating a procedural language and calling it "functional" because it has functions. Sure the concept of functions existed long before compsci but it would be very misleading because "functional programming" is a well-established term. >You would be a fool to think ChatGPT has no NLP or parts-of-speech analysis Pretty sure it doesn't. At least it's not required to. I've run lots of local models and it's just model weights without hardcoded regexps. In fact, I was able to feed grammar rules of an invented language into Claude Sonnet and it was able to construct proper sentences. >text-ada-embedding itself uses POS Do you have a link?