6 ms·
Language is the language model that extends Transformer. Transformer is a base model for any kind of token (words, pixels, etc.). However, currently there is s
by bschmidt1 2y ago
Language is the language model that extends Transformer. Transformer is a base model for any kind of token (words, pixels, etc.).
However, currently there is some language-specific stuff in Transformer that should be moved to Language :) I'm focusing first on language models, and getting into image generation next.
- p1esk 2y agoNo, I mean, a transformer is a very specific model architecture, and your simple language model has nothing to do with that architecture. Unless I’m missing something.
- richrichie 2y agoFor a century, transformer meant a very different thing. Power systems people are justifiably amused.
- p1esk 2y agoAnd it means something else in Hollywood. But we are discussing language models here, aren’t we?
- bschmidt1 2y agoAnd it fits the definition doesn't it since it tokenizes inputs to compute them against pre-trained ones, rather than being based on rules/lookups or arbitrary logic/algorithms? Even in CSS a matrix "transform" is the same concept - the word "transform" is not unique to language models, more a reference to how 1 set of data becomes another by way of computation. Same with tile engines / game dev. Say I wanted to rotate a map, this could be a simple 2D tic-tac-toe board or a 3D MMO tile map, anything in between: Input [ [0, 0, 1], [0, 0, 0], [0, 0, 0] ] Output [ [0, 0, 0], [0, 0, 0], [0, 0, 1] ] The method that takes the input and gives that output is called a "transformer" because it is not looking up some rule that says where to put the new values, it's performing math on the data structure whose result determines the new values. It's not unique to language models. If anything vector word embeddings are much later to this concept than math and game dev. An example of use of word "Transformer" outside language models in JavaScript is Three.js' https://threejs.org/docs/#examples/en/controls/TransformControls https://threejs.org/docs/#examples/en/controls/TransformCont... I used Three.js to build https://www.playshadowvane.com/ https://www.playshadowvane.com/ - built the engine from scratch and recall working with vectors (e.g. THREE Vector3 for XYZ stuff) years before they were being popularized by LLMs.
- p1esk 2y ago[flagged]
- bschmidt1 2y ago[flagged]
- p1esk 2y ago[flagged]
- bschmidt1 2y ago[flagged]
- dang 2y agoYou guys both broke the site guidelines badly in this thread. We have to ban accounts that post like this, so please don't. If you'd please review https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html and stick to the rules when posting here, we'd appreciate it.
- bschmidt1 2y agoI didn't know it was that strict, no offense to the other poster, it was just a little disagreement :)
- dang 2y agoYou guys both broke the site guidelines badly in this thread. We have to ban accounts that post like this, so please don't. If you'd please review https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html and stick to the rules when posting here, we'd appreciate it.
- bschmidt1 2y agoI still call it a transformer because the inputs are tokenized and computed to produce completions, not from lookups or assembling based on rules. > Unless I'm missing something. Only that I said "without taking the LLM approach" meaning tokens aren't scored in high-dimensional vectors, just as far simpler JSON bigrams. I don't think that disqualifies using the term "transformer" - I didn't want to call it a "computer" or a "completer". Have a better word? > JSON instead of vectors I did experiment with a low-dimensional vector approach from scratch, you can paste this into your browser console: https://gist.github.com/bennyschmidt/ba79ba64faa5ba18334b4ae06c857641 https://gist.github.com/bennyschmidt/ba79ba64faa5ba18334b4ae... But the n-gram approach is better, I don't think vectors start to pull away on accuracy until they are capturing a lot more contextual information (where there is already a lot of context inferred from the structure of an n-gram).
- kgeist 2y agoCalling it a "transformer" is misleading when discussing language modelling because it now means a very specific ML architecture while your project seems to be about Markov chains + hardcoded rules using regexps https://github.com/bennyschmidt/llimo/blob/master/models/Chat/index.js https://github.com/bennyschmidt/llimo/blob/master/models/Cha... The idea of tokenizing words and producing completions is not unique to the original transformers, it's a basic idea from NLP. So I'm not sure why you think it should be called a transformer just because it uses tokenized inputs and produces completions as well. It's like saying your new programming language has a "Java-based architecture" simply because they both have classes (and nothing else in common otherwise). >I didn't want to call it a "computer" or a "completer". Have a better word? I've seen projects which also use Markov chains + additional rules ontop, for example there's quite a few projects called "Markov chains with POS tagging": https://github.com/26medias/context-aware-markov-chains https://github.com/26medias/context-aware-markov-chains >not from lookups or assembling based on rules. Not quite sure about "it's not based on rules" when your code has things like: const MATCH_FIRST_MODAL = new RegExp(/IS|AM|ARE|WAS|HAS|HAVE|HAD|MUST|MAY|MIGHT|WERE|WILL|SHALL|CAN|COULD|WOULD|SHOULD|OUGHT|DOES|DID/); or const properNoun = `${part.value} `; if (isPrevNNP) { result += prependArticle(query, properNoun); } Pretty sure your examples in the video are also cherry-picked. The very first example is you asking "where is Paris?" What really happens is, one of the hardcoded regexps transforms it to "Paris is" and then the bigram model repeats the second sentence in the Paris dataset verbatim.