3 ms·
It's literally what it is. You for some reason think transformers are unique to language models - boy are you late to the game https://en.wikipedia.org/wiki/Tra
by bschmidt1 2y ago
It's literally what it is. You for some reason think transformers are unique to language models - boy are you late to the game https://en.wikipedia.org/wiki/Transformation_matrix https://en.wikipedia.org/wiki/Transformation_matrix
A CSS matrix "transform" is the same concept.
Same with tile engines & game dev. Say I wanted to rotate a map:
Input
[
[0, 0, 1],
[0, 0, 0],
[0, 0, 0]
]
Output
[
[0, 0, 0],
[0, 0, 0],
[0, 0, 1]
]
The function is a "transformer" because it is not looking up some rule that says where to put the new values, it's performing math on the data structure whose result determines the new values.
> Not quite sure about "it's not based on rules" when your code has things like:
>
> const MATCH_FIRST_MODAL
Totally irrelevant to the topic. This is the chat interface itself which mostly just parses questions into cursors to be completed. You would be a fool to think ChatGPT has no NLP or parts-of-speech analysis. text-ada-embedding itself uses POS.
> Pretty sure your examples in the video are also cherry-picked
Fantastic detective work, you caught me. But just to confirm - why not just use it yourself? npm i next-token-prediction
Here is an example you can run very easily in Chrome, so you don't have to rely solely on your amazing bullshit detector: https://github.com/bennyschmidt/next-token-prediction/tree/master/examples/ui-autocomplete https://github.com/bennyschmidt/next-token-prediction/tree/m...
Don't forget to log the completions to prove that they aren't broken down by token, and instead just doing key/val lookups or text searches as you said.
> What really happens is, one of the hardcoded regexps transforms it to "Paris is"
The only thing you got right - that questions are transformed into sentences using conventional NLP in order to complete them. This functionality is what makes it a chat bot that you can ask questions.
- kgeist 2y ago>A CSS matrix "transform" is the same concept It's still misleading to call it a transformer in the context of NLP. It doesn't matter what it means in other, non-NLP areas (linear algebra, CSS or gamedev). It's like creating a procedural language and calling it "functional" because it has functions. Sure the concept of functions existed long before compsci but it would be very misleading because "functional programming" is a well-established term. >You would be a fool to think ChatGPT has no NLP or parts-of-speech analysis Pretty sure it doesn't. At least it's not required to. I've run lots of local models and it's just model weights without hardcoded regexps. In fact, I was able to feed grammar rules of an invented language into Claude Sonnet and it was able to construct proper sentences. >text-ada-embedding itself uses POS Do you have a link?
- bschmidt1 2y agoAgain they are the exact same concept. Whether vectors represent tiles in a video game, an object in CSS, matrix algebra you took in school, or the semantics of words used by LLMs, in all cases it's the same meaning of the word "transform". It's not specific to language models at all - which was the thesis of your whole argument. > it's not required to. I've run lots of models Then you must know about skip-gram and how embeddings are trained: https://medium.com/@corymaklin/word2vec-skip-gram-904775613b4c https://medium.com/@corymaklin/word2vec-skip-gram-904775613b... What is meant by "sliding window" or "skip gram" is bigram mapping (or other n-gram). This is ML 101. It's the same training methodology and data structure used in my next-token-prediction lib, and is widely used for training for LLMs. Ask your local AI to explain the basics, or see examples like: https://www.kaggle.com/code/hamishdickson/training-and-plotting-word2vec-with-bigrams https://www.kaggle.com/code/hamishdickson/training-and-plott... > ChatGPT doesn't use parts-of-speech Yes it does, there's not only a huge business in tagging data (both POS and NER) adjacent to AI, but OpenAI specifically famously used African workers on very low wages to tag a bunch of data. ChatGPT uses text-embedding-ada, you'll have to put 2 and 2 together as they don't open source that part. Mistral says: "The preprocessing stage of Text-Embedding-ADA-002 involves applying POS tags to the input text using a separate POS tagger like Spacy or Stanford NLP. These POS tags can be useful for segmenting sentences into individual words or tokens." > I use Claude to make new languages Cool story, has nothing to do with the topic
- kgeist 2y ago>It's not specific to language models at all - which was the thesis of your whole argument. I didn't say that it's unique to LMs. My argument is that saying "my LM is a transformer" is misleading because "transformer" in the context of LMs means a very specific architecture. You're deliberately misusing terms, probably to draw attention to your project. >OpenAI specifically famously used African workers on very low wages to tag a bunch of data Did they tag Polish parts of speech too? Or Ancient Greek? ChatGPT constructs grammatically correct Ancient Greek. I thought they tagged "harmful/non-harmful", not parts of speech? >ChatGPT uses text-embedding-ada [Citation needed] NanoGPT, for example, learns embeddings together with the rest of the network so, as I said, manual tagging is not required. Anyway, looking forward to hearing news about your image generation project. Any news?