3 ms·
The transformer figures out that wants and cash are both verbs (both words can also be nouns). We’ve represented this added context as red text in parentheses,
by version_five 3y ago
The transformer figures out that wants and cash are both verbs (both words can also be nouns). We’ve represented this added context as red text in parentheses, but in reality, the model would store it by modifying the word vectors in ways that are difficult for humans to interpret. These new vectors, known as a hidden state, are passed to the next transformer in the stack.
This isn't really true, it's like when we say that initial layers in CNNs detect big features, it's an oversimplification that gives the wrong impression. I understand the objective of trying to explain a technical subject to a nontechnical audience but I don't think it works at this deep a level.
If you want to learn how llms work, I'd suggest "GPT in 60 lines of numpy": https://jaykmody.com/blog/gpt-from-scratch/ https://jaykmody.com/blog/gpt-from-scratch/
It's more concise and doesn't resort to analogies.
- fritzthedev 3y agoThis is a great writeup. Thank you for sharing.
- binarybits 3y agoWhat isn't really true? The passage is about a hypothetical LLM so obviously the exact steps depicted here don't correspond to any particular LLM. But LLMs undoubtedly modify hidden states to reflect context gleaned from other words, right? I don't understand what point you're making.
- llm_nerd 3y agoIt is reasonably correct. Correct enough for such an intro piece. And FWIW one of the authors is a cognitive scientist at the University of California. Your alternative site serves an entirely different audience/purpose. I'm surprised to see this submission did relatively poorly. HN has seen so many trivial "what is tokenization" articles. This Ars article is actually substantive and gives a very real notion of how LLMs work via their components/successors. A++
- version_five 3y ago> I'm surprised to see this submission did relatively poorly. I think it was too long and too "deep" but at the same time geared towards a nontechnical audience and trying to explain by analogy, so it's going to lose people that don't already understand and annoy people who do. What I like about the link I posted is that it's approachable to anyone who understands python and some basic preliminaries. It provides understanding to an audience that's potentially capable of understanding. The Ars article tries to do to much and I expect will rarely succeed in giving someone who can't understand my article (I posted it I didn't write it) an understanding of transformers.