3 ms·
Thanks for the helpful explanation. What is the context length in this explanation?
by ntonozzi 3y ago
Thanks for the helpful explanation. What is the context length in this explanation?
- two_in_one 3y agomy guess it's N here
- cshimmin 3y agoCorrect. In the context of LLM's, the "items" I refer to would be tokens. As a particle physicist, in the transformers I work with the "items" often instead represent either particles, or detector measurements.
- two_in_one 3y agoThanks, I'm trying to understand all this mechanics. Matrix multiplication is an equivalent to passing through a single fully connected convolution layer. Which means we can probably beef up and make it a small network. Also it's easy to make it work not in fixed windows, but in scrolling, if the output is processed sequentially. Even if not it may make sense too. It can be implemented efficiently. The idea is that in switching window first and last elements are being influenced only from one side. Which is not a good thing in text processing. Only middle elements are influenced from both sides. With scrolling window all elements currently being processed are in the middle. Except for the ends of the dataset.