3 ms·
Correct. In the context of LLM's, the "items" I refer to would be tokens. As a particle physicist, in the transformers I work with the "items" often instead rep
by cshimmin 3y ago
Correct. In the context of LLM's, the "items" I refer to would be tokens. As a particle physicist, in the transformers I work with the "items" often instead represent either particles, or detector measurements.
- two_in_one 3y agoThanks, I'm trying to understand all this mechanics. Matrix multiplication is an equivalent to passing through a single fully connected convolution layer. Which means we can probably beef up and make it a small network. Also it's easy to make it work not in fixed windows, but in scrolling, if the output is processed sequentially. Even if not it may make sense too. It can be implemented efficiently. The idea is that in switching window first and last elements are being influenced only from one side. Which is not a good thing in text processing. Only middle elements are influenced from both sides. With scrolling window all elements currently being processed are in the middle. Except for the ends of the dataset.