3 ms·
Id love an ELI5 for PLE. Im trying to work it into my back of the napikin math for compute vs memory bandwidth limitations on tok/s in PP vs TG work. My attemp
by azath92 24d ago
Id love an ELI5 for PLE. Im trying to work it into my back of the napikin math for compute vs memory bandwidth limitations on tok/s in PP vs TG work.
My attempt at a simplification of this article on it https://sebastianraschka.com/llm-architecture-gallery/per-layer-embeddings/ https://sebastianraschka.com/llm-architecture-gallery/per-la... into a couple of sentences is that they are linear embeddings of the input token space projected per layer, which are then gated by the transformer outputs per layer.
This would mean that the only one set of weights for the ple path needs to be pumped across the memory bandwidth as they are the same linear weights for all layers?
Sheit, maybe im trying to simplify something that i need to look at in detail. but id love to leverage others understanding if possible
- tarruda 24d ago[dead]
- sixothree 24d agoWhile he avoids using the actual PLE acronym, he does actually describe the concept quite well. I think you may enjoy this video. Specifically around 5 minutes into the video is the part you're looking for. https://www.youtube.com/watch?v=1--PzaHafAU https://www.youtube.com/watch?v=1--PzaHafAU
- hadlock 22d ago>Id love an ELI5 for PLE. PLE is, instead of mixture of experts, mixture of associations