3 ms·
This quote seems to fundamentally misunderstand what transformers are doing at all. Technically I suppose you could save all gradient updates from every input t
by ProlificInquiry 4y ago
This quote seems to fundamentally misunderstand what transformers are doing at all. Technically I suppose you could save all gradient updates from every input token, and do some weighted averaging to show what inputs affected the particular output the most, but saving all those gradient updates would be unimaginably space consuming. "Feasible" is doing a lot of work there.
It's very hard for people to get away from the idea that GPT is "copying" something, but that's not what it's doing. The reality is, to get the exact artifact which produced the code in question, you need "Call me Ishmael" from Moby Dick just as much as the Linux kernel source.
- amelius 4y ago> The reality is, to get the exact artifact which produced the code in question, you need "Call me Ishmael" from Moby Dick just as much as the Linux kernel source. Not always. Sometimes it just copies code without modification. It never tells you when it does that, though. So to be on the safe side, better assume that it always does.
- janalsncm 4y agoEven then, the presence of a next token is just as informative as the absence of another during training. That information also gets backpropagated. And during inference, good luck identifying which weights were responsible for a given next token (assuming we’re using greedy decoding, don’t even get started on beam search) let alone which samples contributed to that weight (hint: they all did).