3 ms·
could even try it with a fraction of the attention heads, instead of introducing new tokens
by lucidrains 2y ago
could even try it with a fraction of the attention heads, instead of introducing new tokens
- sdenton4 2y agoAn important piece here is that there's still a training signal making it to the makes weights. See SimSiam for a similar example.
- lucidrains 2y agoindeed, simsiam is a great example of the effectiveness of using stop gradient