3 ms·
There's quite a bit of impressive jargon in that repo and the video! I had to briefly look at the paper abstract, which explains that this is about solving the
by uniqueuid 4y ago
There's quite a bit of impressive jargon in that repo and the video!
I had to briefly look at the paper abstract, which explains that this is about solving the sequence limit of transformer text models:
>While beneficial, the quadratic complexity of self-attention on the input sequence length has limited its application to longer sequences -- a topic being actively studied in the community. To address this limitation, we propose Nyströmformer -- a model that exhibits favorable scalability as a function of sequence length.
That's cool — I'm looking forward to being able to process texts > 512 tokens in the future, and would be especially excited if that were possible for sentence-bert.