3 ms·
I've actually been working on a Unicode only code prediction LM, and it already works pretty well the main issues with Unicode models as also found in the artic
by f_devd 4y ago
I've actually been working on a Unicode only code prediction LM, and it already works pretty well the main issues with Unicode models as also found in the article is the large sequence length required compared a sentencePiece or BPE tokenizer. The current direction to resolve this seems to be to use structured state spaces (S4/DSS) models which scale linearly along sequence length compared to O(n^2) for transformers.
I haven't read the charformer paper yet but if it does a dynamic pooling of the tokens it could be a promising step to have this same functionality in transformer models.