4 ms·
Wouldn't a llm that just tokenized by character be good at it?
by IncreasePosts 1y ago
Wouldn't a llm that just tokenized by character be good at it?
- curioussquirrel 1y agoYes, but it would hurt its contextual understanding and effectively reduce the context window several times.
- viraptor 1y agoOnly in the current most popular architectures. Mamba and RWKV style LLMs may suffer a bit but don't get a reduced context in the same sense.
- curioussquirrel 1y agoYou're right. There was also an experiment in Meta which tokenized bytes directly and it didn't hurt performance much in very small models.
- typpilol 1y agoI asked this in another thread and it would only be better with unlimited compute and memory. Because without those, then the llm has to encode way more parameters and way smaller context windows. In a theoretical world, it would be better, but might not be much better.