4 ms·
> During testing, we have noticed that this capability is unreliable as words are have missing or extra characters. We suspect this may have to do with the T5 t
by gzer0 3y ago
> During testing, we have noticed that this capability is unreliable as words are have missing or extra characters. We suspect this may have to do with the T5 text encoder we used: when the model encounters text in a prompt, it actually sees tokens that represent whole words and must map those to letters in an image.
Ah, interesting remark regarding text rendering.
- ronsor 3y agoThis isn't a new problem. A Google paper from late 2022 mentioned this fact; it went away when they used a byte/character-level text encoder (vs. tokenizing text first).
- mycall 3y agoIt still is an amazing state machine.
- ollin 3y ago"Character-Aware Models Improve Visual Text Rendering" https://arxiv.org/abs/2212.10562 https://arxiv.org/abs/2212.10562 for anyone curious
- tobr 3y agoThis is an aside, but in general, aren’t tokens a pretty bad hack to improve context length? I don’t know enough about the theory to say, but it seems like it should make all kinds of things much more fragile, like how well it can understand misspellings, or languages that don’t map well to the tokens available.