4 ms·
Wouldn't this introduce hidden bias towards word-like content, i.e. making the output seem more coherent than its actually is, masking LLM output quality?
by countWSS 3y ago
Wouldn't this introduce hidden bias towards word-like content,
i.e. making the output seem more coherent than
its actually is, masking LLM output quality?
- thatguysaguy 3y agoIt's not really hidden, that's an intentional feature. All our LMs have been using this (or similar) since GPT-1.
- Der_Einzige 3y agoIt also destroys models ability to follow phonetic or syntactic constraints: https://paperswithcode.com/paper/most-language-models-can-be-poets-too-an-ai-1 https://paperswithcode.com/paper/most-language-models-can-be... I've been waiting for real innovation to come in tokenization algorithms. I think a truly well trained character level model would be awesome. Far easier to guide, and might even know how to "count" it's own output.