4 ms·
Unfortunately, it was trained with a tokenizer which squashes repeated spaces. Basically, this means it can't do code generation or any whitespace-based text fo
by easygenes 3y ago
Unfortunately, it was trained with a tokenizer which squashes repeated spaces. Basically, this means it can't do code generation or any whitespace-based text formatting.
They had one job, and this wasn't it.
- Tostino 3y agoThat is really disappointing to hear... a whole lot of wasted compute.
- ehsanu1 3y agoSource? I do see this note in the link: > Please note that it is advised to avoid using the Hugging Face fast tokenizer for now, as we’ve observed that the auto-converted fast tokenizer sometimes gives incorrect tokenizations. If you are correct, maybe it'll work for languages not reliant on whitespace, paired with a code formatter, but we're definitely coming to expect more from these models these days.
- loudmax 3y agoFrom their project page https://github.com/openlm-research/open_llama https://github.com/openlm-research/open_llama : > For current version of OpenLLaMA models, our tokenizer is trained to merge multiple empty spaces into one before tokenization, similar to T5 tokenizer. Because of this, our tokenizer will not work with code generation tasks (e.g. HumanEval) since code involves many empty spaces. We are planning to open source long context models trained on more code data. Stay tuned. Sounds like they're working on it.
- dmarchand90 3y agoI don't understand the negativity. They are giving a complex piece of technology away for free. They choose to simplify the problem. So what?
- anon373839 3y agoIt was just a very odd decision to make. The project was billed as creating a “drop-in replacement” for the original LLaMA weights, so it’s unclear why they would want to train the model with a different tokenizer.
- denverllc 3y agoFor 5 more tokens (i.e. two space, four space, eight space, etc.), they could have converted up to 32 spaces into log2(N) tokens. e.g. 31 spaces would be 16 + 8 + 4 + 2 + 1 = 5 tokens; 15 spaces would be 8 + 4 + 2 + 1, etc; but for a lot of code the spaces are an even multiple (such as when a tab = 4 spaces or tab = 2 spaces) so even fewer would be required. If you took all of the Python code and queried how many spaces are used, I would wager most lines have four spaces, a few have eight, and even fewer have more than that. If the Python code were written with four spaces = tab, then it wouldn't have cost any more tokens than squashing the repeated spaces (since a "four-space" token consumes the same as a "one-space" token).