3 ms·
The tokenizer differences are major as LLMs are sensitive to whitespace handling. If I am reading the github page properly, OpenLLama failed to learn how to mod
by Vetch 3y ago
The tokenizer differences are major as LLMs are sensitive to whitespace handling. If I am reading the github page properly, OpenLLama failed to learn how to model code properly? Code contains many implicit reasoning tasks.
What other differences are there? The page doesn't mention how numbers are handled. These are two major things that impact model reasoning and numeric ability.
- emadm 3y agoCode is main thing, it has some tradeoffs. It tunes well on code though and the code ai team at stability ai are working on stuff. We can now set and forget runs so will have a better dataset for the next 13b and different tokeniser, this was meant to match as close as possible to be a drop in.