3 ms·
Ask HN: How training of LLM dedicated to code is different from LLM of “text”
Does it go beyond exposing to the LLM only "code"?, or there are extra steps in the training? (like giving the compiler/interpreter rules)?
Since programming are more structured , I think that using grammar that are dedicated to those language might be useful.
- sp332 3y agoThe tokenizer might need tweaking. Base Llama models, for example, are trained on text that has had consecutive spaces reduced to a single space. This is unhelpful for coding where specific amounts of whitespace is at least very nice to have and can even be meaningful. When you talk about a grammar, that's in the decoder, right? You don't need to retrain a model to use one of those.
- transformi 3y agoRegarding tokenizer- You are right in the context of python (but other language are more like pure "English" No?) I'm talking about the "coding grammar" (like for python/CPP..), in order that when he write the code- it should gain benefit from using some syntax checking
- sp332 3y agoYeah, the LLM outputs a distribution of likely next tokens. It is up to the decoder to select one, and it can use a grammar to enforce certain rules on the output. https://github.com/Hellisotherpeople/Constrained-Text-Generation-Studio https://github.com/Hellisotherpeople/Constrained-Text-Genera... or https://github.com/ggerganov/llama.cpp/blob/master/grammars/README.md https://github.com/ggerganov/llama.cpp/blob/master/grammars/... for example.