3 ms·
For the local model it forces valid json structure and formatting tokens are being produced by code rather than generated by an LLM.
by startupsfail 3y ago
For the local model it forces valid json structure and formatting tokens are being produced by code rather than generated by an LLM.
- verdverm 3y agosounds like post-processing made out to be something more? everyone is doing this, it's just part of the pipeline, certainly nothing innovative on that front in guidance
- mmoskal 3y agoIt updates token logits (probabilities) after every token before sampling. I don't think this is very common yet.
- newhouseb 3y agoRight, there are many folks (dozens of us!) yelling about logit processors and building them into various frameworks. The mostly widely accessible form of this is probably BNF grammar biasing in llama.cpp: https://github.com/ggerganov/llama.cpp/blob/master/grammars/README.md https://github.com/ggerganov/llama.cpp/blob/master/grammars/...
- verdverm 3y agoanecdotal counter evidence, I've seen multiple projects / papers manipulating the logits, it's a very common thing to think of doing now to improve performance (by eliminating bad options from consideration)
- Der_Einzige 3y agoStill rare, but I wrote a whole paper last year about what happens when you use this functionality (a lot, including defeating any kind of RLHF!) https://aclanthology.org/2022.cai-1.2.pdf https://aclanthology.org/2022.cai-1.2.pdf