4 ms·
The inference engine (llama.CPP) has full control over the possible tokens during inference. It can "force" the llm to output only valid tokens so that it produ
by tarruda 1y ago
The inference engine (llama.CPP) has full control over the possible tokens during inference. It can "force" the llm to output only valid tokens so that it produces valid json
- kristjansson 1y agoand in fact leverages that control to constrain outputs to those matching user-specified BNFs https://github.com/ggml-org/llama.cpp/tree/master/grammars https://github.com/ggml-org/llama.cpp/tree/master/grammars
- wubrr 1y agoVery cool!
- wubrr 1y agoAhh, I stand corrected, very cool!