4 ms·
We’ve been keeping a close eye on this as well as research is coming out. We’re looking into improving sampling as a whole on both speed and accuracy. Hopefull
by parthsareen 2y ago
We’ve been keeping a close eye on this as well as research is coming out. We’re looking into improving sampling as a whole on both speed and accuracy.
Hopefully with those changes we might also enable general structure generation not only limited to JSON.
- hackernewds 2y agoWho is "we"?
- parthsareen 2y agoI authored the blog with some other contributors and worked on the feature (PR: https://github.com/ollama/ollama/pull/7900 https://github.com/ollama/ollama/pull/7900). The current implementation uses llama.cpp GBNF grammars. The more recent research (Outlines, XGrammar) points to potentially speeding up the sampling process through FSTs and GPU parallelism.
- netghost 2y agoThank you for the details!
- mmoskal 2y agoIf you want avoid startup cost, llguidance [0] has no compilation phase and by far the fullest JSON support [1] of any library. I did a PoC llama.cpp integration [2] though our focus is mostly server-side [3]. [0] https://github.com/guidance-ai/llguidance https://github.com/guidance-ai/llguidance [1] https://github.com/guidance-ai/llguidance/blob/main/parser/src/json/README.md https://github.com/guidance-ai/llguidance/blob/main/parser/s... [2] https://github.com/ggerganov/llama.cpp/pull/10224 https://github.com/ggerganov/llama.cpp/pull/10224 [3] https://github.com/guidance-ai/llgtrt https://github.com/guidance-ai/llgtrt
- parthsareen 2y agoThis looks really useful. Thank you!
- HanClinto 2y agoI have been thinking about your PR regularly, and pondering about how we should go about getting this merged in. I really want to see support for additional grammar engines merged into llama.cpp, and I'm a big fan of the work you did on this.