3 ms·
You're exactly right. The llguidance library [1,2] seems to have emerged as the go-to solution for this by virtue of being >10X faster than its competition. It'
by osaariki 10mo ago
You're exactly right. The llguidance library [1,2] seems to have emerged as the go-to solution for this by virtue of being >10X faster than its competition. It's work from some past colleagues of mine at Microsoft Research based on theory of (regex) derivatives, which we perviously used to ship a novel kind of regex engine for .NET. It's cool work and AFAIK should ensure full adherence to a JSON grammar.
llguidance is used in vLLM, SGLang, internally at OpenAI and elsewhere. At the same time, I also see a non-trivial JSON error rate from Gemini models in large scale synthetic generations, so perhaps Google hasn't seen the "llight" yet and are using something less principled.
1: https://guidance-ai.github.io/llguidance/llg-go-brrr https://guidance-ai.github.io/llguidance/llg-go-brrr
2: https://github.com/guidance-ai/llguidance https://github.com/guidance-ai/llguidance
- red2awn 10mo agoCool stuff! I don't get how all the open source inference framework have this down but the big labs doesn't... Gemini [0] is falsely advertising this: > This capability guarantees predictable and parsable results, ensures format and type-safety, enables the programmatic detection of refusals, and simplifies prompting. [0]: https://ai.google.dev/gemini-api/docs/structured-output?example=recipe#:~:text=This%20capability%20guarantees%20predictable%20and%20parsable%20results%2C%20ensures%20format%20and%20type%2Dsafety%2C%20enables%20the%20programmatic%20detection%20of%20refusals%2C%20and%20simplifies%20prompting. https://ai.google.dev/gemini-api/docs/structured-output?exam...