3 ms·
This is a great writeup! There was a period where reliable structured output was a significant differentiator and was the 'secret sauce' behind some companies s
by iandanforth 1y ago
This is a great writeup! There was a period where reliable structured output was a significant differentiator and was the 'secret sauce' behind some companies success. A NL->SQL company I am familiar with comes to mind. Nice to see this both public and supported by a growing ecosystem of libraries.
One statement surprised me was that the author thinks "models over time will just be able to output JSON perfectly without the need for constraining over time."
I'm not sure how this conclusion was reached. "Perfectly" is a bar that probabilistic sampling cannot meet.
- ninadpathak 1y ago[dead]
- joatmon-snoo 1y agoWe’ve had a lot of success implementing schema-aligned parsing in BAML, a DSL that we’ve built to simplify this problem. We actually don’t like constrained generation as approach - among other issues it limits your ability to use reasoning - and instead the technique we’re using is algorithm-driven error-tolerant output parsing. https://boundaryml.com/ https://boundaryml.com/
- maxdo 1y agoLove your work , thanks ! , 12 factor agent implementation uses your tools too.
- parthsareen 1y agoThank you! Maybe not "perfect" but near-perfect is something we can expect. Models like the Osmosis structure which just structure data inspired some of that thinking (https://ollama.com/Osmosis/Osmosis-Structure-0.6B https://ollama.com/Osmosis/Osmosis-Structure-0.6B). Historically, JSON generation has been a latent capability of a model rather than a trained one, but that seems to be changing. gpt-oss was particularly trained for this type of behavior and so the token probabilities are heavily skewed to conform to JSON. Will be interesting to see the next batch of models!