4 ms·
I already use streaming partial json responses (progressive json) with AI tool calls in production. It’s become a thing, even beyond RSCs, and has many practic
by vinnymac 1y ago
I already use streaming partial json responses (progressive json) with AI tool calls in production.
It’s become a thing, even beyond RSCs, and has many practical uses if you stare at the client and server long enough.
- tough 1y agohow do you do that exactly?
- richin13 1y agoNot the original commenter but I’ve done this too with Pydantic AI (actually the library does it for you). See “Streaming Structured Output” here https://ai.pydantic.dev/output/#streaming-structured-output https://ai.pydantic.dev/output/#streaming-structured-output
- tough 1y agoThanks yes! Im aware of structured outputs, llama.cpp has also great support with GBNF and several languages beyond json. I've been trying to create go/rust ones but its way harder than just json due to all the context/state they carry over
- danenania 1y agoOne way is to eagerly call JSON.parse as fragments are coming in. If you also split on json semantic boundaries like quotes/closing braces/closing brackets, you can detect valid objects and start processing them while the stream continues.
- tough 1y agoInteresting approach! thanks for sharing
- deleted 1y ago[deleted]
- motorest 1y agoCan you offer some detail into why you find this approach useful? From an outsider's perspective, if you're sending around JSON documents so big that it takes so long to parse them to the point reordering the content has any measurable impact on performance, this sounds an awful lot like you are batching too much data when you should be progressively fetching child resources in separate requests, or even implementing some sort of pagination.
- Wazako 1y agoSlow llm generation. A progressive display of a progressive json is mandatory.