3 ms·
I was initially impressed with the landing page, but it does look a bit suspect when things are claimed to be 100x faster without much info on the HW accelerati
by fzysingularity 3y ago
I was initially impressed with the landing page, but it does look a bit suspect when things are claimed to be 100x faster without much info on the HW acceleration or the model sizes.
My best guess is that they're using two approaches to get this running faster:
- structured generation techniques from sglang (https://github.com/sgl-project/sglang https://github.com/sgl-project/sglang) that allow them to generate faster JSON (with look-ahead / pre-fill) with strong guarantees on the output (i.e. 100% reliable, without requiring any retries).
- distilling a gpt-3.5 turbo-esque model from GPT-4 JSON outputs, and using it in conjuction with above to give the additional performance boosts on inference.
It doesn't seem like they're deploying on any custom silicon, nor have they optimized GPU kernels to suggest that the speed ups came there.
- zuck_vs_musk 3y ago> strong guarantees on the output (i.e. 100% reliable, without requiring any retries). Has anyone seen a good JSON library that can handle slightly broken JSON? e.g. trailing commas, unescaped newlines, etc.? I have not found a good one.
- dartos 3y agoYou can force LLMs to generate valid json by using a context free grammar FWIW
- zuck_vs_musk 3y agoPlease elaborate
- kaiokendev 3y agoMatt Rickard has a good entry level blog post about it, from the angle of regex constraining [0]. Context free grammars follow the same principle, except using a finite state machine to restrict the action space. [0]: https://matt-rickard.com/rellm https://matt-rickard.com/rellm
- dartos 3y agoLook at llama.cpp grammars, lmql, or guidance-ai
- reissbaker 3y agoEntirely broken JSON, no — I would be surprised if one existed. If you want slightly laxer semantics like trailing commas, JSON5 [1] is a pretty good spec and is JSON-compatible. I used to use it for LLMs (while telling them to emit JSON — no need to confuse them by explaining JSON5), in order to handle things like trailing commas, but in my experience LLMs have gotten good enough over the last year I mostly don't even bother anymore. 1: https://json5.org/ https://json5.org/
- catlifeonmars 3y agoShould be easy to build on top of a lexer as a pre-parsing pass.
- tlarkworthy 3y agoDirtyjson
- azeirah 3y agoI've been doing a lot of indie work with structured generation and llama.cpp, you can get extremely fast responses with caching and deterministic token skipping. When generating json, calling the llm when you already know that after { "brand": "Toyota" Comes , "year": Is a massive waste. If the data itself is constrained too you can skip most of that too! You'll go down from needing 20 calls to the llm to just three for a simple piece of data like { "brand": "Toyota", "year": 1995 } If they combine these techniques with a model that's specifically trained for structured output, along with a novel inference-time pruning technique that they were talking about in the post I can definitely see them getting these kinds of inference speeds. I'm experimenting with a self hosted api that is fast enough to not even need a gpu for single user use cases (because the latency is good, but not batching). Once I'm done with the finishing touches I'll rent a GPU server for actual hosting.
- Incipient 3y agoHave you seen any blog posts that have some details on this? I can imagine, roughly, the concept but it sounds interesting and I'd like to get a better understanding.
- azeirah 3y agoYeah, dottxt has a post on this, it's a technique starting with a c I believe, something like continuum or something? The op also has something about compressing finite state machines, it's basically the same thing with slightly different details in the implementation.
- fzysingularity 3y agoCheck out outlines by .txt : https://github.com/outlines-dev/outlines https://github.com/outlines-dev/outlines sglang: https://arxiv.org/abs/2312.07104 https://arxiv.org/abs/2312.07104
- Incipient 3y agoThanks!
- alexandd 3y ago[dead]
- dwallin 3y agoI thought they pretty clearly explained where the additional performance came from. If there is only 1 valid schema confirming option you can skip the LLM entirely. If there are only a limited number of possible tokens (eg. only } or ,) then you run it on smaller subset of the model. Between these two you capture a large amount of the actual token count of most json.