5 ms·
> The advantage holds while the DSL stays small and constrained enough that a few in-context examples can convey its usage. There is also a real upfront cost in
by codegladiator 3mo ago
> The advantage holds while the DSL stays small and constrained enough that a few in-context examples can convey its usage. There is also a real upfront cost in designing and maintaining the language and its semantic model. The payoff is therefore concentrated in well-factored, genuinely constrained DSLs backed by a validator.
dsl stays small is doing all the heavy lifting here
the premise is that because of these few existing dsls (like PlantUML mentioned) my "new dsl" will be equally effective. PlantUML has millions of examples in the training data, my new dsls are not (specially if its not json/yaml or just function chain based). as the number of things that can mix and match increase you are basically looking at a whole system prompt just describing the new language.
this brings us to the second part. step 2: after dsl is 'planned' (note they use the java compiler), the dsl need to have a real compiler/executor, not just a validator. because if then you are going to ask the llm to "compile the dsl to implementation" we are back to square 1.
- brookst 3mo agoI’ve had good luck with LLMs and ad hoc DSLs, as well as much less common DSLs like liquidsoap’s stream management DSL. I don’t think there’s magic in DSLs, I just think LLMs respond well to clear, simple structure. Compilation / execution is often true, but not necessary. DSLs can be entirely declarative and used just for gating the stages of a multi-step workflow with checkpoints that have more structure than natural language.
- codegladiator 3mo ago> less common DSLs like liquidsoap’s stream management DSL seems to be on github since 2008 so definitely in the training data. i am not talking about less or more common. either "your dsl" would need to look something like someone elses dsl (at this point is it your dsl?) or you need some way to get your dsls examples in the training data for the llm, or feed it in the prompt. > LLMs respond well to clear, simple structure and what a "clear simple structure" for a dsl is also quite not mentioned. clear and simple would be quite subjective based on the domain, the article says let the llm go in a loop trying to figure out the dsl for you. > checkpoints that have more structure than natural language if llm is at any point in the structured generation part then either you have a deterministic validator/compiler or you are back to reading/reviewing it manually, what can you trust ?
- brookst 3mo agoRe: less common, I was just saying it doesn’t take millions off examples like PlanetUML. > what can you trust I wasn’t clear enough here — you’re responding to DSLs as an interface from non-deterministic LLMs to deterministic external systems. What I meant was using DSLs as intermediate checkpoints in multi-LLM processing. If you just flow natural language through 5 LLM calls, the last one may be getting something very different from what it’s prompt is designed for. But if you make the DSL a contract for handoff, results are much more stable. Perfect and deterministic? No, of course not. Just an improvement and mitigation. But it’s served me well.
- codegladiator 3mo agoits served me okay. > Perfect and deterministic? No, of course not. Just an improvement and mitigation. exactly, not reliable, as the article tries to portray
- brookst 3mo agoInteresting … do you see reliable == infallible? I don’t. But certainly all software is imperfect, and LLMs doubly so.
- codegladiator 3mo agofor me reliable = deterministically repeatable. if a llm has been able to do a task successfully and i ask it to do the same task again, i want the reliability that i get the same outcome (be if failure or success) like if someone says "is this car reliable", i dont expect an infallible car. but if the cars third gear is 'you know sometimes it doesnt work', i wouldnt take that car out of city.
- samatman 2mo agoThis seems too strong to me. I agree with reliable as "repeatable", it's the determinism which seems excess to purpose. If I have a great meal every time I go to a restaurant, that chef is reliably good. I do not require that the tomatoes on the salad be of identical number or placement, or cut along the precise same bias. Nor would I count the sesame seeds on the bun. Are LLMs reliable by that standard? Sometimes yes, for some tasks. The envelope of tasks for which they're reliable is continuing to grow.
- billyp-rva 3mo ago> the premise is that because of these few existing dsls (like PlantUML mentioned) my "new dsl" will be equally effective. PlantUML has millions of examples in the training data, my new dsls are not (*specially if its not json/yaml* or just function chain based) I can confirm that having a DSL that is json/yaml helps a ton. Kind of like static type checking, it eliminates entire swaths of syntactical errors, allowing the LLM to focus on the semantics. > because if then you are going to ask the llm to "compile the dsl to implementation" we are back to square 1. I think this is an edge case; 99% of the time you (and/or the LLM) would have access to the implementor so it wouldn't need to do this.
- codegladiator 3mo ago> DSL that is json/yaml helps a ton it definitely does, and i would say json/yaml is not a dsl. this example of json/yaml keeps coming in the form of "DSL". i would say your configuration is not a dsl, it a declaration. llms are better at declarative stuff ? maybe but there are hardly that many of complex declarative frameworks. PlantUML is a real dsl. not just declarative yaml.
- billyp-rva 3mo ago> and i would say json/yaml is not a dsl But you can have DSLs that are json/yaml, is my point. > PlantUML is a real dsl. PlantUML is a DSL that isn't json/yaml. That doesn't make it better, and you can make the argument that it is worse because the tooling around it won't be as good.
- codegladiator 3mo ago> you can have DSLs that are json/yaml, is my point i disagree on the semantics of "DSL" vs "config in json". these are different. and anyways both of them just pretend to be "reliable" by throwing the responsibility to an upper layer of validator/compiler/interpreter > it is worse because the tooling around it won't be as good plantuml is good because of the the tooling around it. not sure if we are agreeing/disagreeing there, confused by the wording
- efromvt 3mo agoDSLs are a great middle ground for 'use LLM to turn ambiguous spec into something well defined', with the caveat that without discipline they'll inevitably expand until you should just have the agent write whatever the final language is. There's a context tax up front (which will hopefully be less relevant over time) and then you really need a compiler/linter with helpful errors to keep it on the rails, because there is no corrective context in pretraining for something novel. A purely descriptive DSL is just a convention, which is useful, but doesn't inject reliability the same way an enforced syntactic contract does.
- UncleEntity 3mo agoI just have them write the tools to write the DSL's to do the thing then (most of) the sloppy code stays in the generator and if all the different things depend on each other they don't go stale and whatnot. And let them design the DSL themselves for whatever task so it matches their 'internal concept' of how the things work. Worked out pretty well so far but not really practical unless your goal is to make the tools to make the DSLs to make jitting VMs -- https://github.com/dan-eicher/BBQ https://github.com/dan-eicher/BBQ kind of snowballed from "let's parse some binary files" to a way over the top toolkit for playing around with this stuff but, it's fun...
- brrrrrm 3mo ago> you are basically looking at a whole system prompt just describing the new language whats wrong with this? You may be over-indexing on the need for large quantities of examples. These days self-play through RL is far more effective and data (not compute) efficient.
- codegladiator 3mo agonothing wrong if your dsl stays small and you have deterministic validators/compilers to your actual target. if dsl gets large, the number of potential interactions your dsl allow will grow exponentially (unless you are building an one dimensional action layer). and there will be semantic issues unless your dsl is "clear and intuitive", also comprehensive enough to accomodate your ongoing changes, else every change is now 2 changes. so you now you need a comprehensive manual for your agent which needs to be sent in every /completion request.
- marssaxman 3mo agoThe situation is much better than you think. I actually have a little DSL project along these lines: it's meant to be a declarative component in a larger application which enforces certain data guarantees. It is not based on JSON or any other out-of-the-box format. The only source material for this DSL anywhere in the world is on my laptop: a little documentation, a few examples, a partial implementation. This is enough. Codex can not only give me arbitrary examples in this novel DSL on demand, it can see what I'm trying to do with the compiler and extend it for me.