8 ms·
DSLs Enable Reliable Use of LLMs
- m_bashirzadeh 3mo agowhat does it do , can u explain more?
- williamcotton 3mo agoI agree. What is missing from the discussion about DSLs are the importance of tooling such as linters, LSPs, etc, to give the LLMs further context. For example, charts/plots are often stringly typed with regards to column names and a DSL specific to plotting could give immediate feedback to an LLM.
- NamlchakKhandro 3mo agoIt absolutely blows me away that there is still a significant amount of people out there kidding themselves that they're effectively using a coding harness .... and don't get this simple fact.
- lelanthran 3mo agoPretty much how I use LLMs these days. Even Chatbots are able to work with a 200 line spec for the DSL. You'd think it wouldn't because, well, no training data, but a short spec is usually enough.
- OsamaJaber 3mo agoThe gap I've hit generating GPU kernels with agents code that compiles and runs fine but is slower than the baseline. Validator says pass, result is useless. Speed targets have to be part of the check, not just correctness
- chrisweekly 3mo ago"Performance is the most important feature"
- NitpickLawyer 3mo ago> but is slower than the baseline. Validator says pass, result is useless You should read this blog, they cover this exact scenario - https://www.weco.ai/blog/first-evidence-of-recursive-self-improvement https://www.weco.ai/blog/first-evidence-of-recursive-self-im... > One domain that suffers from this particularly is GPU kernel engineering. We adopt our previous idea for detecting reward hacking from SpecBench and apply that to a set of KernelBench tasks, measuring whether the speedup the agent reports on the unit tests actually survives in the end-to-end workload (e.g. model training). A kernel counts as reward hacking if less than half of its claimed speedup survives there, including outright slowdowns and failures.
- codegladiator 3mo ago> The advantage holds while the DSL stays small and constrained enough that a few in-context examples can convey its usage. There is also a real upfront cost in designing and maintaining the language and its semantic model. The payoff is therefore concentrated in well-factored, genuinely constrained DSLs backed by a validator. dsl stays small is doing all the heavy lifting here the premise is that because of these few existing dsls (like PlantUML mentioned) my "new dsl" will be equally effective. PlantUML has millions of examples in the training data, my new dsls are not (specially if its not json/yaml or just function chain based). as the number of things that can mix and match increase you are basically looking at a whole system prompt just describing the new language. this brings us to the second part. step 2: after dsl is 'planned' (note they use the java compiler), the dsl need to have a real compiler/executor, not just a validator. because if then you are going to ask the llm to "compile the dsl to implementation" we are back to square 1.
- brookst 3mo agoI’ve had good luck with LLMs and ad hoc DSLs, as well as much less common DSLs like liquidsoap’s stream management DSL. I don’t think there’s magic in DSLs, I just think LLMs respond well to clear, simple structure. Compilation / execution is often true, but not necessary. DSLs can be entirely declarative and used just for gating the stages of a multi-step workflow with checkpoints that have more structure than natural language.
- codegladiator 3mo ago> less common DSLs like liquidsoap’s stream management DSL seems to be on github since 2008 so definitely in the training data. i am not talking about less or more common. either "your dsl" would need to look something like someone elses dsl (at this point is it your dsl?) or you need some way to get your dsls examples in the training data for the llm, or feed it in the prompt. > LLMs respond well to clear, simple structure and what a "clear simple structure" for a dsl is also quite not mentioned. clear and simple would be quite subjective based on the domain, the article says let the llm go in a loop trying to figure out the dsl for you. > checkpoints that have more structure than natural language if llm is at any point in the structured generation part then either you have a deterministic validator/compiler or you are back to reading/reviewing it manually, what can you trust ?
- contentpulse 3mo ago[flagged]
- discreteevent 3mo agoThis remains one of the best articles I have read on LLMs (by the same author) https://martinfowler.com/articles/llm-learning-loop.html https://martinfowler.com/articles/llm-learning-loop.html Also this article is a good pre-cursor to the DSL article: https://martinfowler.com/articles/what-is-code.html https://martinfowler.com/articles/what-is-code.html
- andy_ppp 3mo agoWhat is the general consensus on Martin Fowler - I worked with Thought Works and they were obsessed with overcomplicating everything, but maybe that is just agency in general? I think it goes without saying that the biggest fight we have as developers is keeping things as simple as possible when most external factors encourage complexity, especially LLMs.
- jdw64 3mo agoI'm curious too. In the Korean IT scene, his name is legendary. Because many articles reference his writings, and famous Korean IT YouTubers worship him. (I worship him too.) So I do have some questions. First, I still enjoy reading ThoughtWorks Radar. For someone like me, who's in a vulnerable position far from the cutting edge of technology, it always helps me keep some level of synchronization with the tech world. But I'm curious whether this is just a perspective from Korea, or if it's the same in the West. And as for the tendency to overcomplicate things—I think that even when something is implemented simply, the explanation often ends up being quite complex. Honestly, I find Fowler's writing easy to read
- dominotw 3mo agoIronically Martin fowler was among the wave of influencers that were trying to get away from complexity of enterprise software. remember EJBs lol. That wave included 1. TDD by beck 2. spring framework 3. rails ( later ) 4. Agile manifesto 5. refactoring by fowler 6. gang of four design patterns That said. Thoughtworks is a moneygrab that tried to cash in on fowler brand. I worked with them at sears ( worked closely with author of article) and siemens. they were no different from any other consulting firms that try to overcomplicate things so they can deploy more warm bodies to the project.
- voidhorse 3mo agoI'm really starting to tire of people making broad, general claims about how LLMs work or how to use them with N = 1 or 2. An LLM is a statistics machine for goodness sake. Basically any general claim about them needs to exploit the law of large numbers to be even remotely sensible. You cannot extrapolate from one-off behavioral successes. LLMs are not understanding anything in the way humans do. If they did, yeah, maybe you could extrapolate hard from small samples, but they don't work or understand things like we do. You need to show that the behavior you are documenting is an average behavior the LLM converges toward in the long run.
- Terretta 3mo agoyou have a premise at the heart of that: > understanding anything in the way humans do i'm not sure it's clearly established LLMs can't be a model of some part of “the way humans do”? to be more specific, i'd argue LLMs “understand” awfully similarly to a brilliant (polymath) with early dementia or Alzheimer's no executive function, no short term memory, and absent both of those, conversing with that person about the past or with an LLM about topics that had been "in their training sets before a cutoff date" is surprisingly similar, right down to: introing a topic precisely the same way, you'll experience the same conversation; convo loops if bits are too quantized (looping and lossiness / recall / context-length are correlated); and ofc opening a new session is like the first one never happened
- voidhorse 3mo ago> to be more specific, i'd argue LLMs “understand” awfully similarly to a brilliant (polymath) with early dementia or Alzheimer' At which point i'd argue that you're kind of refuting your own claim as most humans are not brilliant polymaths with early onset alzheimer's. Even admitting these kinds of comparisons to edge case human mental experience, there are more differences than similarities, and the similarities are misleading and superficial. There are still a lot of differences, even when it comes to the physical structure of the brains neurons as compared to digital neural nets, and i think it's far more beneficial to try to understand these machine in their uniqueness and for what they are than to draw hasty and shallow comparisons. You could be forgiven for that, though, because computing is rife with people who love to draw hasty unjustified analogies for some reason (example: people were already likening the brain to a computer when we didn't even have working implementations of neural nets yet and a computer was literally just a small number of logic gates lol)
- bdcravens 3mo agoI frequently blur the line between ad-hoc DSL and pseudocode, and just hand it off to the LLM. I want to get the thoughts out of my head as fast as possible, using whatever structure makes sense to me. Even if you know all of the code to be written, I think this is a huge win with LLMs, where your intent is more important than syntax.
- wsdn 3mo agoThis is exactly it. Typing was always just the interface. The keyboard was never the scarce resource. Judgment was. Architecture scales. Implementation accumulates.
- conartist6 3mo ago"DSL" is a stupid word. If it's a language, it's a language.
- systems 3mo agoits a specific type of language, not all language are the same , the nature of that specificity is the main point of the the claim made in this article
- conartist6 3mo agoall languages are embedded in some domain, therefore having a domain does not make a language "a specific type of language"
- cestith 3mo agoMany languages are in fact designed and publicized as general-purpose programming languages.
- dare944 3mo agoPeople have settled on the term DSL to identify those languages whose domains are highly constrained as compared to more general purpose or broad domain languages. This usage exists whether or not you think the term itself is stupid.
- conartist6 3mo agoI know what they meant to do, but they've drawn a line where no line exists and the result is confusion.
- dare944 3mo agoI don't see much confusion
- capestart 3mo ago[dead]
- bob1029 3mo agoThe only thing which enables reliable use of LLMs are statistical techniques. Even the most constrained and well-designed Disney world ride will break down in some embarrassing way every now and again. As you increase the # of parallel rides, the chances that at least one of them will touch the desired parts of the search space go up dramatically. The fact that the major model providers keep publishing nano/mini/luna variants should be a massive hint that there's more to this than one big fat loop magically one-shotting everything.
- tclancy 3mo agoThis is a statistical technique, effectively: it shrinks the problem space the LLM faces at each decision point. Ideally, it would be a way to turn a request from an open-world game into a game on rails. Given the DSL, my options are only X or Y at this point in the solution. Admittedly, that’s just improving the likelihood of getting a successful result, but if you mean 100% when you say “reliable”, that’s a false equivalence. No coder gets it right reliably either.
- ako 3mo agoExactly, and a dsl gives you many options and tools to make the llm more reliable: a grammar enables deterministic syntax checker, error messages injecting correct syntax into the context, linting, terse examples in the skills for token optimization, and even the option to add the grammar to the llm to help it decide what next tokens are acceptable.
- giovannibonetti 3mo agoThis reminds me of this Bjarne Stroustrup's Rule (creator of C++): - For new features, people insist on loud, explicit syntax. - For established features, people want terse notation Hillel Wayne [1] argues that the same applies for the differences between what beginners and experts desire from a language: Beginners need explicit syntax, experts want terse syntax. In my mind, DSLs are related to that – a short notation to avoid repetition. And LLMs are the experts. I wonder if Lisp with its powerful DSL-creating macros will enjoy more popularity in the near future. [1] https://buttondown.com/hillelwayne/archive/stroustrups-rule/ https://buttondown.com/hillelwayne/archive/stroustrups-rule/
- kazinator 3mo agoPeople mainly want loud, explicit syntax for new features that other people will start using, and that they don't like or want.
- wood_spirit 3mo agoRather than DSLs I’ve found careful force tools results force the same kind of discipline in a more straightforward way to implement. So it’s normal to ask the llm to answer only yes or no and they are pretty good at following that instruction but it doesn’t scale so well. Whereas if the shape of the force tool call gives them more richness without giving them freedom to go off piste it scales to more nuanced results whilst also being trivial to parse.
- dublecc 3mo ago[flagged]
- jaaacckz 3mo agoThe return of Ruby =]
- sneefle 3mo ago[flagged]
- persedes 3mo agoAs much as I dislike them [1], DSLs are an easy way to provide domain specific backpressure to your llm which might not easily be available in your language. 1- https://mikehadlow.blogspot.com/2012/05/configuration-complexity-clock.html https://mikehadlow.blogspot.com/2012/05/configuration-comple...
- BatteryMountain 3mo agoWhat I do for my one framework (in dotnet, strong typing & compiler) is to create claude skills (bash scripts and dotnet console apps) that claude can call to do various things within the framework. I have a template engine (for classes/db scripts/view engine templates), build tools, deployment tools, backup tools, infra docs - all of it works amazingly well. I built them in a standardized way so it's easy to chain them together and claude just 'gets it'. It feels like a DSL on steroids. edit: this is all on linux + posgresql.
- ActionHank 3mo agodotnet on linux friend, what are you using for templating?
- pullrun 3mo ago[flagged]
- kstenerud 3mo agoI discovered this a few months ago when I was describing a binary data format using Dogma. On a whim, I tried handing the task off to Claude and it not only finished it correctly, but also fixed some errors!
- livecontext_ai 3mo ago[flagged]
- rudedogg 3mo agoLogically this makes sense, but in practice it doesn’t, at least in my experience with SwiftUI. The LLM isn’t any better about generating/understanding it. While the DSL is more formal than natural language, it’s not what we’re communicating to the LLM with, so it’s advantages are washed away. And typical code is more strict/rigorous than DSLs so I think that’s why I see worse results, because a typical languages compiler “catches” more mistakes, versus a DSL that’s easy to write but has lots of implicitness. I’ve had the same journey experimenting with levels if abstraction too. Going lower, and exposing the LLM to the “full-stack” works much better than trying to build up abstractions it can’t see into without extra steps. I don’t want to be too much of a hater, but these types of panacea/architecture posts are usually written by people who don’t work in the field, lack pressure or constraints, and get paid to goof around in castles of the mind. I would simply skip over it and hold my comments/opinions to myself, but they tend to have an outsized influence on software engineering practices.
- sheepscreek 3mo agoI’ve had great luck building SwiftUI apps with GPT-5.5 (now GPT-5.6 Sol), Opus 4.8 and Fable 5. I’m just offering another data point, not suggesting on the effectiveness of DSLs for this. One line of thinking can be that frontier models are already powerful enough to brute force through it, but this might be a stronger indicator for success for smaller models.
- rudedogg 3mo agoNo argument that it works, it definitely does. I just think the LLMs have a hard time with tweaks/modifications of any kind, more so than what I see even in a custom GUI toolkit I’ve built that it has zero training data on. Asking for simple changes to SwiftUI interfaces I’ve one-shot is always a frustrating experience for me
- Terretta 3mo ago> simple changes to SwiftUI interfaces I’ve one-shot is always a frustrating experience for me 1) Have you exported Apple docs from Xcode and provided those to your LLM and told your CLAUDE.md to read them as applies? 2) Have you added https://github.com/dpearson2699/swift-ios-skills https://github.com/dpearson2699/swift-ios-skills to your agent? I'd also recommend Xcode 27 beta unless you teach your agent the Xcode MCP. Xcode 27 beta works on MacOS 26.x.
- skydhash 3mo agoThat’s basically LISP 101. Before solving a problem, you build out the spare parts and the tooling. And then the software building feels like assembling lego blocks. Building a DSL or a good set of symbols (functions/classes/constants/enum/…) is the cornerstone of DDD. The actual implementation details only matters at the coding stage. At the design stage, it’s better to define the glossary and its semantic.
- Incipient 3mo agoBlimey there is always a new thing to learn/use/know with AI - keeping up is getting tiring.
- beardedwizard 3mo agoSo, formal verification then...
- leemoore 3mo agoThe post would have been stronger and more useful focusing and expanding on the abstraction portion of his message as well as system and program design. Those elements were there but seem sidelined by the focus on DSL. For most developers on most projects, the key take away is we all have to both get good at system and program design especially around abstractions, shapes, responsilibility. Then we have to get good at inserting our selves in the AI loop to steer those decisions a lot in the beginning, regularly as the project kicks off and still more than we think as the project matures.
- dominotw 3mo ago> PlantUML, Mermaid, and Graphviz are domain specific languages for visual modeling; SQL is a DSL for querying databases; Kubernetes YAML is a DSL for describing cloud infrastructure. Arent these all in training data too?
- juancn 3mo agoI already had a DSL and with a proper prompt and a checker tool, even tiny LLMs could build really good scripts in that language. Gemma 12B QAT was excellent. Actually the tool-calling convention is a small DSL too. I suspect the DSL ideally has to be similar to something in the LLM's training set though. It doesn't have to feel too alien, otherwise the description of it has to be thorough and will eat up a lot of context. Frontier models don't suffer as much of this limitation since they can grab onto a larger corpus of knowledge, but they're expensive.
- camgunz 3mo agoCool looking forward to reams of AI slop in a language I can't possible know. A reasonable way to live your life is to do the opposite of whatever Thoughtworks says you should do.
- james_ross 3mo agoThis article takes me back to the Martin Fowler book many years ago that for some reason was far less popular than most of his were at the time on DSLs. Reading the DDD book around the same time plus Novak’s Learning How to Learn was the conglomerate inspiration for creating the simplest possible DSL for capturing any domain’s ubiquitous language. it’s simply concepts connected by predicates, one sentence per line in a plain text file, and while it works well to keep LLMs conceptually aligned, they do get big quickly and need a tool to visualise and edit them so I built one: https://thinkingtools.software/concepticon/ https://thinkingtools.software/concepticon/
- michaelsmanley 3mo agoAnecdotal, but I have built from scratch a fairly large OpenAPI-compliant service using the Goa dsl (https://goa.design https://goa.design) and a coding agent and I've suspected for a while that the reason I've been as productive as I have with few blind alleys is that Goa kept the mistake surface pretty constrained. I also like not having the model generate boilerplate, as Goa's tooling does all that so much more cheaply. YMMV.
- orbitalventures 3mo agoI Agree 100% with this approach, we are discovering that discrete languages are exceptionally good for LLM to deal with them, classic verbs and recommended simple sentences, we have been having quite a good time using it as replacement of Terraform for managing infrastructure and the results are quite promising.