10 ms·
Kotlin creator's new language: talk to LLMs in specs, not English
- rcvassallo83 7mo agoIts early for April fools
- dang 7mo ago"Don't be snarky." "Please don't post shallow dismissals, especially of other people's work. A good critical comment teaches us something." https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- rcvassallo83 7mo agoThanks for the reminder
- gritzko 7mo agoSo is it basically Markdown? The landing does not articulate, unfortunately, what the key contribution is.
- matthewkayin 7mo agoI tried looking through some of the spec samples, and it was not clear what the "language" was or that there was any syntax. It just looks like a terse spec.
- oceanwaves 7mo agoIn my building and research of Simplex, specs designed for LLM consumption don't need a formalized syntax as much as they just need an enforced structure, ideally paired with a linter. An effective spec for LLMs will bridge the gap between natural language and a formal language. It's about reducing ambiguity of intent because of the weaknesses and inconsistencies of natural language and the human operator.
- pjmlp 7mo agoI think stuff like Langflow and n8n are more likely to be adopted, alongside with some more formal specifications.
- ljlolel 7mo agoGetting so close to the idea. We will only have Englishscripts and don’t need code anymore. No compiling. No vibe coding. No coding. Https://jperla.com/blog/claude-electron-not-claudevm
- pure-orange 7mo agothis will have to compile to something tho? So there will always be code
- ljlolel 7mo agoNo!
- lich_king 7mo agoWe built LLMs so that you can express your ideas in English and no longer need to code. Also, English is really too verbose and imprecise for coding, so we developed a programming language you can use instead. Now, this gives me a business idea: are you tired of using CodeSpeak? Just explain your idea to our product in English and we'll generate CodeSpeak for you.
- theK 7mo agoDamn, I am the product A-GAIN?
- amelius 7mo agoCOBOL?
- souvlakee 7mo agoNo joke. I'm 100% sure that if it's successful, we will find CC's skill to write specs for CodeSpeak.
- lucasoshiro 7mo agoYeah. It's hard to express and understand nested structures in a natural language yet they are easy in high-level programming languages. E.g. "the dog of first son of my neighbour" vs "me.neighbour.sons[0].dog", "sunny and hot, or rainy but not cold" vs "(sunny && hot) || (rainy && !cold)". In the past maths were expressed using natural language, the math language exists because natural language isn't clear enough.
- lich_king 7mo agoDid you mean AbstractNeighborDispatcherFactory?
- Sharlin 7mo agoI'm sure that this time the language will be simple and English-like enough that execs can use it directly, similarly to COBOL and SQL.
- cesarvarela 7mo agoInstead of using tabs, it would be much better to show the comparison side by side. Also, the examples feel forced, as if you use external libraries, you don't have to write your own "Decode RFC 2047"
- roxolotl 7mo agoThis doesn't seem particularly formal. I still remain unconvinced reducing is really going to be valuable. Code obviously is as formal as it gets but as you trend away from that you quickly introduce problems that arise from lack of formality. I could see a world in which we're all just writing tests in the form of something like Gherkin though.
- xhkkffbf 7mo agoWhile I'm also a bit skeptical, I think some formalism could really simplify everything. The programming world has lots of words that mean close to the same thing (subroutine, method, function, etc. ). Why not choose one and stick to it for interactions with the LLM? It should save plenty of complexity.
- tasuki 7mo ago> I could see a world in which we're all just writing tests in the form of something like Gherkin though. Yes, and the implementation... no one actually cares about that. This would be a good outcome in my view. What I see is people letting LLMs "fill in the tests", whereas I'd rather tests be the only thing humans write.
- newsoftheday 7mo ago> Yes, and the implementation... no one actually cares about that. There has been a profession in place for many decades that specifically addresses that...Software Engineering.
- hrmtst93837 7mo ago[flagged]
- eterps 7mo ago> I could see a world in which we're all just writing tests in the form of something like Gherkin though. That works great in practice, Gherkin even has a markdown dialect [1]. If you combine it with a tool like aico [2] you can have a really effective development workflow. [1] https://github.com/cucumber/gherkin/blob/main/MARKDOWN_WITH_GHERKIN.md https://github.com/cucumber/gherkin/blob/main/MARKDOWN_WITH_... [2] https://github.com/jurriaan/aico https://github.com/jurriaan/aico
- theoriginaldave 7mo agoI for one can't wait to be a confident CodeSpeak programmer /sarc Does this make it a 6th generation language?
- h4ch1 7mo agoYou can basically condense this entire "language" into a set of markdown rules and use it as a skill in your planning pipeline. And whatever codespeak offers is like a weird VCS wrapper around this. I can already version and diff my skills, plans properly and following that my LLM generated features should be scoped properly and be worked on in their own branches. This imo will just give rise to a reason for people to make huge 8k-10k line changes in a commit.
- stephbook 7mo ago> And whatever codespeak offers is like a weird VCS wrapper around this. I'm still getting used to the idea that modern programs are 30 lines of Markdown that get the magic LLM incantation loop just right. Seems like you're in the same boat.
- amelius 7mo agoI want to see an LLM combined with correctness preserving transforms. So for example, if you refactor a program, make the LLM do anything but keep the logic of the program intact.
- whalesalad 7mo agohttps://en.wikipedia.org/wiki/Literate_programming https://en.wikipedia.org/wiki/Literate_programming
- tonipotato 7mo agoThe problem with formal prompting languages is they assume the bottleneck is ambiguity in the prompt. In my experience building agents, the bottleneck is actually the model's context understanding. Same precise prompt, wildly different results depending on what else is in the context window. Formalizing the prompt doesn't help if the model builds the wrong internal representation of your codebase. That said curious to see where this goes.
- slfnflctd 7mo agoTwo pieces of advice I keep seeing over & over in these discussions-- 1) start with a fresh/baseline context regularly, and 2) give agents unix-like tools and files which can be interacted with via simple pseudo-English commands such as bash, where they can invoke e.g. "--help" to learn how to use them. I'm not sure adding a more formal language interface makes sense, as these models are optimized for conversational fluency. It makes more sense to me for them to be given instructions for using more formal interfaces as needed.
- kleiba 7mo agoI cannot read light on black. I don't know, maybe it's a condition, or simply just part of getting old. But my eyes physically hurt, and when I look up from reading a light-on-black screen, even when I looked at only for a short moment, my eyes need seconds to adjust again. I know dark mode is really popular with the youngens but I regularly have to reach for reader mode for dark web pages, or else I simply cannot stand reading the contents. Unfortunately, this site does not have an obvious way of reading it black-on-white, short of looking at the HTML source (CTRL+U), which - in fact - I sometimes do.
- embedding-shape 7mo agoDo you sit in a bright room? Right now, during the night, I see your comment like this: https://i.imgur.com/c7fmBns.png https://i.imgur.com/c7fmBns.png, but during the day when the room is bright, I also see everything with light themes/background colors, otherwise it is indeed hard to see properly.
- kleiba 7mo agoUnfortunately, in my case, it's not a matter of lighting conditions.
- skydhash 7mo agoWhen it’s dark (I can’t stand bright rooms at night), I lower the brightness of my screens instead of going for dark mode. I have astigmatism and any tiny bright spot is hard to focus on. It’s easier when the bright part is large and the dark parts are small (black on white is best).
- newsoftheday 7mo agoSame for me, has been my whole life. I complain about it all the time. It's well documented that people can read black on light far better and with less eye strain than light on black; yet there seems to be a whole generation of developers determined to force us all to try and read it. Even the media sites like Netflix, Prime, etc. force it. At least Tubi's is somewhat more readable. Sometimes a site will include a button or other UI element to choose a light theme but I find it odd that so many sites which are presumed to be designed by technically competent people, completely ignore accessibility concerns.
- xvedejas 7mo agoWe already have a language for talking to LLMs: Polish https://www.zmescience.com/science/news-science/polish-effective-prompting-ai/ https://www.zmescience.com/science/news-science/polish-effec...
- the_duke 7mo agoThis doesn't make too much sense to me. * This isn't a language, it's some tooling to map specs to code and re-generate * Models aren't deterministic - every time you would try to re-apply you'd likely get different output (without feeding the current code into the re-apply and let it just recommend changes) * Models are evolving rapidly, this months flavour of Codex/Sonnet/etc would very likely generate different code from last months * Text specifications are always under-specified, lossy and tend to gloss over a huge amount of details that the code has to make concrete - this is fine in a small example, but in a larger code base? * Every non-trivial codebase would be made up of of hundreds of specs that interact and influence each other - very hard (and context - heavy) to read all specs that impact functionality and keep it coherent I do think there are opportunities in this space, but what I'd like to see is: * write text specifications * model transforms text into a *formal* specification * then the formal spec is translated into code which can be verified against the spec 2 and three could be merged into one if there were practical/popular languages that also support verification, in the vain of ADA/Spark. But you can also get there by generating tests from the formal specification that validate the implementation.
- davedx 7mo agoMy process has organically evolved towards something similar but less strictly defined: - I bootstrap AGENTS.md with my basic way of working and occasionally one or two project specific pieces - I then write a DESIGN.md. How detailed or well specified it is varies from project to project: the other day I wrote a very complete DESIGN.md for a time tracking, invoice management and accounting system I wanted for my freelance biz. Because it was quite complete, the agent almost one-shot the whole thing - I often also write a TECHNICAL-SPEC.md of some kind. Again how detailed varies. - Finally I link to those two from the AGENTS. I also usually put in AGENTS that the agent should maintain the docs and keep them in sync with newer decisions I make along the way. This system works well for me, but it's still very ad hoc and definitely doesn't follow any kind of formally defined spec standard. And I don't think it should, really? IMO, technically strict specs should be in your automated tests not your design docs.
- the_duke 7mo ago
- jajuuka 7mo agoWe created programming languages to direct programs. Then created LLM's to use English to direct programs. Now we've create programming languages to direct LLM's. What is old is new again!
- mft_ 7mo agoConceptually, this seems a good direction. The other piece that has always struck me as a huge inefficiency with current usage of LLMs is the hoops they have to jump through to make sense of existing file formats - especially making sense of (or writing) complicated semi-proprietary formats like PDF, DOC(X), PPT(X), etc. Long-term prediction: for text, we'll move away from these formats and towards alternatives that are designed to be optimal for LLMs to interact with. (This could look like variants of markdown or JSON, but could also be Base64 [0] or something we've not even imagined yet.) [0] https://dnhkng.github.io/posts/rys/ https://dnhkng.github.io/posts/rys/
- pessimizer 7mo agoIf LLMs can't deal with those legacy file formats, I don't trust them to be able to deal with anything. The idea that LLMs are so sophisticated that we have a need to dumb down inputs in order to interact with them is self-contradictory.
- layer8 7mo agoWhile I agree, the parent also talks about efficiency. If a different format increases efficiency, that could be reason enough to switch to it, even if understanding doesn’t improve and already was good before.
- mft_ 7mo agoThank you, yes, efficiency was entirely my point. :) Humans are far more efficient when they interact with information that's in a format that suits their abilities or preferences; it seems pretty obvious that in some ways the same would likely be true for LLMs.
- deleted 7mo ago[deleted]
- fallkp 7mo ago"Coming soon: Turning Code into Specs" There you have it: Code laundering as a service. I guess we have to avoid Kotlin, too.
- newsoftheday 7mo agoI avoid Kotlin as a principal, any language that can't get the type and variable name in the correct order; I avoid them completely.
- oytis 7mo agoThen of course we are going to ask LLMs to generate specifications in this new language
- kittikitti 7mo agoThe intent of the idea is there, and I agree that there should be more precise syntax instead of colloquial English. However, it's difficult to take CodeSpeak seriously as it looks AI generated and misses key background knowledge. I'm hoping for a framework that expands upon Behavior Driven Development (BDD) or a similar project-management concept. Here's a promising example that is ripe for an Agentic AI implementation, https://behave.readthedocs.io/en/stable/philosophy/#the-gherkin-language https://behave.readthedocs.io/en/stable/philosophy/#the-gher...
- Brajeshwar 7mo agoSo, back to a programming language, albeit “simplified.”
- oceanwaves 7mo agohttps://thinkwright.ai/simplex https://thinkwright.ai/simplex
- deleted 7mo ago[deleted]
- Cpoll 7mo ago> The spec is the source of truth This feels wrong, as the spec doesn't consistently generate the same output. But upon reflection, "source of truth" already refers to knowledge and intent, not machine code.
- newsoftheday 7mo ago> not machine code Actually, computers, being machines, do equate machine code and source of truth.
- le-mark 7mo agoThis concept is assuming a formalized language would make things easier somehow for an llm. That’s making some big assumptions about the neuro anatomy if llms. This [1] from the other day suggests surprising things about how llms are internally structured; specifically that encoding and decoding are distinct phases with other stuff in between. Suggesting language once trained isn’t that important. [1] https://news.ycombinator.com/item?id=47322887 https://news.ycombinator.com/item?id=47322887
- abreslav 7mo agoWe are not trying to make things easier for LLMs. LLMs will be fine. CodeSpeak is built for humans, because we benefit from some structure, knowing how to express what we want, etc.
- lifis 7mo agoAs far as I can tell it's not a new language, but rather an alternative workflow for LLM-based development along with a tool that implements it. The idea, IIUC, seems to be that instead of directly telling an LLM agent how to change the code, you keep markdown "spec" files describing what the code does and then the "codespeak" tool runs a diff on the spec files and tells the agent to make those changes; then you check the code and commit both updated specs and code. It has the advantage that the prompts are all saved along with the source rather than lost, and in a format that lets you also look at the whole current specification. The limitation seems to be that you can't modify the code yourself if you want the spec to reflect it (and also can't do LLM-driven changes that refer to the actual code), and also that in general it's not guaranteed that the spec actually reflects all important things about the program, so the code does also potentially contain "source" information (for example, maybe your want the background of a GUI to be white and it is so because the LLM happened to choose that, but it's not written in the spec). The latter can maybe be mitigated by doing multiple generations and checking them all, but that multiplies LLM and verification costs. Also it seems that the tool severely limits the configurability of the agentic generation process, although that's just a limitation of the specific tool.
- deleted 7mo ago[deleted]
- souvlakee 7mo agoAs far as I can tell C is not a new language, but rather an alternative workflow for assembly development along with a tool that implements it.
- abreslav 7mo agoI second that :)
- abreslav 7mo ago> The limitation seems to be that you can't modify the code yourself if you want the spec to reflect it Eventually, we'll end up in a world where humans don't need to touch code, but we are not there yet. We are looking into ways to "catch up" the specs with whatever changes happen in the code not through CodeSpeak (agents or manual changes or whatever). It's an interesting exercise. In the case of agents, it's very helpful to look at the prompts users gave them (we are experimenting with inspecting the sessions from ~/.claude). More generally, `codespeak takeover` [1] is a tool to convert code into specs, and we are teaching it to take prompts from agent sessions into account. Seems very helpful, actually. I think it's a valid use case to start something in vibe coding mode and then switch to CodeSpeak if you want long-term maintainability. From "sprint mode" to "marathon mode", so to speak [1] https://codespeak.dev/blog/codespeak-takeover-20260223 https://codespeak.dev/blog/codespeak-takeover-20260223
- AlexC04 7mo agothis is really exciting and dovetails really closely with the project I'm working on. I'm writing a language spec for an LLM runner that has the ability to chain prompts and hooks into workflows. https://github.com/AlexChesser/ail https://github.com/AlexChesser/ail I'm writing the tool as proof of the spec. Still very much a pre-alpha phase, but I do have a working POC in that I can specify a series of prompts in my YAML language and execute the chain of commands in a local agent. One of the "key steps" that I plan on designing is specifically an invocation interceptor. My underlying theory is that we would take whatever random series of prose that our human minds come up with and pass it through a prompt refinement engine: > Clean up the following prompt in order to convert the user's intent > into a structured prompt optimized for working with an LLM > Be sure to follow appropriate modern standards based on current > prompt engineering reasech. For example, limit the use of persona > assignment in order to reduce hallucinations. > If the user is asking for multiple actions, break the prompt > into appropriate steps (**etc...) That interceptor would then forward the well structured intent-parsed prompt to the LLM. I could really see a step where we say "take the crap I just said and turn it into CodeSpeak" What a fantastic tool. I'll definitely do a deep dive into this.
- tamimio 7mo agoAs someone who hates writing (and thus coding) this might be a good tool, but how’s is it different from doing the same in claude? And I only see python, what about other languages, are they also production grade?
- yellow_lead 7mo agoSo, just a markdown file?
- ppqqrr 7mo agoi’ve been doing this for a while, you create an extra file for every code file, sketch the code as you currently understand it (mostly function signatures and comments to fill in details), ask the LLM to help identify discrepancies. i call it “overcoding”. i guess you can build a cli toolchain for it, but as a technique it’s a bit early to crystallize into a product imo, i fully expect overcoding to be a standard technique in a few years, it’s the only way i’ve been able to keep up with AI-coded files longer than 1500 lines
- WillAdams 7mo agoThis raises a question --- how well do LLMs understand Loglan? https://www.loglan.org/ https://www.loglan.org/ Or Lojban? https://mw.lojban.org/ https://mw.lojban.org/
- CodeCompost 7mo agoYes I'm also one of those LLM skeptics but actually this looks interesting.
- montjoy 7mo agoSo, instead of making LLMs smarter let’s make everything abstract again? Because everyone wants to learn another tool? Or is this supposed to be something I tell Claude, “Hey make some code to make some code!” I’m struggling to see the benefit of this vs. just telling Claude to save its plan for re-use.
- herrington_d 7mo agoIsn't the case study.... too contrived and trivial? The largest code change is 800 lines so it can readily fit in a model's context. However, there is no case for more complicated, multi-file changes or architecture stuff.
- leksak 7mo agoI think I prefer Tracey https://github.com/bearcove/tracey https://github.com/bearcove/tracey
- sriramgonella 7mo ago[flagged]
- mempko 7mo agoI think the magic sauce in this project is the fact that they convert diffs in spec to diffs in code, which is likely more stable than just regenerating the whole thing.
- skydhash 7mo agoThe thing is, such exploration can be done on a whiteboard or a moodboard. Once it’s we settled on a process, we code it and let the computer take over. I really believe the struggle is knowledge and communication of ideas, not the coding part (which is fairly easy IMO).
- booleandilemma 7mo agoAlas, I thought I invented this. https://news.ycombinator.com/item?id=47284030 https://news.ycombinator.com/item?id=47284030
- pshirshov 7mo agoFrom what I was able to understand during the interview there, it's not actually a language, more like an orchestrator + pinning of individual generated chunks. The demo I've briefly seen was very very far from being impressive. Got rejected, perhaps for some excessive scepticism/overly sharp questions. My scepticism remains - so far it looks like an orchestrator to me and does not add enough formalism to actually call it a language. I think that the idea of more formal approach to assisted coding is viable (think: you define data structures and interfaces but don't write function bodies, they are generated, pinned and covered by tests automatically, LLMs can even write TLA+/formal proofs), but I'm kinda sceptical about this particular thing. I think it can be made viable but I have a strong feeling that it won't be hard to reproduce that - I was able to bake something similar in a day with Claude.
- _doctor_love 7mo agoI find it weird that this comment is gray but it's the only one in the thread so far that mentions TLA+ which is highly relevant here.
- b4rtaz__ 7mo agoA few days ago I released https://github.com/b4rtaz/incrmd https://github.com/b4rtaz/incrmd , which is similar to Codespeak. The main difference is that the specification is defined at the *project* level. I'm not sure if having the specification at the *file* level is a good choice, because the file structure does not necessarily align with the class structure, etc.
- uday_singlr 7mo agoWe tend to obsess over abstractions, frameworks, and standards, which is a good thing. But we already have BDD and TDD, and now, with english as the new high-level programming language, it is easier than ever to build. Focusing on other critical problem spaces like context/memory is more useful at this point. If the whole purpose of this is token compression, I don't see myself using it.
- ternaryoperator 7mo agoAgreed. There is definitely a similarity to BDD.
- phplovesong 7mo agoThis is pretty lame. I WANT to write code, something that has a formal definition and express my ideas in THAT, not some adhoc pseudo english an LLM then puts the cowboy hat on and does what the hotness of the week is. Programming is in the end math, the model is defined and, when done correctly follows common laws.
- seanmcdirmid 7mo agoI've done something similar for queries. Comments: * Yes, this is a language, no its not a programming language you are used to, but a restricted/embellished natural language that (might) make things easier to express to an LLM, and provides a framework for humans who want to write specifications to get the AI to write code. * Models aren't deterministic, but they are persistent (never gonna give up!). If you generate tests from your specification as well as code, you can use differential testing to get some measure (although not perfect) of correctness. Never delete the code that was generated before, if you change the spec, have your model fix the existing code rather than generate new code. * Specifications can actually be analyzed by models to determine if they are fully grounded or not. An ungrounded specification is going to not be a good experience, so ask the model if it thinks your specification is grounded. * Use something like a build system if you have many specs in your code repository and you need to keep them in sync. Spec changes -> update the tests and code (for example).
- etothet 7mo agoUnder "Prerequisites"[0] I see: "Get an Anthropic API key". I presume this is temporary since the project is still in alpha, but I'm curious why this requires use of an API at all and what's special about it that it can't leverage injecting the prompt into a Claude Code or other LLM coding tool session. [0]: https://codespeak.dev/blog/greenfield-project-tutorial-20260209 https://codespeak.dev/blog/greenfield-project-tutorial-20260...
- frizlab 7mo agoThe next step will be to formalize all the instructions possible to give to a processor and use that language!
- nunobrito 7mo agoExactly as necessary as Kotlin itself.
- haspok 7mo agoI would just like to point out the fun fact that instead of the brave new MD speak, there is still a `codespeak.json` to configure the build system itself... ...which seems to suggest that the authors themselves don't dogfood their own software. Please tell me that Codespeak was written entirely with Codespeak! Instead of that json, which is so last year, why not use an agent to create an MD file to setup another agent, that will compile another MD file and feed it to the third agent, that... It is turtles, I mean agents, all the way down!
- sutterd 7mo agoI am trying a similar spec driven development idea in a project I am working on. One big difference is that my specifications are not formalized that much. Tney are in plain language and are read directly by the LLM to convert to code. That seems like the kind of thing the LLM is good at. One other feature of this is that it allows me to nudge the implmentation a little with text in the spec outside of the formal requirements. I view it two ways, as spec-to-code but also as a saved prompt. I haven't spent enough time with it to say how successfuly it is, yet.
- issamG 7mo agoDo you save these "prompts" so you can improve, and in turn improve the code. to me Spec Driven Development is more than a spec to generate code, structured or not.
- sutterd 7mo agoThe spec contains formal, numbered items which are requirements and also serve to make tests (these are spec tests, additional implementation tests are also allowed by the implementer). When I said "they are not formalized as much", I mean I am not as strict on the spec format as CodeSpeak is, where their spec can be parsed with a tool. For me it is up to the LLM to use the spec itself. I have additional text beyond the requirement items which also influences how the LLM implements the code. I did this because it is too tough, for me at least, to prompt the LLM just based on strict requirements. This is perhaps cheating according to what you might call SDD. I'm just trying to be practical. The idea in the end is that this spec implies the code and maintaining the spec is the same as maintaining the code. Strictly speaking this won't be true, but I am hoping it still works anyway.
- koolala 7mo agoLooks like JSON like YAML. It is still English. Was hoping for something like Lojban.
- good-idea 7mo ago"Shrink your codebase 5-10x" "[1] When computing LOC, we strip blank lines and break long lines into many"
- jasonjmcghee 7mo agoI don't think this is the gotcha you think it is... I imagine this is before and after- not just after. As in, they aren't just making lines long and removing whitespace (something models love to do when you ask it to remove lines of code)
- good-idea 7mo agoYep, you're right, I read this too fast - it's also breaking long lines into many and I read this in reverse. I just imagined how much I could reduce my own LOC by adjusting the print width on my prettier settings..
- aplomb1026 7mo ago[flagged]
- semessier 7mo agoit's not a new question if the as-is programming languages are optimal for LLMs: a language for LLM use would have to strongly typed. But that's about it for obvious requirements.
- ivanjermakov 7mo agoAnother great way to shrink your codebase 10x? Rewrite it in APL. If less code means less information, what are we gonna do when missing information was important?
- mgax 7mo agoGood code is the specification.
- iLoveOncall 7mo agoThe tweet I saw a few weeks ago about LLMs enabling building stupid ideas that would have never been built otherwise particularly resonates with this one.
- deleted 7mo ago[deleted]
- deleted 7mo ago[deleted]
- paxys 7mo agoI read through the thing and don't quite understand what this adds that the dozens of LLM coding wrappers don't already do. You write a markdown spec. The script takes it and feeds it to an LLM API. The API generates code. Okay? Where is this "next-generation programming language" they talk about?
- hmokiguess 7mo agoI'm gonna be honest here, I opened this website excited thinking this was a sort of new paradigm or programming language, and I ended up extremely confused at what this actually is and I still don't understand. Is it a code generator tool from specs? Ugh. Why not push for the development of the protocol itself then?
- weezing 7mo agoI'll stick to Polish
- taintlord 7mo ago[dead]
- petetnt 7mo agoBuddy invented RobotFramework, great job.
- wuweiaxin 7mo agoThe pattern we keep converging on is to treat model calls like a budgeted distributed system, not like a magical API. The expensive failures usually come from retries, fan-out, and verbose context growth rather than from a single bad prompt. Once we started logging token use per task step and putting hard ceilings on planner depth, costs became much more predictable.
- BrianFHearn 7mo ago[flagged]
- colordrops 7mo agoIsn't that the point though? In the development loop, you'd diagnose why it's not building what you expect, so you flush out those previous implicit or even subconscious edge cases, undocumented behaviors, and tribal knowledge and codify them into the spec. It would actually end up being a lot easier to maintain than a bunch of undocumented spaghetti.
- deleted 7mo ago[deleted]
- sornaensis 7mo agoThis seems like a step backwards. Programming Languages for LLMs need a lot of built in guarantees and restrictions. Code should be dense. I don't really know what to make of this project. This looks like it would make everything way worse. I've had good success getting LLMs to write complicated stuff in haskell, because at the end of the day I am less worried about a few errant LLM lines of code passing both the type checking and the test suite and causing damage. It is both amazing and I guess also not surprising that most vibe coding is focused on python and javascript, where my experience has been that the models need so much oversight and handholding that it makes them a simple liability. The ideal programming language is one where a program is nothing but a set of concise, extremely precise, yet composable specifications that the _compiler_ turns into efficient machine code. I don't think English is that programming language.
- ucyo 7mo agoLiterally the first example on the main page declared as code.py would result in an indentation error :)
- riantogo 7mo agoWhen we understand that AI allows the spec to be in English (or any natural language), we might stop attempting to build "structured english" for spec.
- pcblues 7mo agoA formal way for a senior to tell AI (clueless junior) to do a senior's job? Once again, who checks and fixes the output code? Of course an expert would throw it out and design/write it properly so they know it works.
- niam 7mo agoThe title writer might be doing the project a disservice by using the term "formal" to describe it, given that the project talks a lot about "specs". I mistook it to imply something about formal specification. My quick understanding is that isn't really trying to utilize any formal specification but is instead trying to more-clearly map the relationship between, say, an individual human-language requirement you have of your application, and the code which implements that requirement.
- giantg2 7mo agoThis is basically what I talked about maybe a year ago. Glad to see someone is taking it on.
- temp123789246 7mo agoOne requirement for a programming language to be “good” is that doing this, with sufficient specificity to get all the behavior you want, will be more verbose than the code itself.
- photios 7mo ago> codespeak login Instant tab close!
- phyzix5761 7mo agoHuh? I don't see a login. All the code examples and set up instructions are available to anyone visiting the page.
- photios 7mo agoIt's in the Quickstart article here: https://codespeak.dev/blog/greenfield-project-tutorial-20260209 https://codespeak.dev/blog/greenfield-project-tutorial-20260...
- neopointer 7mo agoThe next step is to use AI to edit the spec... /s
- oofbaroomf 7mo agoUgh, I just wish there was a deterministic and formal way to tell a computer what I want...
- haolez 7mo agoFeels like writing tests before writing code, but with LLMs :)
- cube2222 7mo agoThis is actually... pretty cool? Definitely won't use it for prod ofc but may try it out for a side-project. It seems that this is more or less: - instead of modules, write specs for your modules - on the first go it generates the code (which you review) - later, diffs in the spec are translated into diffs in the code (the code is *not* fully regenerated) this actually sounds pretty usable, esp. if someone likes writing. And wherever you want to dive deep, you can delve down into the code and do "microoptimizations" by rolling something on your own (with what seems to be called here "mixed projects"). That said, not sure if I need a separate tool for this, tbh. Instead of just having markdown files and telling cause to see the md diff and adjust the code accordingly.
- abreslav 7mo agoWe'd love to hear your feedback! Feel free to come to our discord to ask questions/share experience: https://l.codespeak.dev/discord https://l.codespeak.dev/discord
- jaredklewis 7mo agoTHe HN title seems very misleading to me. How is this, in any sense of the word, "formal?" I don't see that particular word used to describe this tool on the web page itself. The site does describe it as a "programming language," which feels like a novel use of the term to me. The borders around a term like "programming language" are inherently fuzzy, but something like "code generation tool" better describes CodeSpeak IMHO.
- dang 7mo agoOk we've deformalized the title above.
- Garlef 7mo agoI think this is 100% the right direction: Instead of imperatively letting the agents hammer your codebase into shape through a series of prompts, you declare your intent, observe the outcome and refine the spec. The agents then serve as a control plane, carrying out the intent.
- abreslav 7mo agoVery much agree. I like the imperative vs declarative angle you take here. Thank you!
- vybandz 7mo agoSounds like crap.
- siscia 7mo agoWhat I found more useful is an extra step. Spec to tests, and then red tests to code and green tests. LLMs works on both translation steps. But you end up with an healthy amount of tests. I tagged each tests with the id of the spec so I do get spec to test coverage as well. Beside standard code coverage given by the tests.
- abreslav 7mo agoVery much agree on coverage. We're actually doing something in that area: https://codespeak.dev/blog/coverage-20260302 https://codespeak.dev/blog/coverage-20260302 For now, it's only about test coverage of the code, but the spec coverage is coming too.
- siscia 7mo agoI think you guys are doing pretty much everything right.
- abreslav 7mo agoWhen you translate spec to tests (if those are traditional unit tests or any automated tests that call the rest of the code), that fixes the API of the code, i.e. the code gets designed implicitly in the test generation step. Is this working well in your experience?
- siscia 7mo agoYes it is passable. Good enough that I don't review it. Granted, it is a personal project that I care only to the point that I want it to work. There are no money on the line. Nothing professional. I believe that part of the secret is that I force CC to run the whole est suites after it change ANY file. Using hooks. It makes iteration slower because it kinda forces it to go from green to green. Or better from red to less red (since we start in red). But overall I am definitely happy with the results. Again, personal projects. Not really professional code.
- 7mo ago
- stephbook 7mo ago> It’s one of those things that crackpots keep trying to do, no matter how much you tell them it could never work. If the spec defines precisely what a program will do, with enough detail that it can be used to generate the program itself, this just begs the question: how do you write the spec? Such a complete spec is just as hard to write as the underlying computer program, because just as many details have to be answered by spec writer as the programmer. Joel Spolsky, stackoverflow.com founder, Talk at Yale: Part 1 of 3 https://www.joelonsoftware.com/2007/12/03/talk-at-yale-part-1-of-3/ https://www.joelonsoftware.com/2007/12/03/talk-at-yale-part-...
- CamperBob2 7mo agoWhat he misses is that it's much easier to change the spec than the code. And if the cost of regenerating the code is low enough, then the code is not worth talking about.
- deleted 7mo ago[deleted]
- mwarkentin 7mo agoIs it? If the spec is as detailed as the code would be? If you make a change to one part of the spec do you now have inconsistencies that the LLM is going to have to resolve in some way? Are we going to have a compiler, or type checker type tools for the spec to catch these errors sooner?
- CamperBob2 7mo agoIt IS a compiler. You might as well ask if the machine-language output of a C compiler is as detailed as the C code was. To anticipate your objection: you can get over determinism now, or you can get over it later. You will get over it, though, if you intend to stay in this business.
- discreteevent 7mo ago
- nunez 7mo agoBDD reincarnated!!!!
- jonstaab 7mo agoIs it open source? This is a cool idea, but I'm pretty sure it's probably just a thin wrapper around claude. I also couldn't install it on my headless dev box because it relies on a localhost callback. Well, I'm looking forward to the first open source version in about 10 minutes.
- modernerd 7mo agoSo it has "two-way conversion": `codespeak build` — takes the spec and turns it into code via LLM, like a non-deterministic compiler. `codespeak takeover` — reads a file and creates a spec from it. You can progressively opt in ("mixed mode") so it only touches files you allow it to (and makes new ones if needed). Pros: - Formalised version of the "agentic engineering" many are already doing, but might actually get people to store their specs and decisions in a concise way that seems more sane than committing your entire meandering chat session. - Encouraging people to review spec and code side-by-side at a file level seems reasonable. Could even build an IDE/plugin around that concept to auto-load/navigate the spec and code side-by-side like their examples: https://codespeak.dev/shrink-factor/markitdown-eml https://codespeak.dev/shrink-factor/markitdown-eml. If tokens per second for popular models continues to improve, could even update the spec by hand and see the code regenerate live on the fly, perhaps via `codespeak watch`. - Reduces the code you have to write by 5-10x. Largely by convincing you not to write it any more. Our graphics cards write the code for us in this timeline and many people are even happy about it. - As models improve, could optionally re-run `build` against the same original spec. (Why do that if the output already produces the intended result and the test suite still passes? Presumably for simpler code. Or faster output. Or lower memory use. Or simply _different_ bugs.) - Moves programming back toward structured thinking backed by a committed artifact and a solid two-word command you can run, instead of actively having conversations with far away GPUs like that's normal now. - Could theoretically swap out the build target language if you grow to trust the build process to be your babelfish/specfish. Kind of Haxe with Markdown. Cons: - Seems to be gated by their login, can't bring your own model? - Suspect the labs can all clone this concept very easily. `claude build` and `claude spec`? The idea of a non-deterministic 'build' command had me cringing at first. But formalising a process many are using anyway that currently feels pretty sloppy perhaps isn't so terrible. If nothing else, writing `build` is a lot quicker and maintains a whisker of self-respect. At least compared to typing, "please take this spec and adapt the Python accordingly" followed 2 minutes later by, "I updated the spec to deal with the edge-case you missed, try again but don't miss anything this time".
- codethief 7mo ago> Case studies: > - Encoding auto-detection and normalization for beautifulsoup4 I was kinda expecting to see the name "chardet" pop up here. :-)
- Steinmark 7mo ago[dead]
- p0u4a 7mo agoWe're entering the area of stochastic software, everybody's a gambler
- labrador 7mo agoSeems like a lot of busy work to me
- deleted 7mo ago[deleted]
- ncr100 7mo agoFrom an inclusivity perspective, can more people than "programmers" be enlisted to write specs? We are putting people out of work. Why not employ MORE people to do LESS, by sharing the responsibility? A group activity, perhaps? Eg make room in this spec > program development workflow for, say, ... Tech Writers. Add them to the development team to ensure the language is right for the LLM ahead of time!
- liampulles 7mo agoI think I want to know exactly what SQL ends up hitting the DB, and I want to fine tune it precisely. This is the same issue I've had with ORMs - I get that they make it easier to generate functionality at speed, but ultimately I want control over the biggest performance lever I have available to me.
- dakial1 7mo agoSomewhat related but I always wondered if I asked a LLM to create a new language with full focus on LLM coding efficiency, ignoring the need for humans to read it, what would it come with? Binary? ...and I obviously asked Gemini about it and it replied: "A language optimized exclusively for Large Language Model (LLM) efficiency would prioritize Token Density, Context Window Management, and Architectural Alignment. It would not be binary, as standard LLM architectures (Transformers) process discrete tokens from a predefined vocabulary, not raw bits." Example of it: Feature Human-Readable (Python/C++) LLM-Native (Hypothetical) -------------------------------------------------------------------------- Logic if (x > 10) { return true; } ¿x10† Memory int\* ptr = malloc(sizeof(int)); §m4 Tokens Used ~10-15 2-3
- phpnode 7mo agoThe issue with these LLM-targeting DSLs is that you have to waste a bunch of your context window explaining the grammar and semantics to the LLM, whereas they already speak existing programming languages because they've seen so much existing code. This usually negates the benefits of the DSL.
- unsaved159 7mo agoNot clear to me why need this. You can just write a markdown spec without any side projects, then tell an agent to code it.
- peter_d_sherman 7mo agoThere's an interesting set of ideas here! If we look at the history of programming languages, we see the idea of Templating occuring over and over again, in different contexts, i.e., C's macros, C++ Templates, embedding PHP code snippets into an otherwise mostly HTML file, etc., etc. Templating can involve aspects of meta-code (code about the code), interpretation proxying (which engine/compiler/system/parser/program/subsystem/? is responsible for interpreting a given section of text), etc., etc. Here we see this idea as another level of proxied/layered abstraction/indirection, in this case between an AI/LLM and the underlying source code... Is this a good idea? Will all code be written like this, using this pattern or a similar one, in the future? I for one don't know (it's too early to tell!) but one thing is for sure, and that's that this new "layer" certainly contains an interesting set of ideas! I will definitely be watching to see more about how this pattern plays out in future software development...