13 ms·
We should revisit literate programming in the agent era
- sublinear 7mo ago> This is especially important if the primary role of engineers is shifting from writing to reading. This was always the primary role. The only people who ever said it was about writing just wanted an easy sales pitch aimed at everyone else. Literate programming failed to take off because with that much prose it inevitably misrepresents the actual code. Most normal comments are bad enough. It's hard to maintain any writing that doesn't actually change the result. You can't "test" comments. The author doesn't even need to know why the code works to write comments that are convincing at first glance. If we want to read lies influenced by office politics, we already have the rest of the docs.
- c0rp4s 7mo agoYou're right that you can't test comments, but you can test the code they describe. That's what reproducibility bundles do in scientific computing ;; the prose says "we filtered variants with MAF < 0.01", and the bundle includes the exact shell command, environment, and checksums so anyone can verify the prose matches reality. The prose becomes a testable claim rather than a decorative comment. That said, I agree the failure mode of literate programming is prose that drifts from code. The question is whether agents reduce that drift enough to change the calculus.
- ares623 7mo agoI don't buy that. Writing is taking a bad rap from all this. Writing _is_ a form of more intense reading. Reading on steroids, as they say. If reading is considered good, writing should be considered better.
- bigyabai 7mo agoWriting in that draft style is really only useful because a) you read the results and b) you write an improved version at the end. Drafting forever is not considered "better" because someone (usually you) has to sift through the crap to find the good parts. This is especially pronounced in the programming workplace, where the most "senior" programmers are asked to stop programming so they can review PRs.
- 8note 7mo ago> You can't "test" comments. I'm thinking that we're approaching a world where you can both test for comments and test the comments themselves.
- senderista 7mo agoNow that would be really interesting: prompt an LLM to find comments that misrepresent the code! I wonder how many false positives that would bring up?
- ccosky 7mo agoI have a Claude Code skill for adding, deleting and improving comments. It does a decent job at detecting when comments are out of date with the code and updating them. It's not perfect, but it's something.
- hrmtst93837 7mo ago[flagged]
- perrygeo 7mo agoConsidering LLMs are models of language, investing in the clarity of the written word pays off in spades. I don't know whether "literate programming" per se is required. Good names, docstrings, type signatures, strategic comments re: "why", a good README, and thoughtfully-designed abstractions are enough to establish a solid pattern. Going full "literate programming" may not be necessary. I'd maybe reframe it as a focus on communication. Notebooks, examples, scripts and such can go a long way to reinforcing the patterns. Ultimately that's what it's about: establishing patterns for both your human readers and your LLMs to follow.
- crazygringo 7mo agoYeah, I think what is needed is somewhere between docstrings+strategic comments, and literate programming. Basically, it's incredibly helpful to document the higher-level structure of the code, almost like extensive docstrings at the file level and subdirectory level and project level. The problem is that major architectural concepts and decisions are often cross-cutting across files and directories, so those aren't always the right places. And there's also the question of what properly belongs in code files, vs. what belongs in design documents, and how to ensure they are kept in sync.
- amelius 7mo agoAlso: "Bad programmers worry about the code. Good programmers worry about data structures and their relationships." -- Linus Torvalds
- rustybolt 7mo agoI have noticed a trend recently that some practices (writing a decent README or architecture, being precise and unambiguous with language, providing context, literate programming) that were meant to help humans were not broadly adopted with the argument that it's too much effort. But when done to help an LLM instead of a human a lot of people suddenly seem to be a lot more motivated to put in the effort.
- zdragnar 7mo agoIn my years of programming, I find that humans rarely give documentation more than a cursory glance up until they have specific questions. Then they ask another person if one is available rather than read for the answer. The biggest problem is that humans don't need the documentation until they do. I recall one project that extensively used docblock style comments. You could open any file in the project and find at least one error, either in the natural language or the annotations. If the LLM actually uses the documentation in every task it performs- or if it isn't capable of adequate output without it- then that's a far better motivation to document than we actually ever had for day to day work.
- ijk 7mo agoI have discovered that the measure of good documentation is not whether your team writes documentation, but is instead determined by whether they read it.
- suzzer99 7mo agoThe other problem is that documentation is always out of date, and one wrong answer can waste more time than 10 "I don't knows".
- 1718627440 7mo agoI think this really depends on culture. If you target OS APIs or the libc, the documentation is stellar. You have several standards and then conceptual documentation and information about particular methods all with historic and current and implementation notes, then there is also an interactive hypertext system. I solve 80% of my questions with just looking at the official documentation, which is also installed on my computer. For the remaining I often try to use the WWW, but these are often so specific, that it is more successful to just read the code. Once I step out of that ecosystem, I wonder how people even cope with the lack of good documentation.
- gervwyk 7mo agoFor me this is where a config layer shines. Develop a decent framework and then let the agents spin out the configuration. This allows a trusted and tested abstraction layer that does not shift and makes maintenance easier, while making the code that the agents generate easier to review and it also uses much less tokens. So as always, just build better abstractions.
- cyanydeez 7mo agowhen do you think we'll get to build real software?
- jauntywundrkind 7mo agoI fully agree. (Seeing how good Figment2 is for layered config in rust is wildly eye opening, has been a revelatory experience.) Sometimes what we manage with config is itself processing pipelines. A tool like darktable has a series of processing steps that are run. Each of those has config, but the outer layer is itself a config of those inner configs. And the outer layer is a programmable pipeline; it's not that far apart from thinking of each user coming in and building their own http handler pipeline, making their own bespoke computational flow. I guess my point is that computation itself is configuration. XSLT probably came closest to that sun. But we see similar lessons everywhere we look.
- macintux 7mo agoI work with a project that is heavily configuration-driven. It seems promising, but in reality: - Configuration is massively duplicated, across repositories - No one is willing to rip out redundancy, because comprehensive testing is not practical - In order to understand the configuration, you have to read lots of code, again across multiple repositories (this in particular is a problem for LLM assistance, at least the way we currently use it) I love the idea, but in practice it’s currently a nightmare. I think if we took a week we could clean things up a fair bit, but we don’t have a week (at least as far as management is concerned), and again, without full functional testing, it’s difficult to know when you’ve accidentally broken someone else’s subsystem
- 7mo ago
- anotheryou 7mo agobut doesn't "the code is documentation" work better for machines? and don't we have doc-blocks?
- zdragnar 7mo agoCode doesn't express intent, only the implementation. Docblocks are fine for specifying local behavior, but are terrible for big picture things.
- anotheryou 7mo agoright you are :) does literate code have a place for big pic though?
- palata 7mo agoWell many times it does. bool isEven(number: Int) { return number % 2 == 0 } I would say this expresses the intent, no need for a comment saying "check if the number is even". Most of the code I read (at work) is not documented, still I understand the intent. In open source projects, I used to go read the source code because the documentation is inexistent or out-of-date. To the point where now I actually go directly to the source code, because if the code is well written, I can actually understand it.
- zdragnar 7mo agoIn your example, the implementation matches the intention. That is not the same thing. bool isWeekday(number: Int) { return number % 2 == 0 } With this small change, all we have are questions: Is the name wrong, or the behavior? Is this a copy / paste error? Where is the specification that tells me which is right, the name or the body? Where are the tests located that should verify the expected behavior? Did the implementation initially match the intent, but some business rule changed that necessitated a change to the implantation and the maintainer didn't bother to update the name? Both of our examples are rather trite- I agree that I wouldn't bother documenting the local behavior of an "isEven" function. I probably would want a bit of documentation at the callsite stating why the evenness of a given number is useful to know. Generally speaking, this is why I tend to dislike docblock style comments and prefer bigger picture documentation instead- because it better captures intent.
- librasteve 7mo agoI dont know Org, but Rakudoc https://docs.raku.org/language/pod https://docs.raku.org/language/pod is useful for literate programming (put the docs in the code source) and for LLM (the code is "self documenting" so that in the LLM inversion of control, the LLM can determine how to call the code). https://podlite.org https://podlite.org is this done in a language neutral way perl, JS/TS and raku for now. Heres an example: #!/usr/bin/env raku =begin pod =head1 NAME Stats::Simple - Simple statistical utilities written in Raku =head1 SYNOPSIS use Stats::Simple; my @numbers = 10, 20, 30, 40; say mean(@numbers); # 25 say median(@numbers); # 25 =head1 DESCRIPTION This module provides a few simple statistical helper functions such as mean and median. It is meant as a small example showing how Rakudoc documentation can be embedded directly inside Raku source code. =end pod unit module Stats::Simple; =begin pod =head2 mean mean(@values --> Numeric) Returns the arithmetic mean (average) of a list of numeric values. =head3 Parameters =over 4 =item @values A list of numeric values. =back =head3 Example say mean(1, 2, 3, 4); # 2.5 =end pod sub mean(*@values --> Numeric) is export { die "No values supplied" if @values.elems == 0; @values.sum / @values.elems; } =begin pod =head2 median median(@values --> Numeric) Returns the median value of a list of numbers. If the list length is even, the function returns the mean of the two middle values. =head3 Example say median(1, 5, 3); # 3 say median(1, 2, 3, 4); # 2.5 =end pod sub median(*@values --> Numeric) is export { die "No values supplied" if @values.elems == 0; my @sorted = @values.sort; my $n = @sorted.elems; return @sorted[$n div 2] if $n % 2; (@sorted[$n/2 - 1] + @sorted[$n/2]) / 2; } =begin pod =head1 AUTHOR Example written to demonstrate Rakudoc usage. =head1 LICENSE Public domain / example code. =end pod
- cadamsdotcom 7mo agoTest code and production code in a symmetrical pair has lots of benefits. It’s a bit like double entry accounting - you can view the code’s behavior through a lens of the code itself, or the code that proves it does what it seems to do. You can change the code by changing either tests or production code, and letting the other follow. Code reviews are a breeze because if you’re confused by the production code, the test code often holds an explanation - and vice versa. So just switch from one to the other as needed. Lots of benefits. The downside is how much extra code you end up with of course - up to you if the gains in readability make up for it.
- senderista 7mo agoThe "test runbook" approach that TFA describes sounds like doctest comments in Python or Rust.
- stephbook 7mo agoTake it to the logical conclusion. Track the intended behavior in a proper issue tracking software like Jira. Reference the ticket in your version control system. Boring and reliable, I know. If you need guides to the code base beyond what the programming language provides, just write a directory level readme.md where necessary.
- andyferris 7mo agoI think the externality of issue tracking systems like Jira (or even GitHub) cause friction. Literate programming has everything in one place. I’d like to have a good issue tracking system inside git. I think the SQLite version management system has this functionality but I never used it. One thing to solve is that different kinds of users need to interact with it in different kinds of ways. Non-programmers can use Jira, for example. Issues are often treated as mutable text boxes rather than versioned specification (and git is designed for the latter). It’s tricky!
- jauntywundrkind 7mo agoOne of the things I love most about WebMCP is the idea that it's a MCP session that exists on the page, which the user already knows. Most of these LLM things are kind of separate systems, with their own UI. The idea of agency being inlayed to existing systems the user knows like this, with immediate bidirectional feedback as the user and LLM work the page, is incredibly incredibly compelling to me. Series of submissions (descending in time): https://news.ycombinator.com/item?id=47211249 https://news.ycombinator.com/item?id=47211249 https://news.ycombinator.com/item?id=47037501 https://news.ycombinator.com/item?id=47037501 https://news.ycombinator.com/item?id=45622604 https://news.ycombinator.com/item?id=45622604
- jph00 7mo agoNearly all my coding for the last decade or so has used literate programming. I built nbdev, which has let me write, document, and test my software using notebooks. Over the last couple of years we integrated LLMs with notebooks and nbdev to create Solveit, which everyone at our company uses for nearly all our work (even our lawyers, HR, etc). It turns out literate programming is useful for a lot more than just programming!
- mkl 7mo agoThis seems to be the best link? https://solve.it.com/ https://solve.it.com/ The name is quite hard to search for, as it's used by a lot of different things. Jeremy it's pretty hard to understand what this is from the descriptions, and the two videos are each ~1 hour long. Please consider showing screenshots and one or two short videos.
- moehj 7mo ago[dead]
- amelius 7mo agoWe need an append-only programming language.
- cfiggers 7mo agoInteresting and semi-related idea: use LLMs to flag when comments/docs have come out of sync with the code. The big problem with documentation is that if it was accurate when it was written, it's just a matter of time before it goes stale compared to the code it's documenting. And while compilers can tell you if your types and your implementation have come out of sync, before now there's been nothing automated that can check whether your comments are still telling the truth. Somebody could make a startup out of this.
- spawarotti 7mo agoThere is at least one startup doing it already (I'm not affiliated with it in any way): https://promptless.ai/ https://promptless.ai/
- cfiggers 7mo agoThanks for the pointer. That looks more to me like it's totally synthesizing the docs for me. I can see someone somewhere wanting that. I would want a UX more like a compiler warning. "Comment on line 447 may no longer be accurate." And then I go fix it my own dang self.
- gogopromptless 7mo agoHa, this is funny (also sad for me because I failed to explain on website clearly) because you have described exactly what it does as an example of what it can't do. The core loop is more like a truffle-hunting pig than a ghostwriter. Promptless watches for signal that your product is behaving differently from the live documentation. It watches PRs opened/merging, Slack threads, support tickets. Then like a pig alerting on a truffle it shows up like "hey, this section over here doesn't match what the code/product does anymore." Now of course we'll also generate a first draft of a suggested fix, but I want to say 40% of tech writers just like knowing when things changed. Its a proper union find algorithm, where every suggestion links back to the source that triggered it, but multiple source do get linked up to just a single canonical suggestion. So you don't get duplicate alerts if people keep talking for weeks about a fix going out in the next release. Obviously I've got some more work to do on the website again but c'est la vie.
- deleted 7mo ago[deleted]
- charcircuit 7mo ago>I don't have data to support this With there being data that shows context files which explain code reduces the performance of them, it is not straightforward that literate programming is better so without data this article is useless.
- trane_project 7mo agoI think full literate programming is overkill but I've been doing a lighter version of this: - Module level comments with explanations of the purpose of the module and how it fits into the whole codebase. - Document all methods, constants, and variables, public and private. A single terse sentence is enough, no need to go crazy. - Document each block of code. Again, a single sentence is enough. The goal is to be able to know what that block does in plain English without having to "read" code. Reading code is a misnomer because it is a different ability from reading human language. Example from one of my open-source projects: https://github.com/trane-project/trane/blob/master/src/scheduler.rs https://github.com/trane-project/trane/blob/master/src/sched...
- avatardeejay 7mo agoSomething in this realm covers my practice. I just keep a master prompt for the whole program, and sparsely documented code. When it's time to use LLM's in the dev process, they always get a copy of both and it makes the whole process like 10x as coherent and continuous. Obvi when a change is made that deviates or greatly expands on the spec, I update the spec.
- grapheneposter 7mo agoI do something similar with quality gates. I have a bunch of markdown files at the ready to point agents to for various purposes. It lets me leverage LLMs at any stage of the dev process and my clients get docs in their format without much maintenance from myself. As you said once you get it down it becomes a very coherent process that can be iterated on in its own right. I am currently fighting the recursive improvement loop part of working with agents.
- rednafi 7mo agoI think a lighter version of literate programming, coupled with languages that have a small API surface but are heavy on convention, is going to thrive in this age of agentic programming. A lighter API footprint probably also means a higher amount of boilerplate code, but these models love cranking out boilerplate. I’ve been doing a lot more Go instead of dynamic languages like Python or TypeScript these days. Mostly because if agents are writing the program, they might as well write it in a language that’s fast enough. Fast compilation means agents can quickly iterate on a design, execute it, and loop back. The Go ecosystem is heavy on style guides, design patterns, and canonical ways of doing things. Mostly because the language doesn’t prevent obvious footguns like nil pointer errors, subtle race conditions in concurrent code, or context cancellation issues. So people rely heavily on patterns, and agents are quite good at picking those up. My version of literate programming is ensuring that each package has enough top-level docs and that all public APIs have good docstrings. I also point agents to read the Google Go style guide [1] each time before working on my codebase.This yields surprisingly good results most of the time. [1] https://google.github.io/styleguide/go/ https://google.github.io/styleguide/go/
- username223 7mo ago> The Go ecosystem is heavy on style guides, design patterns, and canonical ways of doing things. Go was designed based on Rob Pike's contempt for his coworkers (https://news.ycombinator.com/item?id=16143918 https://news.ycombinator.com/item?id=16143918), so it seems suitable for LLMs.
- Arubis 7mo agoAnecdotally, Claude Opus is at least okay at literate emacs. Sometimes takes a few rounds to fix its own syntax errors, but it gets the idea. Requiring it to TDD its way in with Buttercup helps.
- aplomb1026 7mo ago[flagged]
- akater 7mo agoThe question posed is, “With agents, does it become practical to have large codebases that can be read like a narrative, whose prose is kept in sync with changes to the code by tireless machines?” It's not practical to have codebases that can be read like a narrative, because that's not how we want to read them when we deal with the source code. We jump to definitions, arriving at different pieces of code in different paths, for different reasons, and presuming there is one universal, linear, book-style way to read that code, is frankly just absurd from this perspective. A programming language should be expressive enough to make code read easily, and tools should make it easy to navigate. I believe my opinion on this matters more than an opinion of an average admirer of LP. By their own admission, they still mostly write code in boring plain text files. I write programs in org-mode all the time. Literally (no pun intended) all my libraries, outside of those written for a day job, are written in Org. I think it's important to note that they are all Lisp libraries, as my workflow might not be as great for something like C. The documentation in my Org files is mostly reduced to examples — I do like docstrings but I appreciate an exhaustive (or at least a rich enough) set of examples more, and writing them is much easier: I write them naturally as tests while I'm implementing a function. The examples are writen in Org blocks, and when I install a library of push an important commit, I run all tests, of which examples are but special cases. The effect is, this part of the documentation is always in sync with the code (of course, some tests fail, and they are marked as such when tests run). I know how to sync this with docstrings, too, if necessary; I haven't: it takes time to implement and I'm not sure the benefit will be that great. My (limited, so far) experience with LLMs in this setting is nice: a set of pre-written examples provides a good entry point, and an LLM is often capable of producing a very satisfactory solution, immediately testable, of course. The general structure of my Org files with code is also quite strict. I don't call this “literate programming”, however — I think LP is a mess of mostly wrong ideas — my approach is just a “notebook interface” to a program, inspired by Mathematica Notebooks, popularly (but not in a representative way) imitated by the now-famous Jupyter notebooks. The terminology doesn't matter much: what I'm describing is what the silly.business blogpost is largerly about. The author of nbdev is in the comments here; we're basically implementing the same idea. silly.business mentions tangling which is a fundamental concept in LP and is a good example of what I dislike about LP: tangling, like several concepts behing LP, is only a thing due to limitations of the programming systems that Donald Knuth was using. When I write Common Lisp in Org, I do not need to tangle, because Common Lisp does not have many of the limitations that apparently influenced the concepts of LP. Much like “reading like a narrative” idea is misguided, for reasons I outlined in the beginning. Lisp is expressive enough to read like prose (or like anything else) to as large a degree as required, and, more generally, to have code organized as non-linearly as required. This argument, however, is irrelevant if we want LLMs, rather than us, read codebases like a book; but that's a different topic.
- pjmlp 7mo agoI rather go with formal specifications, and proofs.
- deleted 7mo ago[deleted]
- palata 7mo agoI am not convinced. - Natural languages are ambiguous. That's the reason why we created programming languages. So the documentation around the code is generally ambiguous as well. Worse: it's not being executed, so it can get out of date (sometimes in subtle ways). - LLMs are trained on tons of source code, which is arguably a smaller space than natural languages. My experience is that LLMs are really good at e.g. translating code between two programming languages. But translating my prompts to code is not working as well, because my prompts are in natural languages, and hence ambiguous. - I wonder if it is a question of "natural languages vs programming languages" or "bad code vs good code". I could totally imagine that documenting bad code helps the LLMs (and the humans) understand the intent, while documenting good code actually adds ambiguity. What I learned is that we write code for humans to read. Good code is code that clearly expresses the intent. If there is a need to comment the code all over the place, to me it means that the code is maybe not as good as it should be :-). Of course there is an argument to make that the quality of code is generally getting worse every year, and therefore there is more and more a need for documentation around it because it's getting hard to understand what the hell the author wanted to do.
- hosh 7mo agoI don’t have my LLMs generate literate programming. I do ask it to talk about tradeoffs. I have full examples of something that is heavily commented and explained, including links to any schemas or docs. I have gotten good results when I ask an LLM to use that as a template, that not everything in there needs to be used, and it cuts down on hallucinations by quite a bit.
- bottd 7mo ago> If there is a need to comment the code all over the place, to me it means that the code is maybe not as good as it should be :-) If good code was enough on its own we would read the source instead of documentation. I believe part of good software is good documentation. The prose of literate source is aimed at documentation, not line-level comments about implementation.
- WillAdams 7mo ago
- nailer 7mo ago> Literate programming is the idea that code should be intermingled with prose such that an uninformed reader could read a code base as a narrative Have you tried naming things properly? A reader that knows your language could then read your code base as a narrative.
- ajkjk 7mo agoI've had the same thought, maybe more grandiosely. The idea is that LLM prompts are code -- after all they are text that gets 'compiled' (by the LLM) into a lower-level language (the actual code). The compile process is more involved because it might involve some back-and-forth, but on the other hand it is much higher level. The goal is to have a web of prompts become the source of truth for the software: sort of like the flowchart that describes the codebase 'is' the codebase.
- Copyrightest 7mo agoOne problem with this is that there isn't really a "current prompt" that completely describes the current source code; each source file is accompanied by a full chat log, including false starts and misunderstandings. It's sort of like reading a git history instead of the actual file.
- ajkjk 7mo agotrue, but that just means that's the problem to solve. probably the ideal architecture isn't possible right now. But I sorta imagine that you could later on take the full transcript of that conversation and expect any LLM to implement more or less the same thing based on it, so that eventually it becomes a full 'spec'. And maybe there is a way to trim the parts out of it that are not needed... like to automatically produce an initial prompt which is equivalent to the results of a longer session, but is precise enough so as to not need clarification upon reprocessing it. Something like that? I'm not sure if that's something that already exists.
- sarchertech 7mo ago> But I sorta imagine that you could later on take the full transcript of that conversation and expect any LLM to implement more or less the same thing based on it Why would you think this though? There are an infinite number of programs that can satisfy any non-trivial spec. We have theoretical solutions to LLM non-determinism, we have no theoretical solutions to prompt instability especially when we can’t even measure what correct is.
- koolala 7mo agoLeft to right APL style code seems like it could be words instead of symbols.
- whatgoodisaroad 7mo agoit could be fun to make a toy compiler that takes an arbitrary literate prompt as input and uses an LLM to output a machine code executable (no intermediate structured language). could call it llmllvm. perhaps it would be tremendously dangerous
- wewewedxfgdf 7mo agoWhat we need is comments that LLMs simply do not delete. We need metadata in source code that LLMs don't delete and interpreters/compilers/linters don't barf on.
- rudhdb773b 7mo agoI'd love to see what Tim Daly could with LLMs on Axiom's code base.
- ljlolel 7mo agoEveryone is circling getting rid of the code and just having Englishscript https://jperla.com/blog/claude-electron-not-claudevm https://jperla.com/blog/claude-electron-not-claudevm
- arikrahman 7mo agoI have instructed my LLMs to at least provide a comment per function, but prompt it to comment when it takes out things additionally, and why it opted to choose a particular solution. DistroTube also loves declarative literate programming approach, often citing how his one document configuration with nix configures his whole system.
- hsaliak 7mo agoI explored this in std::slop (my clanker) https://github.com/hsaliak/std_slop https://github.com/hsaliak/std_slop. One of it's differentiating features of this clanker i that it only has a single tool call, run_js. The LLM produces js scripts to do it's work. Naturally, i tried to teach it to add comments for these scripts and incorporate literate programming elements. This was interesting because, every tool call now 'hydrated' some free form thinking, but it comes at output token cost. Output Tokens are expensive! In GPT-5.4 it's ~180 dollars per Million tokens! I've settled for brief descriptions that communicate 'why' as a result. The code is documentation after all.
- prpl 7mo agoIt has always been been possible to program literately in programming languages - not to the extent that you can in Web, but good code can read like a story and obviate comments
- cmontella 7mo agoI agree with this. I've been a fan of literate programming for a long time, I just think it is a really nice mode of development, but since its inception it hasn't lived up to its promise because the tooling around the concept is lacking. Two of the biggest issues have been 1) having to learn a whole new toolchain outside of the compiler to generate the documents 2) the prose and code can "drift" meaning as the codebase evolves, what's described by the code isn't expressed by the prose and vice versa. Better languages and tooling design can solve the first problem, but I think AI potentially solves the second. Here's the current version of my literate programming ideas, Mechdown: https://mech-lang.org/post/2025-11-12-mechdown/ https://mech-lang.org/post/2025-11-12-mechdown/ It's a literate coding tool that is co-designed with the host language Mech, so the prose can co-exist in the program AST. The plan is to make the whole document queryable and available at runtime. As a live coding environment, you would co-write the program with AI, and it would have access to your whole document tree, as well as live type information and values (even intermediate ones) for your whole program. This rich context should help it make better decisions about the code it writes, hopefully leading to better synthesized program. You could send the AI a prompt, then it could generate the code using live type information; execute it live within the context of your program in a safe environment to make sure it type checks, runs, and produces the expected values; and then you can integrate it into your codebase with a reference to the AI conversation that generated it, which itself is a valid Mechdown document. That's the current work anyway -- the basis of this is the literate programming environment, which is already done. The docs show off some more examples of the code, which I anticipate will be mostly written by AIs in the future: https://docs.mech-lang.org/getting-started/introduction.html https://docs.mech-lang.org/getting-started/introduction.html
- catlifeonmars 7mo agoWe actually have had literate programming for a while, it just doesn’t look exactly how it was envisioned: Nowadays, it’s common for many libraries to have extensive documentation, including documentation, hyperlinks and testable examples directly inline in the form of comments. There’s usually a well defined convention for these comments to be converted into HTML and some of them link directly back to the relevant source code. This isn’t to say they’re exactly what is meant by literate programming, but I gotta say we’re pretty damn close. Probably not much more than a pull request away for your preferred languages’ blessed documentation generator in fact. (The two examples I’m using to draw my conclusions are Rust and Go).
- openclaw01 7mo ago[dead]
- teleforce 7mo agoNot sure if the author know about CUE, here's the HN post from early this year on literate programming with CUE [1]. CUE is based of value-latticed logic that's LLM's NLP cousin but deterministic rather than stochastic [2]. LLMs are notoriously prone to generating syntactically valid but semantically broken configurations thus it should be used with CUE for improving literate programming for configs and guardrailing [3]. [1] CUE Does It All, But Can It Literate? (22 comments) https://news.ycombinator.com/item?id=46588607 https://news.ycombinator.com/item?id=46588607 [2] The Logic of CUE: https://cuelang.org/docs/concept/the-logic-of-cue/ https://cuelang.org/docs/concept/the-logic-of-cue/ [3] Guardrailing Intuition: Towards Reliable AI: https://cue.dev/blog/guardrailing-intuition-towards-reliable-ai/ https://cue.dev/blog/guardrailing-intuition-towards-reliable...
- jimbokun 7mo agoThis does seem exciting at first glance. Just write the narrative part of literate programming and an LLM generates the code, then keep the narrative and voila! Literate programming without the work of generating both. However I see two major issues: Narrative is meant to be consumed linearly. But code is consumed as a graph. We navigate from a symbol to its definition, or from definition to its uses, jumping from place to place in the code to understand it better. The narrative part of linear programming really only works for notebooks where the story being told is dominant and the code serves the story. Second is that when I use an LLM to write code, the changes I describe usually require modifying several files at once. Where does this “narrative” go relative to the code. And yes, these two issues are closely related.
- Agent_Builder 7mo ago[dead]
- yuppiemephisto 7mo agoI do a form of literate programming for code review to help read AI code. I use [Lean 4](lean-lang.org) and its doc tool [Verso](https://github.com/leanprover/verso/ https://github.com/leanprover/verso/) and have it explain the code through a literate essay. It is integrated with Lean and gets proper typechecking etc which I find helpful.
- JEONSEWON 7mo ago[flagged]
- eisbaw 7mo agohttps://github.com/eisbaw/litterate_bitorrent https://github.com/eisbaw/litterate_bitorrent 800 pages, noweb extracts rust. Made by claude in a ralph loop over 1-2 days. yes, it downloads actual torrents.
- DennisL123 7mo agoIf agents can already read and rewrite code, literate programming might actually be unnecessary. Instead of maintaining prose alongside code, you could let agents generate explanations on demand. The real requirement becomes writing code in a form that is easily interpretable and transformable by the next agent in the chain. In that model, code itself becomes the stable interface, while prose is just an ephemeral view generated whenever a human (or another agent) needs it.
- tacone 7mo agoI am already doing that. For performance, I am just caching the latest explanation alongside the code.
- jasfi 7mo agoI wrote something similar where you specify the intent in Markdown at the file level. That can also be done by an AI agent. Each intent file compiles to a source file. It works, but needs improvement. Any feedback is welcome! https://intentcode.dev https://intentcode.dev https://github.com/jfilby/intentcode https://github.com/jfilby/intentcode
- beernet 7mo agoLiterate programming sounds great in a blog post, but it falls apart the moment an agent starts hallucinating between the prose and the actual implementation. We’re already struggling with docstrings getting out of sync; adding a layer of philosophical "intent" just gives the agent more room to confidently output garbage. If you need a wall of text to make an agent understand your repo, your abstractions are probably just bad. It feels like we're trying to fix a lack of structural clarity with more tokens.
- rorylaitila 7mo agoEven on the latest models, LLMs are not deterministic between "don't do this thing" and "do this thing". They are both related to "this thing" and depending on other content in the context and seed, may randomly do the thing or not. So to get the best results, I want my context to be the smallest possible truthful input, not the most elaborated. More is not better. I think good names on executable source code and tightest possible documentation is best for LLMs, and probably for people too.
- octoclaw 7mo ago[dead]
- s3anw3 7mo agoI think the tension between natural language and code is fundamentally about information compression. Code is maximally compressed intent — minimal redundancy, precise semantics. Prose is deliberately less compressed — redundant, contextual, forgiving — because human cognition benefits from that slack. Literate programming asks you to maintain both compression levels in parallel, which has always been the problem: it's real work to keep a compressed and an uncompressed representation in sync, with no compiler to enforce consistency between them. What's interesting about your observation is that LLMs are essentially compression/decompression engines. They're great at expanding code into prose (explaining) and condensing prose into code (implementing). The "fundamental extra labor" you describe — translating between these two levels — is exactly what they're best at. So I agree with your conclusion: the economics have changed. The cost of maintaining both representations just dropped to near zero. Whether that makes literate programming practical at scale is still an open question, but the bottleneck was always cost, not value.
- fhub 7mo agoWe were taught Literate Programming and xtUML at university. In both courses, the lecturers (independently) tried to convince us that these technologies were the future. I also did an AI/ML course. That lecturer lamented that the golden era was in the past.
- threethirtytwo 7mo agoShould be extremely low effort to try this out with an agent. The thing is, I feel an agent can read code as if it was english. It doesn't differentiate one as hard and the other as much more readable as we do. So it could end up just increasing the token burn amount just to get through a program because it has to run through the literate part as well as the actual code part.
- melisgl 7mo ago- For many languages, we can get away with Untangled LP. See, e.g. https://quotenil.com/untangling-literate-programming.html https://quotenil.com/untangling-literate-programming.html - Introducing redundancies (of code, tests, documentation) is our primary tool to increase our confidence in the correctness of the solution: See, e.g. https://quotenil.com/multifaceted-development.html https://quotenil.com/multifaceted-development.html - Untangled LP has been a good idea even before LLMs. It's even better now, as LLMs can maintain documentation and check it against the code.
- jarnm0 7mo agoI agree that we should revisit literate programming, but I don't think using LLMs to generate or summarize code is ever going to be the ultimate solution. You want something that is unambiguous and computable but that also non-technical people can work with - a programming language which reads like natural language. In 2021 I started to "solve programming in natural language" by building a platform which enables creating these kinds of domain-specific (projectional) programming languages which can look exactly like (structured) natural language. The idea was to enable domain/business experts to manage the business rules in different kinds of systems. The platform works and the use-cases are there, but I haven't been able to commercialize it yet. I didn't initially build it for LLMs, but after the release of GPT 3.5 it became obvious that these structured natural languages would be the perfect link between non-technical people, LLMs and deterministic logic. So now I have enabled the platform to instruct LLMs to work with the languages with very good results and are trying to commercialize for LLM use-cases. There absolutely is synenergies in combining literate programming and LLMs! I've written a bit more about it here: - https://www.linkedin.com/pulse/how-i-accidentally-built-context-hypergraph-platform-jarno-montonen-rnxqf/ https://www.linkedin.com/pulse/how-i-accidentally-built-cont... - https://www.linkedin.com/pulse/llms-structured-natural-languages-jarno-montonen-0hv3f https://www.linkedin.com/pulse/llms-structured-natural-langu... (P.S. Looking for a co-founder, feel free to reach out in LinkedIn if this resonates!)
- vicchenai 7mo ago[dead]
- ChicagoDave 7mo agoI think we’re on the verge of readable code and human-edited code disappearing. There is a paradigm shift coming. Ephemeral code.
- monsieurbanana 7mo agoThose two are not linked. I could buy that maybe human-readable code will be the minority. But what does ephemeral code even means? That we will throw everything out of the window at every release cycle and recreate from scratch with llms based on specs? That's not happening
- gombosg 7mo agoI think you're right, ephemeral code would be the concept that you have (I'm hand-waving) "the spec", that specifies what the code should be doing and the AI could regenerate the code any time based on it. I'm also baffled by this concept and fundamentally believe that code _should be_ the ground truth (the spec), hence it should be human readable. That's what "clean code" would be about, choosing tools and abstractions so that code is consumable for humans and easy to reason about, debug and extend. If we let go of that and rely on LLMs entirely... not sure where that would land, since computers ultimately execute the code - and the company is liable for the results of that code being executed -, not the plain language "specs".
- ChicagoDave 7mo agoBy ephemeral I mean we no longer care about code as an asset. If a feature is broken or requires changes, we can perform a clean organ transplant. The actual code doesn’t matter anymore. Its testable functionality is what matters.
- xmcqdpt2 7mo agoI'd much much rather the model write the code blocks than the prose myself. In my experience LLM can produce pretty decent code, but the writing is horrible. If anything I would prefer an agentic tool where you don't even see the slop. I definitely would rather it not be committed.
- macey 7mo agoI agree it's worth revisiting. Actually I wrote about this recently, I didn't realise there was a precedent here. https://tessl.io/blog/how-to-capture-intent-with-coding-agents/ https://tessl.io/blog/how-to-capture-intent-with-coding-agen... > As a benefit, the code base can now be exported into many formats for comfortable reading. This is especially important if the primary role of engineers is shifting from writing to reading. Underrated point. Also, whether we like it or not, people without engineering backgrounds will be closer to code in the future. That trend isn't slowing down. The inclusion of natural language will make it easier for them to be productive and learn.
- CloakHQ 7mo ago[dead]
- mattrathbun 7mo ago[dead]
- oliver_dr 7mo ago[dead]
- gwbas1c 7mo agoOne thing I've discovered with an LLM is that I can ask it to search through my codebase and explain things to me. It saves a lot of time when I need to understand concepts that would otherwise require a few hours of reading and digging.
- CharlieDigital 7mo agoThe easiest thing to do is to have the LLM leave its own comments. This has several benefits because the LLM is going to encounter its own comments when it passes this code again. > - Apply comments to code in all code paths and use idiomatic C# XML comments > - <summary> be brief, concise, to the point > - <remarks> add details and explain "why"; document reasoning and chain of thought, related files, business context, key decisions. > - <params> constraints and additional notes on usage > - inline comments in code sparingly where it helps clarify behavior (I have something similar for JSDoc for JS and TS) Several things I've observed: 1. The LLM is very good at then updating these comments when it passes it again in the future. 2. Because the LLM is updating this, I can deduce by proxy that it is therefore reading this. It becomes a "free" way to embed the past reasoning into the code. Now when it reads it again, it picks up the original chain-of-thought and basically gets "long term memory" that is just-in-time and in-context with the code it is working on. Whatever original constraints were in the plan or the prompt -- which may be long gone or otherwise out of date -- are now there next to the actual call site. 3. When I'm reviewing the PR, I can now see what the LLM is "thinking" and understand its reasoning to see if it aligns with what I wanted from this code path. If it interprets something incorrectly, it shows up in the `<remarks>`. Through the LLM's own changes to the comments, I can see in future passes if it correctly understood the objective of the change or if it made incorrect assumptions.
- solarkraft 7mo agoHow do you deal with the comments sometimes being relatively noisy for humans? I tend to be annoyed by comments overly referring to a past correction prompt and not really making sense by themselves, but then again this IS probably the highest value information because these are exactly the things the LLM will stumble on again.
- CharlieDigital 7mo ago> How do you deal with the comments sometimes being relatively noisy for humans? To extents, that is a function of tweaking the prompt to get the level of detail desired and signal/vs noise produced by the LLM. e.g. constraining the word count it can use for comments. We have a small team of approvers that are reviewing every PR and for us, not being able to see the original prompt and flow of interactions with the agent, this approach lets us kind of see that by proxy when reviewing the PR so it is immensely useful. Even for things like enum values, for example. Why is this enum here? What is its use case? Is it needed? Having the reasoning dumped out allows us to understand what the LLM is "thinking". (Of course, the biggest benefit is still that the LLM sees the reasoning from an earlier session again when reading the code weeks or months later).
- trixn86 7mo agoI don't think that agents actually benefit from comments that describe what the code does at all. In my experience in the best case they don't really improve response quality and in the worst case they drastically reduce it. This is just noise that does not help the AI understand the context any better. This has already been true for a trained developer and it is even more so true for AI agents. Natural language is in almost every way less efficient in providing context and AI has no problem at all to infer intent from good code. The challenge is rather to make the AI produce good code which needs a strict harness and rules. Another good addition is semantic indexing of the codebase to help the AI find code using semantic search (which is what some agents already do quite successfully). The only context I consistently found to be useful is about project-specific tool calling. Trying to provide natural language context about the project itself always proved to be ambiguous, inaccurate and out-of-date. Agents are very good at reading code and code is the best way to express context unambiguously.
- empath75 7mo agoYou can have perfectly good code, which is perfectly easy to understand which nevertheless _does not do what you intended to do_. That is why tests exist, after all.
- frakt0x90 7mo agoMaybe for literate programming, we can switch from common, ambiguous human languages like English and Spanish to [Lojban](https://en.wikipedia.org/wiki/Lojban https://en.wikipedia.org/wiki/Lojban)! That way our human language will be unambiguous which will translate to machine code much better. We'll call this the de facto "language for programming". Improvements and other variants may pop up in the future as new needs arise. All that is old is new again.
- ontouchstart 7mo agoLiterate programming in the sense of Donald Knuth is more about the chain of thoughts of the programmer than documenting code with comments or doc strings.
- robertwer 7mo agoThere seems to be some evidence that literate programming style comments help humans to comprehend code they don't know. I found a paper investigating this. Some folks from Google tested 1) how good LLMs can update existing code with LP style comments and 2) if that helps humans to better understand that enhanced code. (see the 2024 arXiv paper "Natural Language Outlines for Code"). If I remember correctly they had systematically tested how good humans could understand the enhanced code compared to no comments at all and they also tested different flavours of comments (line level, block level etc.). The conclusion was: if you use the right amount of comments in the right style (intent explaining the purpose of the code on block level, not every line), it's very beneficial.