12 ms·
Interesting read, and some interesting ideas, but there's a problem with statements like these: > Sean proposes that in the AI future, the specs will become th
by maltalex 1y ago
Interesting read, and some interesting ideas, but there's a problem with statements like these:
> Sean proposes that in the AI future, the specs will become the real code. That in two years, you'll be opening python files in your IDE with about the same frequency that, today, you might open up a hex editor to read assembly.
> It was uncomfortable at first. I had to learn to let go of reading every line of PR code. I still read the tests pretty carefully, but the specs became our source of truth for what was being built and why.
This doesn't make sense as long as LLMs are non-deterministic. The prompt could be perfect, but there's no way to guarantee that the LLM will turn it into a reasonable implementation.
With compilers, I don't need to crack open a hex editor on every build to check the assembly. The compiler is deterministic and well-understood, not to mention well-tested. Even if there's a bug in it, the bug will be deterministic and debuggable. LLMs are neither.
- beefnugs 1y agosounds like a good nudge to make tests better
- oblio 1y agoIf the tests are written by the AI, who watches the watchers? :-)
- ozim 1y agoThe fun part is that specs already are non-deterministic. If you spend time to write out requirements in English in a way that cannot be misinterpreted in any way you end up with programming language.
- diavolodeejay 1y agoSo… COBOL?
- jlawson 1y agoSpecs are ambiguous but not necessarily non-deterministic. The same entity interpreting the spec in exactly the same way will resolve the ambiguities the same way each time. Human and current AI interpretation of specs is non-deterministic process. But, if we wanted to build a deterministic AI we could.
- oblio 1y ago> But, if we wanted to build a deterministic AI we could. Is this bold proposal backed by any theory?
- reverius42 1y agoGiven that they all use pseudo-random (and not actually random) numbers, they are "deterministic" in the sense that given a fixed seed, they will produce a fixed result... But perhaps that's not what was meant by deterministic. Something like an understandable process producing an answer rather than a pile of linear algebra?
- catoc 1y agoI was thinking the exact same thing: if you don’t change the weights, use identical “temperature” etc, the same prompt will yield the same output. Under the hood it’s still deterministic code running on a deterministic machine
- block_dagger 1y agoThis is incorrect. Temperature would need to be zero to get same result.
- jimbo808 1y agoHumans don't make mistakes nearly as much, the mistakes they do make are way more predictable (they're easier to spot in code review), and they don't tend to make the kinds of catastrophic mistakes that could sink a business. They also tend to cause codebases to rapidly deteriorate, since even very disciplined reviewers can miss the kinds of strange and unpredictable stuff an LLM will do. Redundant code isn't evident in a diff, and things like tautological tests, or useless tests where they're mocking everything and only actually testing the mocks. Or they'll write a bunch of redundant code because they really just aggressively avoid code re-use unless you are very specific. The real problem is just that they don't have brains, and can't think. They generate text that is optimized to look the most right, but not to be the most right. That means they're deceptive right off the bat. When a human is wrong, it usually looks wrong. When an LLM is wrong, it's generating the most correct looking thing it possibly could while still being wrong, with no consideration for actual correctness. It has no idea what "correctness" even means, or any ideas at all, because it's a computer doing matmul. They are text summarization/regurgitation, pattern matching machines. They regurgitate summaries of things seen in their training data, and that training data was written by humans who can think. We just let ourselves get duped into believing the machine is the where the thinking is coming from and not the (likely uncompensated) author(s) whose work was regurgitated for you.
- ACCount37 1y ago>The real problem is just that they don't have brains, and can't think. That would have had more weight if you haven't just described junior developer behavior beforehand. "LLMs can't think" is anthropocentric cope. It's the old AI effect all over again - people would rather die than admit that there's very little practical difference between their own "thinking" and that of an AI chatbot.
- gnulinux996 1y ago> That would have had more weight if you haven't just described junior developer behavior beforehand. Effectively telling that junior developers "don't have brains" is in very bad taste and offensively wrong. > people would rather die than admit that there's very little practical difference between their own "thinking" and that of an AI chatbot. Would you like to elaborate on this? I was told that McDonalds employees would have been replaced by now, self-driving cars will be driving the streets and new medicines would have been discovered. It's been a couple of years that "AI" is out, and no singularity yet.
- qcnguy 1y agoNot really, code even in high level languages is always lower level than English just for computer nonsense reasons. Example: "read a CSV file and add a column containing the multiple of the price and quantity columns". That's about 20 words. Show me the programming language that can express that entire feature in 20 words. Even very English-like languages like Python or Kotlin might just about do it, if you're working in something else like C++ then no. In practice, this spec will expand to changes to your dependency lists (and therefore you must know what library is used for CSV parsing in your language, the AI knows this stuff better than you), then there's some file handling, error handling if the file doesn't exist, maybe some UI like flags or other configuration, working out what the column names are, writing the loop, saving it back out, writing unit tests. Any reasonable programmer will produce a very similar PR given this spec but the diff will be much larger than the spec.
- cess11 1y agoIverson languages could do that quite succinctly.
- sarchertech 1y ago>"read a CSV file and add a column containing the multiple of the price and quantity columns" This is an underspecification if you want to reliably repeatably produce similar code. The biggest difference is that some developers will read the whole CSV into memory before doing the computations. In practice the difference between those implementation is huge. Another big difference is how you represent the price field. If you parse them as floats and the quantity is big enough, you'll end up with errors. Even if quantity is small, you'll have to deal with rounding in your new column. You didn't even specific the name of the new column, so the name is going to be different every time you run the LLM. What happens if you run this on a file the program has already been ran on? And these are just a few of the reasonable ways of fitting that spec but producing wildly different programs. Making a spec that has a good chance of producing a reasonably similar program each time looks more like: “Read input.csv (UTF-8, comma-delimited, header row). Read it line by line, do not load the entire file into memory. Parse the price and quantity columns as numbers, stripping currency symbols and thousands separators; interpret decimals using a dot (.). Treat blanks as null and leave the result null for such rows. Compute per-row line_total = round(Decimal(price) * Decimal(quantity), 2). Append line_total as the last column (name the column "Total") without reordering existing columns, and write to output.csv, preserving quoting and delimiter. Do not overwrite existing columns. Do not evaluate or emit spreadsheet formulas.” And even then you couldn't just check this in and expect the same code to be generated each time, you'd need a large test suite--just to constraint the LLM. And even then the LLM would still occasionally find ways to generate code that passes the tests but does thing you don't want it to.
- aabhay 1y agoNot only that but they’re lossy. A hex representation is strictly more information as long as comments are included or generated.
- physicsguy 1y ago> With compilers, I don't need to crack open a hex editor on every build to check the assembly. The tooling is better than just cracking open the assembly but in some areas people do effectively do this, usually to check for vectorization of hot loops, since various things can mean a compiler fails to do it. I used to use Intel VTune to do this in the HPC scientific world.
- Ozzie_osman 1y ago> This doesn't make sense as long as LLMs are non-deterministic. I think we will find ways around this. Because humans are also non-deterministic. So what do we do? We review our code, test it, etc. LLMs could do a lot more of that. Eg, they could maintain and run extensive testing, among other ways to validate that behavior matches the spec.
- maltalex 1y agoIf you're reviewing the code, then you're no longer "opening python files with the same frequency that you open up a hex editor to read assembly".
- lysecret 1y agoI agree this whole spec based approach is misguided. Code is the spec.
- fragmede 1y ago> there's no way to guarantee that the LLM will turn it into a reasonable implementation. There's also no way to guarantee that you're not going to get hit by a meteor strike tomorrow. It doesn't have to be provably deterministic at a computer science PhD level for people without PhDs to say eh, it's fine. Okay, it's not deterministic. What does that mean in practice? Given the same spec.md file, at the layer of abstraction where we're no longer writing code by hand, who cares, because of a lack of determinism, if the variable for the filename object is called filename or fname or file or name as long as the code is doing something reasonable? If it works, if it passes tests, if we presume that the stoichastic parrot is going to parrot out its training data sufficiently close each time, why is it important? As far as compilers being deterministic, there's a fascinating detail we ran into with Ksplice. They're not. They're only sufficiently enough that we trust them to be fine. There was this bug we kept tripping, back in roughly 2006, where GCC would swap registers used for a variable, resulting in the Ksplice patch being larger than it had to be, to include handling the register swap as well. The bug has since been fixed, exposing the details of why it was choosing different registers, but unfortunately I don't remember enough details about it. So don't believe me if you don't want to, but the point is, we trust the c compiler, given a function that takes in variables a, b, c, d, that a, b, c, and d will be map them to r0, r1, r2, or r3. We don't actually care what the order that mapping goes, so long as it works. So the leap, that some have made, and others have not, is that LLMs aren't going to randomly flip out and delete all your data. Which is funny, because that's actually happened on replit. Despite that, despite the fact that LLMs still hallucinate total bullshit and goes off the rail; some people trust LLMs enough to convert a spec to working code. Personally, I think we're not there yet and won't be while GPU time isn't free. (Arguably it is already because anybody can just start typing into chat.com, but that's propped up by VC funding. That isn't infinite, so we'll have to see where we're at in a couple of years.) That addresses the determinism part. The other part that was raised is debuggable. Again, I don't think we're at a place where we can get rid of generated code any time soon, and as long as code is being generated, then we can debug it using traditional techniques. As far as debugging LLMs themselves, it's not zero. They're not mainstream yet, but it's an active area of research. We can abliterate models and fine tune them (or whatever) to answer "how do you make cocaine", counter to their training. So they're not total black boxes. Thus, even if traditional software development dies off, the new field is LLM creation and editing. As with new technologies, porn picks it up first. Llama and other downlodable models (they're not open source https://www.downloadableisnotopensource.org/ https://www.downloadableisnotopensource.org/ ). Downloadable models have been fine tuned or whatever to generate adult content, despite being trained not to. So that's new jobs being created in a new field.
- deleted 1y ago[deleted]
- musebox35 1y ago"The prompt could be perfect, but there's no way to guarantee that the LLM will turn it into a reasonable implementation." I think it is worse than that. The prompt, written in natural language, is by its very nature vague and incomplete, which is great if you are aiming for creative artistry. I am also really happy that we are able to search for dates using phrases like "get me something close to a weekend, but not on Tuesdays" on a booking website instead of picking dates from a dropdown box. However, if natural language was the right tool for software requirements, software engineering would have been a solved problem long ago. We got rightfully excited with LLMs, but now we are trying to solve every problem with it. IMO, for requirements specification, the situation is similar to earlier efforts using formal systems and full verification, but at the exact opposite end. Similar to formal software verification, I expect this phase to end up as a partially failed experiment that will teach us new ways to think about software development. It will create real value in some domains and it will be totally abandoned in others. Interesting times...
- weego 1y agoWe also have to pretend that anyone has ever been any good at writing descriptive, detailed, clear and precise specs or documentation. That might be a skillset that appears in the workforce, but absolutely not in 2 years. A technical writer that deeply understands software engineering so they can prompt correctly but is happy not actually looking at code and just goes along with whatever the agent generates? I don't buy it. This seems like a typical engineer forgets people aren't machines line of thinking.
- FitchApps 1y agoThis. Even with Junior Devs, implementation is always more or less deterministic (based on ones abilities/skills/aptitude). With AI models, you get totally different implementations even when specifically given clear directions via prompt.
- dcre 1y ago“This doesn't make sense as long as LLMs are non-deterministic.” I think this is a logical error. Non-determinism is orthogonal to probability of being correct. LLMs can remain non-deterministic while being made more and more reliable. I think “guarantee” is not a meaningful standard because a) I don’t think there can be such a thing as a perfect prompt, and b) humans do not meet that standard today.
- phil294 1y agoNeither are humans, so this argument doesn't really stand.
- maltalex 1y ago> Neither are humans, so this argument doesn't really stand. Even when we give a spec to a human and tell them to implement it, we scrutinize and test the code they produce. We don't just hand over a spec and blindly accept the result. And that's despite the fact that humans have a lot more common sense, and the ability to ask questions when a requirement is ambiguous.
- apwell23 1y agobut ppl writing that already knew that . so why are they writing this kind of stuff. what the fuck is even going on?