6 ms·
Claude Is Not a Compiler
- desktopentree 3mo ago[flagged]
- customguy 3mo agoNevermind "comments", has anyone with a name and reputation claimed they are?
- formerly_proven 3mo agoSpec-driven development is sold on the idea that you regenerate modules or entire services when you change the spec. It's a really dumb idea but was briefly popular in enterprise, probably still is in some companies.
- tossandthrow 3mo agoA compiler does not transalte specs to code. Specs are denotational by nature. A program could synthesize a program that adheres to the specs. But that is not what we understand by a compiler that generally has to preserve operational semantics.
- lukasschwab 3mo agoThe title is a reference back to "Is Claude Code a Compiler?" by the same author, linked from the first sentence: https://commaok.xyz/ai/is-claude-a-compiler/ https://commaok.xyz/ai/is-claude-a-compiler/ And then claim is attributed right there: > The hands-down highlight was the talk by Erik Schluntz: Vibe coding in prod. > Among other things, he drew an analogy between LLMs and compilers.
- customguy 3mo agoDidn't read the article because the conclusion is obvoius enough to me as is, never heard of Erik Schluntz before, and drawing an analogy between two things means very little anyway. "this bike is about as heavy as that table" doesn't require the rebuttal that bikes aren't tables, for example. so far I mostly saw it spewed forth on HN
- frollogaston 3mo agoHeard it a bunch of times from my superiors at work, if that counts.
- bakugo 3mo agoI've heard the "if you hate vibe coding you must also hate compilers" strawman so many times, I'm sure it's been said by someone with a "name and reputation" at least once.
- couchand 3mo agoThe argument being made here is really incredible when you unpack it. 1) The construction of the Empire State Building was particularly effective due to the depth of human-to-human collaboration. 2) Isn't it great that we can burn a bunch of dinosaur blood to convince ourselves that we don't need other humans?
- kaonwarb 3mo agoI believe you are projecting your personal misgivings onto the argument actually being made. May be better to write your own rebuttal instead?
- johnea 3mo agoI'm pretty sure they just did...
- badrequest 3mo agoA rebuttal would spend some time trying to make a case, this is a petulant complaint.
- johnea 3mo agoOne person's petulant complaint is another person's rebuttal. I think what you're really saying is you don't agree with OP's perspective...
- doug_durham 3mo agoThe whole resource argument against AI is quickly becoming obsolete. Time to move on to a new argument.
- luckydata 3mo agonot in any way that matters, no. this is a strawman and a really bad one too.
- ryaniscool 3mo ago
- eigencoder 3mo agoI really like what exe.dev is doing. And their philosophy of how cloud services could work is amazing. I have a monthly subscription and it's really cool to just vibe code a little website with Shelley and have it running in minutes. I feel like I have so much fun using their service, and spend so much less time thinking about what things will cost, or what cloud service to use. It's just so much simpler and the friction is so much less compared to other cloud services.
- varenc 3mo agoAgreed. Big love for exe.dev. For quick low stakes projects, using their Shelly assistant is the simplest way to give an LLM a fully persistent VM with web services I've found.
- torginus 3mo ago> But our VMs start fast, so fast that even if we created the DNS entries before creating the VM, our users still had to sit around waiting for DNS to propagate, which occasionally took minutes, not seconds. To my (limited) understanding, this is not a good idea, and is an unfixable problem from the server side. Companies, VPNs or ISPs or routers, often use their own DNS servers, and those can have caching logic, which means it doesn't matter how fast your own DNS implementation is, as the users lookup request wont hit your DNS server, it'll hit an intermediate cache.
- deleted 3mo ago[deleted]
- tibordp 3mo agoPresumably these are new subdomains, if a caching DNS resolvers never saw the domain yet, it would simply query the authoritative server. There's a small nuance here that DNS spec also caches negative (NXDOMAIN) results, but the TTL of those is controllable by the SOA record on the domain that contains it. exe.xyz has a fairly short 10 second negative TTL, so even if you tried to resolve it before it was set up in the authoritative DNS, it wouldn't stay negative-cached very long (though some DNS caches put a lower bound on the TTL and cache it for longer) > dig +short soa exe.xyz ns1.exe.dev. hostmaster.exe.dev. 1854440 86400 7200 1209600 10
- eigencoder 3mo agoUsually when you create the VM, it gets its own DNS subdomain that looks like `my-new-vm-name.exe.xyz`. (I'm not affiliated with exe.dev, I just use their product). So I doubt there would be any cached requests for that particular subdomain before the VM is created.
- dspillett 3mo agoNXDOMAIN results and other errors have their own TTL value though, so if someone queries too early in the process you could be waiting a little while for their cache to check again. The value for this is set in the SOA record for your DNS zone. For example, example.org/example.com seems to have TTLs set at 300s (5 minutes) for A records, but the negative caching value is 1800s (30 minutes). If your VM domain is set the same way, you could see a premature lookup causing issues for 30 minutes instead of just 5. Even the 5 minutes could be a pain if your VM spin-up process expects to be able to lookup via the name very early and will fall over if it can't.
- vivzkestrel 3mo ago- it is a very very very very fancy autocomplete
- eraserj 3mo agoI'll go a step further and say Claude is an Interpreter. It even has just-in-time compilation: it can directly follow a spec and generate small snippets live. I'm pretty sure we'll soon see services (admittedly highly inefficient) built on Claude-as-a-backend.
- tovej 3mo agoIt's not. It's a cargo cult developer that produces more code than you can review and confidently pushes bugs to prod.
- whiplash451 3mo agoWhy inefficient? If Claude continuously makes the system more efficient based on demand, there's no reason for the system to be less efficient than if it was managed by humans.
- DauntingPear7 2mo agowords have meaning
- qwertox 3mo agoWhen Claude does its thinking I'm often reminded of https://xkcd.com/303/ https://xkcd.com/303/
- frollogaston 3mo agoI don't want to hear this ever again. No Patrick, mayonnaise is not an instrument.
- fcarraldo 3mo agoit is when used to make music
- frollogaston 3mo agoClaude doesn't even make machine code though. Or it can if you really want to, but that's now how people use it. If you wanna include transpilers or interpreters, fine but the answer is still no, otherwise our git commits would just be prompts that get built either in CI or JIT. People deliver code, not prompts. Maybe that could change some day, but even then, it wouldn't be a useful designation. It's like saying "oh technically the Americas are an island." There's a reason we still distinguish between a transpiler and compiler.
- deleted 3mo ago[deleted]
- armchairhacker 3mo agoA compiler is an algorithm and Claude isn't*. A compiler almost never produces a wrong output, even when compiling an extremely complicated program. But a compiler must be clearly defined and is limited to input/commands it's defined for. A compiler will "correctly" process input into unintuitive output, but that's not always what the user intended (e.g. omitting large sections of code that are undefined behavior). Claude makes mistakes, but can process any input/command expressible in training data, and has more potential to infer when something is unintentional. Tools like Claude are great at building algorithms, then tests to verify them. The algorithms they build aren't just compilers, but any software that can be verified. Even though Claude may not be a compiler, it has the potential to (at least if it improves) create production-grade software that can be verified, like other compilers. But maybe not good UX without human input (expanded to any subjective experience, like video games). * Not a "source language to target language" deductive algorithm. Technically Claude is an algorithm to predict the next token, but acting as a compiler or anything else it's an inductive heuristic, because it guesses (https://stackoverflow.com/q/2334225 https://stackoverflow.com/q/2334225)
- tester756 3mo ago>A compiler almost never produces a wrong output back log of compiler's bugs can be pretty large
- doubled112 3mo ago> almost never But not never. "Almost" is doing a lot of work there.
- armchairhacker 3mo agoIn a production-grade compiler like LLVM they are rare enough that it compiles giant projects like Linux and Chromium without detectable issues. Maybe eventually Claude can, since it can write tests, but compilers / (deductive) algorithms have other advantages like efficiency and predictability.
- 3mo ago
- dataviz1000 3mo ago[flagged]
- xav_authentique 3mo ago> I read a vanishingly small amount of the actual code. This sentiment kind of saddens me. I'm all for burning tokens to write throwaway code just to prototype a solution, but I don't get not reading (or at least familiarizing yourself with) the code that you will deploy to prod.
- Jtarii 3mo agoAI psychosis.
- IronWolve 3mo agoAnd claude isnt a replacement for the chain of humans, you dont know what you dont know, thats why we have a person in a role to gatekeep each step. Anyone using claude can see it misses steps that an engineer wouldnt, it can make bad choices that a sysadmin wouldnt, it can pick the wrong order a project manager wouldnt, it can pick wrong cost models that a bean counter wouldnt, etc etc etc But, I'm pretending it is and living with those mistakes.
- qudat 3mo ago> I’d say that, in all the ways that matter, I understand the code. Sure, if I had to hand-edit it now, there’d be a serious learning curve. But I won't have to. Famous last words, but point is taken. The non deterministic nature of an llm breaks the metaphor that they are like a compiler. However, it’s not foreign to compilers to receive feedback from the running program (PGOs), so there are still parallels to the feedback we provide LLMs that guide their “optimization”. I think even calling LLMs a non-deterministic compiler isn’t accurate either so ultimately I agree the metaphor doesn’t quite work. I do think LLMs are like compilers in terms of how they changed how we build programs from a historical context. But that’s about it.
- sarchertech 3mo agoThe non-determinism is theoretical solvable. Chaos (prompt instability) isn’t. When you change a single word in the input, you get a completely different output. Because of this LLMs will never be compilers.
- trjordan 3mo ago> Claude wasn’t just a compiler here. I never handed off a task and let an agent make a bunch of decisions in order to reduce it to practice. > I’d say that, in all the ways that matter, I understand the code. I think the dissonance here is really important, and not a bad thing at all. A lot of the decision _were_ handed off the the AI, but they weren't the decisions the author cared about. This is a big selling point of AI! If something is doable with a computer, it’ll figure it out. 30 minutes and 200m tokens later, it’ll take any idea and declare “the feature is fully implemented.” The hard part is figuring out where to inject that friction, so you can see where it's making decisions for you that matter. The author approached this by incrementally building the thing, reviewing and poking and prodding at every step. A week of attention following a bunch of design discussions is fast, but that's still not trivially cheap. I want to see us talk more about the decision exhaust of agents, because the better the models get, the more decisions we'll want them to make. I wrote a bit more here: https://tern.sh/blog/compiler-never-says-no/ https://tern.sh/blog/compiler-never-says-no/
- gritzko 3mo agoLLM-is-a-compiler is indeed a simplistic approach. I wrote a rebuttal to the yesterday's Cursor post, may reuse it here https://replicated.live/blog/follow-up https://replicated.live/blog/follow-up The idea that a 835-page spec "just exists" and we run an LLM to implement it is completely flawed. Specs do not appear out of nowhere, they co-evolve with the code. If you have the code, why do you want to generate it again? Good software is made as a product of numerous feedback loops and LLMs let you operate those loops faster. They do not supplement the entire process though. In the end, a good product is a barrel of distilled feedback.
- wrs 3mo agoIndeed, this is a much more realistic and useful framing of what to do with the current capabilities of LLMs. They work so fast, at such a high level, that you can speedrun the co-evolution process and end up with a spec and code that works. You can actually use second-system (fourth? fifth? I lost track) syndrome to your advantage because it’s so cheap to throw things away. And they don’t just write the system, they also write the test harness, the benchmarks, the failure simulation, the statistics…all the stuff that seems like “overhead” but lets you drive that evolution with data. I’ve kind of drifted into this mode a bit at a time over the last year, but hadn’t stepped back to make such a coherent explanation of it.
- atomicnumber3 3mo agoSomething I have always (even pre-LLMs) found funny to think about is - all code possible to run on a computer already exists. It's some permutation of all the bits of available memory. It's in there somewhere. So, suppose you want a specific program. 1. Some huge % of those possible programs are obviously not the one you want (most don't even compile). 2. Remaining programs might look similar to the one you want, but are buggy enough to be completely unusable, unreadable to the point of intractability, and so on. 3. Remaining programs look substantially correct but upon using for > a few mins you note major bugs that make it still not correct enough. 4. Remaining programs look substantially correct and seem to generally do your will but contain a long-tail showstopper subtle bug that corrupts saved data, or makes all the output subtly incorrect, etc. 5. Remaining programs might be useful, even if they're mildly annoying. 6. This process can probably continue for several thousand iterations until you finally find "it." The program you wanted. Or... one of them. There's probably still 10k+ candidate programs left at this stage. Our job has always been to get to 5 and aspire to 6. Indeed most of software development is just doing 5->6 in a loop. I think LLMs help us get to somewhere between 3 and 5 faster than we used to. And a big problem with them is that programs in 3 vs 5 all already look substantially correct and there's no way to know if you're getting 3 or 5. Generally, the above is not how it _felt_ to write software pre-LLMs, it was just a cute way to imagine what you're doing. Now it's weirdly apropos.
- dataviz1000 3mo agoClaude is more like a Probabilistic Turing Machine. [0] It's correctness isn't deterministic but rather a distribution. It is predictable. There are lots of places in computer science that determinism isn't necessarily the best. - UDP video calls let packets vanish or arrive corrupted, because waiting for retransmissions would freeze the picture. - Stochastic gradient descent picks a random mini-batch of training data and treats it as the whole dataset, because computing the true gradient on every step would make training infeasible. - Speculative execution in modern CPUs guesses which branch a program will take and rolls back when wrong, because waiting for the real answer would leave half the silicon idle. What we want is something like a Las Vegas algorithm, a probabilistic machine with a cheap verifier. [0] https://en.wikipedia.org/wiki/Probabilistic_Turing_machine https://en.wikipedia.org/wiki/Probabilistic_Turing_machine
- skydhash 3mo agoDeterminism isn’t important, but there are usually threshold of acceptable performance in all off those cases. Machine Learning was all about improving accuracy even when you know that there’s probably an error somewhere. But the current LLM hypers refuses to acknowledge that the answer to a prompt may be wrong (even though the providers have a warning line about it). Instead of assessing the risk and provide corrective methods, all they do is pushing to use it more and more everywhere.
- kloud 3mo agoSpecs might become one solution for coping with the need to review increased volume of code. A spec is a higher level of abstraction than code, which is a higher level of abstraction than machine code. The industry made the transition to higher-level once, paradigm is changing so it might happen again. The workflow I imagine is either deriving specs from the conversation or reverse engineering the code to spec, review and edit the spec which should be tighter and much more compressed, then deterministically compile to code. Of course we don't want to be spec-first only, that would be going back waterfall, but doing iterations back and forth. Now, Claude is not a compiler because it is closed and non-deterministic (they do opaque processing on server, hiding reasoning tokens), but LLMs might be. We refer to a piece of code from npm/pip by name to get some code by downloading it. We then have lockfiles with hashes to ensure integrity. Currently we are vibing it, but in the future we could refer to a piece of code by prompt/spec and getting the code by inferring it. To ensure integrity, the lockfile would be hashes of open weights and inference code (and ironing out implementation details like non-determinism due to GPU scheduling, etc.).
- skydhash 3mo ago> A spec is a higher level of abstraction than code, which is a higher level of abstraction than machine code No it’s not. Just like a quick doodle is not an higher level representation of the Mona Lisa. Sure for someones that knows the latter, it can suggest it. Or for someone that doesn’t know it, it may provide some basis of conversation. But it’s not the real thing. You can’t provide a doodle of something and expect an artist that hasn’t seen it to paint it. Doodling’s value is that it lets you iterate on ideas without the accidental constraints that comes with the implementation (choosing paints, finding references materials, deciding colors,…). Some necessary choices are kept for later, while you decide on the most important ones. So a spec is useful in designing software, but it’s severely lacking in implementing it. And the decisions made in implementation are as important as the ones made in the design. Even more after a while in production.
- toonvanvr 3mo agoIt's funny how you describe something very close to what I was attempting to design. It was meant to trickle down from tickets to deterministic code. Data wise, it should be like a pyramid of tickets diluting layer by layer into leaf nodes which were "implementable" as statements in code. I'm not sure if that description makes sense read by someone else. I think what made sense was envisioning a nanoswarm of LLMs (anticipating ASIC performance) diluting specs into semantic logic nodes, but a conversion from these to deterministic code made the vibe coded experiment come to halt. Your lock approach could be a shortcut to that. Totally off-topic: it's quite intriguing that you can feel once friction starts building during a design phase. Suddenly everything slows down. I wonder if it's quantifiable and therefore can identify "wrong" design choices made by either humans or LLMs. (Hoping some claw bot pics this up to finish my idea on my github tix repo in the initial-design branch, as main is empty ~ MPL2)
- js8 3mo agoI still think it is (in the source-code generating mode) a compiler, just like LiquidHaskell is. (I haven't used LH but it essentially can automatically supply functions based on the type conditions you specify.) In my view, reasoning LLMs have learned a large number of close to true sentences of "logic" in natural language. These rules are not consistent (unlike typing rules of LH that allow creation of provably correct programs), but they work in practice better than LH. Essentially, they represent an unsound formal specification of natural language, together with many useful (almost)tautologies that helps solving constraint problems (just like LH does, but again unlike LLM, provably correctly). LLMs proved that human language can be formalized reasonably close to soundness. I believe we can have a sound formalization of human language, and that kind of formal language will be once superior to LLMs. It will also not be as complicated as LLMs and require only fraction of compute to use.
- jumploops 3mo ago> By the time I was ready to build a keeper, I had accumulated a scar-tissue document that was empirically sufficient to guide an agent through most of the important decisions, at every layer, ranging from high level goals through architecture down to the occasional low level detail, such as the exact shape of the data type for load-bearing concurrent caches. Waterfall is dead, long live waterfall!
- davidpapermill 3mo agoI've heard something along the lines of "Claude is like a compiler: source code is the new object code, you don't look at that anymore" many times. And I don't really think this is true. Compilers are usually deterministic, and whilst we can find edge cases, it's nothing like an AI agent writing all the code for you. I think you have two choices, given the Claude is a code generator and not a compiler: (a) you review most or all of code to make sure it makes sense, or (b) you trust but verify via a strong test suite, potentially also created by Claude. The problem with (a) is that you lose a lot of the speed-up. The problem with (b) is that you have no human oversight and the code may be incomplete, badly designed, or plain wrong. Currently we review all code because correctness is extremely important to what we do, but that comes at a cost. I don't know what the answer is here, in general. Does trust build over time? Do the models just get so good we can trust them to make zero mistakes?
- QuercusMax 2mo agoPart of the value of LLMs and humans is nondeterminism. Pair nondeterministic output with strictly verified results (proper tests) and you can create a useful working system. Humans can't build a system perfectly, and agents definitely can't. And agents (like humans) will build a different system every time even with the same prompt. Even if it's just trivial differences like array vs linked-list, there are still differences. The value in having executable code is that it won't change its behavior unless you deliberately modify it. I don't think there are any shortcuts. Models will get smarter and smarter and make fewer mistakes, but you'll still want to produce real source code to execute because it's consistent (you can sell it as a product, set it and forget it, etc.), and most importantly: running real code is orders of magnitude faster than having an AI either run the process "manually", or have it write the code out multiple times.
- davidpapermill 2mo agoThanks, that's really helpful. What do you think happens in terms of testing/reviewing going forward? I'd really appreciate your thoughts on that.
- fzeroracer 3mo ago> I’d say that, in all the ways that matter, I understand the code. Sure, if I had to hand-edit it now, there’d be a serious learning curve. But I won't have to. And more importantly, I can reason about the system, share perspectives with my colleagues, and guide agents on future work. If there's a serious learning curve to editing code then you don't understand the code. We used to call that 'on-boarding' when you brought a new engineer on the team as they got up to speed as to how the codebase worked.
- agentultra 3mo agoThis seems to be grasping for analogies to make sense of the work this person was doing. A compiler does a lot more than source code translation. There are specifications that tell us the de jure specifications of the language, if there is one, and then we have to recognize the de facto implementations of said language. The often disagree and leave much on the table. Some times on purpose, such as implementation details, and other times by omission. Users of this compiler expect a deterministic compilation of the source text into the target code but it’s rarely 1:1. There are optimization passes, inlining, barriers, etc. And then there are the run-time effects of executing a program! I think the analogy gets a little weak because natural language is not a precise enough language to specify discrete systems. What it sounds like the author is doing is bypassing decision points with other people and delaying making decisions themselves until the LLM agent forces them to? Which is a fine approach but I don’t think the analogy with a compiler is necessary. People seem to have a hard enough time understanding branch prediction and thread barriers.
- _dwt 3mo agoThere is a lot of rehashing of LLM arguments in the comments here, but I'd love to read an actual distributed systems expert's take on the DNS architecture the author and/or Claude ended up with. One of my personal fears about LLM coding is that some of those implicitly outsourced decisions turn out to matter more than expected. Does the "timeline" stamp make sense? (I know there's not much detail to work from.)
- znkr 3mo agoTimeline from the article sounds a lot like a logical clock, which is a well known primitive in distributed systems: https://en.wikipedia.org/wiki/Logical_clock https://en.wikipedia.org/wiki/Logical_clock
- overgard 3mo agoI'm kind of getting to the point where if I know something is vibe coded, I just won't use it. It's not an anti-AI thing, it's just a quality thing. Pretty much every piece of vibe coded software I've used has been bad in some regard. The worst ones are the ones that aren't obviously bad but rather subtly bad in dangerous ways (the article about OpenCode yesterday definitely made me nope out of using that)
- borzi 3mo agoThe entire "Don't Read the Code" argument is just insane to me and ironically when you see who the people are that are advocating it, it's often the same people that talked about microservices, the latest java script framework, no-code or any other methodology that was in vogue at the time just to avoid focusing on solving the real problems; The illusion that they are just one "smart step" away from achieving a 10x productivity boost that will make everything else trivial...
- veloxxn 3mo ago[flagged]
- redlewel 3mo agothe shill is strong with this post
- nunez 3mo ago> exe.dev VMs have nice domain names: vm-name.exe.xyz. When we start a new VM, we add a CNAME entry or three. Easy, right? > But our VMs start fast, so fast that even if we created the DNS entries before creating the VM, our users still had to sit around waiting for DNS to propagate, which occasionally took minutes, not seconds. > We did the obvious thing: We wrote our own DNS server, so that DNS always immediately matched the source of truth. And life was good. > But latency matters, so we added regions. And just like that, DNS became the long pole again, because all DNS was served out of Oregon. Also, deployments caused tiny DNS outages. To fix this, all we needed now was a geographically distributed but fully consistent DNS server. > We did what a sensible engineer does when faced with a hard problem: cheat. We vibe-engineered a distributed DNS server tuned to our specific needs. I'm sure they have already considered this, but what problem does this solve that couldn't be solved by using dnsmasq, unbound or something like that? Why reinvent the wheel?
- philipwhiuk 3mo agoThat doesn't require vibe-coding, which is what they are selling as the solution to every problem.
- tpetry 3mo agoThe question is more what this really solves. They said they built all this because DNS propagation was too slow. But secondary dns servers by e.g. your isp will still have old records when you change something. So what did they earn? Every is correct within their network pretty fast. But for a normal user there isnt any difference because 8.8.8.8 or 1.1.1.1 still cache old records?
- vuciuc 3mo agoyo dawg! I heard you like distributed services So I made a distributed service for your distributed name service so you can distribute your names to the service that distributes them on the onternet
- zX41ZdbW 2mo agoWhy not a wildcard DNS pointing to a few Anycast IP addresses with proxies?
- cadamsdotcom 2mo agoHeadline true, article false. AI exists at the boundary of codification. When you understand something well enough to express it in a deterministic way, it is no longer worthwhile to keep asking AI for that thing over and over - you should codify it. But wait, I hear you say? AI has knowledge! And we can rely on that knowledge and sprinkle on top some brief instructions and it'll get things right and work out the details itself! That is spec driven development. And this is great one single time. Except, knowledge shifts. You are outsourcing part of what you've codified to society as a whole, in the form of language model knowledge. You're also trusting current and future models to always produce something that meets the requirements, and never be quantized under you, and produce the same result next time despite being a probabilistic system with no memory, and and and. Better to freeze everything in place in your codebase so you get predictable results - in other words, codify it! Let's store the desired state of that button in code in a very complicated thing that's existed for decades called a string, so it always says "Finish" and never says "Done" - even if we regenerate the app. That's pretty unexciting and doesn't justify tens of billions of investment. Which means you won't hear it from anyone with tokens to sell. But it's what you have to do to make something people want.
- valentynkit 2mo ago[dead]