6 ms·
"When was the last time you reviewed the machine code produced by a compiler?" Compilers will produce working output given working input literally 100% of my t
by atomicnumber3 8mo ago
"When was the last time you reviewed the machine code produced by a compiler?"
Compilers will produce working output given working input literally 100% of my time in my career. I've never personally found a compiler bug.
Meanwhile AI can't be trusted to give me a recipe for potato soup. That is to say, I would under no circumstances blindly follow the output of an LLM I asked to make soup. While I have, every day of my life, gladly sent all of the compiler output to the CPU without ever checking it.
The compiler metaphor is simply incorrect and people trying to say LLMs compile English into code insult compiler devs and English speakers alike.
- jen729w 8mo ago> Meanwhile AI can't be trusted to give me a recipe for potato soup. Because there isn’t a canonical recipe for potato soup.
- Jensson 8mo agoThat is not the issue, any potato soup recipe would be fine, the issue is that it might fetch values from different recipes and give you an abomination.
- D-Machine 8mo agoThis exactly, I cook as passion, and LLMs just routinely very clearly (weighted) "average" together different recipes to produce, in the worst case, disgusting monstrosities, or, in the best case, just a near-replica of some established site's recipe.
- antonvs 8mo ago> ... some established site's recipe. At least with the LLM, you don't have to wade through paragraph after paragraph of "I remember playing in the back yard as a child, I would get hungry..." In fact LLMs write better and more interesting prose than the average recipe site.
- D-Machine 8mo agoIt's not hard to scroll to the bottom of a page, IMO, but regardless, sites like you are mentioning have trash recipes in most cases. I only go with resources where the text is actual documentation of their testing and/or the steps they've made, or other important details (e.g. SeriousEats, Whats Cooking America / America's Test Kitchen, AmazingRibs, Maangchi for Korean, vegrecipesofindia, Modernist series, etc) or look for someone with some credibility (e.g. Kenji Lopez, other chef on YouTube). In this case the text or surrounding content is valuable and should not be skipped. A plain recipe with no other details is generally only something an amateur would trust. If you need a recipe, you don't know how to make it by definition, so you need more information to verify that the recipe is done soundly. There is also no reason to assume / trust that the LLMs summary / condensation of various recipes is good, because cooking isn't something where you can semantically condense or even mathematically combine various recipes together to get one good one. It just doesn't work like that, there is just one secret recipe that produces the best dish, and LLMs don't know how to judge quality of recipes, mostly. I've never had an LLM produce something better or more trustworthy than any of those sites I mentioned, and have had it just make shit up when dealing with anything complicated (i.e. when trying to find the optimal ratio of starch to flour for Korean fried chicken, it just confidently claimed 50/50 is best, when this is obviously total trash to anyone who has done this). The only time I've ever found LLMs useful for cooking is when I need to cook something obscure that only has information in a foreign language (e.g. icefish / noodlefish), or when I need to use it for search about something involving chemistry or technique (it once quickly found me a paper proving that baking soda can indeed be used to tenderize squid - but only after I prompted it further to get sources and go beyond its training data, because it first hallucinated some bullshit about baking soda only working on collagen or something, which is just not true at all). So I would still never trust or use the quantities it gives me for any kind of cooking / dish without checking or having the sources, instead I would rely on my own knowledge and intuitions. This makes LLMs useless for recipes in about 99% of cases.
- antonvs 8mo ago> This makes LLMs useless for recipes in about 99% of cases. But you spent a whole lot of time essentially describing how 99% of recipe sites are also useless. What's the difference?
- lebuin 8mo agoThere's also no canonical way to write software, so in that sense generating code is more similar to coming up with a potato soup recipe than compiling code.
- LiamPowell 8mo ago> Compilers will produce working output given working input literally 100% of my time in my career. In my experience this isn't true. People just assume their code is wrong and mess with it until they inadvertently do something that works around the bug. I've personally reported 17 bugs in GCC over the last 2 years and there are currently 1241 open wrong-code bugs. Here's an example of a simple to understand bug (not mine) in the C frontend that has existed since GCC 4.7: https://gcc.gnu.org/bugzilla/show_bug.cgi?id=105180 https://gcc.gnu.org/bugzilla/show_bug.cgi?id=105180
- grey-area 8mo agoThese are still deterministic bugs, which is the point the OP was making. They can be found and solved once. Most of those bugs are simply not that important, so they never get attention. LLMS on the other hand are non-deterministic and unpredictable and fuzzy by design. That makes them not ideal when trying to produce output which is provably correct - sure you can output and then laboriously check the output - some people find that useful, some are yet to find it useful. It's a little like using Bitcoin to replace currencies - sure you can do that, but it includes design flaws which make it fundamentally unsuited to doing so. 10 years ago we had rabid defenders of these currencies telling us they would soon take over the global monetary system and replace it, nowadays, not so much.
- zx8080 8mo ago> It's a little like using Bitcoin to replace currencies [...] At least, Bitcoin transactions are deterministic. Not many would want to use a AI currency (mostly works; always shows "Oh, you are 100% right" after losing one's money).
- grey-area 8mo agoSure bitcoin is at least deterministic, but IMO (an that of many in the finance industry) it's solving entirely the wrong problem - in practice people want trust and identity in transactions much more than they want distributed and trustless. In a similar way LLMs seem to me to be solving the wrong problem - an elegant and interesting solution, but a solution to the wrong problem (how can I fool humans into thinking the bot is generally intelligent), rather than the right problem (how can I create a general intelligence with knowledge of the world). It's not clear to me we can jump from the first to the second.
- keyle 8mo agoYou're correct, and I believe this is only a matter of time. Over time it has been getting better and will keep doing so.
- bigstrat2003 8mo agoMaybe. But it's been 3 years and it still isn't good enough to actually trust. That doesn't raise confidence that it will ever get there.
- keyle 8mo agoYou need to put this revolution in scale with other revolutions. How long did it take for horses to be super-seeded by cars? How long did powertool take to become the norm for tradesmen? This has gone unbelievably fast.
- grey-area 8mo agoI think things can only be called revolutions in hindsight - while they are going on it's hard to tell if they are a true revolution, an evolution or a dead-end. So I think it's a little premature to call Generative AI a revolution. AI will get there and replace humans at many tasks, machine learning already has, I'm not completely sure that generative AI will be the route we take, it is certainly superficially convincing, but those three years have not in fact seen huge progress IMO - huge amounts of churn and marketing versions yes, but not huge amounts of concrete progress or upheaval. Lots of money has been spent for sure! It is telling for me that many of the real founders at OpenAI stepped away - and I don't think that's just Altman, they're skeptical of the current approach. PS Superseded.
- suddenlybananas 8mo ago>super-seeded Cute eggcorn there.
- antonvs 8mo ago*superseded It comes from the Latin "supersedēre", which taken literally, means "sit on top of". "Super" = above, on top of. "Sedēre" = to sit. "Super" is already familiar to English speakers. "Sedēre" is the root of words like sedentary, sedan, sedate, reside, and preside. The more metaphorical meaning of "supersede" as "replace" developed over time and across languages, but the literal meaning is already fairly close.
- anematode 8mo agoI'm trying to track down a GCC miscompilation right now ;)
- keyle 8mo agoI feel for you :D
- andai 8mo agoThis is obviously besides the point but I did blindly follow a wiener schnitzel recipe ChatGPT made me and cooked for a whole crew. It turned out great. I think I got lucky though, the next day I absolutely massacred the pancakes.
- D-Machine 8mo agoI genuinely admire your courage and willingness (or perhaps just chaos energy) to attempt both wiener schnitzel and pancakes for a crew, based on AI recipes, despite clearly limited knowledge of either.
- bonesss 8mo agoRecent experiments with LLM recipes (ChatGPT): missed salt in a recipe to make rice, then flubbed whether that type of rice was recommended to be washed in the recipe it was supposedly summarizing (and lied about it, too)… Probabilistic generation will be weighted towards the means in the training data. Do I want my code looking like most code most of the time in a world full of Node.js and PHP? Am I better served by rapid delivery from a non-learning algorithm that requires eternal vigilance and critical re-evaluation or with slower delivery with a single review filtered through an meatspace actor who will build out trustable modules in a linear fashion with known failure modes already addressed by process (ie TDD, specs, integration & acceptance tests)? I’m using LLMs a lot, but can’t shake the feeling that the TCO and total time shakes out worse than it feels as you go.
- andai 8mo agoThere was a guy a few months ago who found that telling the AI to do everything in a single PHP file actually produced significantly better results, i.e. it worked on the first try. Otherwise it defaulted to React, 1GB of node modules, and a site that wouldn't even load. >Am I better served For anything serious, I write the code "semi-interactively", i.e. I just prompt and verify small chunks of the program in rapid succession. That way I keep my mental model synced the whole time, I never have any catching up to do, and honestly it just feels good to stay in the driver's seat.
- 8mo ago
- pcl 8mo ago”I've never personally found a compiler bug.” I remember the time I spent hours debugging a feature that worked on Solaris and Windows but failed to produce the right results on SGI. Turns out the SGI C++ compiler silently ignored the `throw` keyword! Just didn’t emit an opcode at all! Or maybe it wrote a NOP. All I’m saying is, compilers aren’t perfect. I agree about determinism though. And I mitigate that concern by prompting AI assistants to write code that solves a problem, instead of just asking for a new and potentially different answer every time I execute the app.
- Ygg2 8mo agoCompilers don't change output assemby based on what markdown you provide them via .claude. Or what tone of voice in prompt you gave them. Or if Saturn is in Aries or Sagittarius.
- bostik 8mo agoEverything more complex than a hello-world has bugs. Compiler bugs are uncommon, but not that uncommon. (I must have debugged a few ICEs in my career, but luckily have had more skilled people to rely on when code generation itself was wrong.) Compilers aren't even that bad. The stack goes much deeper and during your career you may be (un)lucky enough to find yourself far below compilers: https://bostik.iki.fi/aivoituksia/random/developer-debugging-progression.html https://bostik.iki.fi/aivoituksia/random/developer-debugging... NB. I've been to vfs/fs depths. A coworker relied on an oscilloscope quite frequently.
- nneonneo 8mo agoI had a fun bug while building a smartwatch app that was caused by the sample rate of the accelerometer increasing when the device heated up. I had code that was performing machine learning on the accelerometer data, which would mysteriously get less accurate during prolonged operation. It turned out that we gathered most of our training data during shorter runs when the device was cool, and when the device heated up during extended use, it changed the frequencies of the recorded signals enough to throw off our model. I've also used a logic analyzer to debug communications protocols quite a few times in my career, and I've grown to rather like that sort of work, tedious as it may be. Just this week I built a VFS using FUSE and managed to kernel panic my Mac a half-dozen times. Very fun debugging times.
- bombolo 8mo ago[dead]
- senko 8mo ago> Compilers will produce working output given working input literally 100% of my time in my career. I've never personally found a compiler bug. First compilers were created in the fifties. I doubt those were bug-free. Give LLMs some fifty or so years, then let's see how (un)reliable they are.
- wtetzner 8mo agoWhat I don't understand about these arguments is that the input to the LLMs is natural language, which is inherently ambiguous. At which point, what does it even mean for an LLM to be reliable? And if you start feeding an unambiguous, formal language to an LLM, couldn't you just write a compiler for that language instead of having the LLM interpret it?
- senko 8mo ago1) Determinism isn't the same as reliability. Compilers are deterministic (modulo bugs), but most things in life are not, but can still be reliable. The opposite also holds: "npm install && npm run build" can work today and fail in a year (due to ecosystem churn) even though every single component in that chain is deterministic. 2) Reliability is a continuum, not a discreet yes/no. In practice, we want things to be reliable enough (where "enough" is determined per domain). I don't presume this will immediately change your mind, but hopefully will open your eyes to looking at this a bit differently.
- wtetzner 8mo ago> I don't presume this will immediately change your mind I'm not saying that AI isn't useful. I'm just claiming it's not analogous to a compiler. If it was, you would treat your prompts as source code, and check them into source control. Checking the output of an LLM into source control is analogous to committing the machine code output from a compiler into source control. My question still stands though. What does it mean for a tool to be reliable when the input language is ambiguous? This isn't just about the LLM being nondeterministic. At some point those ambiguities need to be resolved, either by the prompter, or the LLM. But the resolution to those ambiguities doesn't exist in the original input.
- rootnod3 8mo agoAbsolutely this. I am tired of that trope. Or the argument that "well, at some point we can come up with a prompt language that does exactly what you want and you just give it a detailed spec." A detailed spec is called code. It's the most round-about way to make a programming language that even then is still not deterministic at best.
- wtetzner 8mo agoAnd at the point that your detailed specification language is deterministic, why do you need AI in the middle?
- rootnod3 8mo agoExactly the point. AI is absolutely BS that just gets peddled by shills. It does not work. It might work for some JS bullcrao. But take existing code and ask it to add capsicum next to an ifdef of pledge. Watch the mayhem unfold.
- bayindirh 8mo agoCompilers and processors are deterministic by design. LLMs are non-deterministic by design. It's not apples vs. oranges. They are literally opposite of each other.
- Scene_Cast2 8mo agoJust to nitpick - compilers (and, to some extent, processors) weren't deterministic a few decades ago. Getting them to be deterministic has been a monumental effort - see build reproducibility.
- idopmstuff 8mo ago> Meanwhile AI can't be trusted to give me a recipe for potato soup. This just isn't true any more. Outside of work, my most common use case for LLMs is probably cooking. I used to frequently second guess them, but no longer - in my experience SOTA models are totally reliable for producing good recipes. I recognize that at a higher level we're still talking about probabilistic recipe generation vs. deterministic compiler output, but at this point it's nonetheless just inaccurate to act as though LLMs can't be trusted with simple (e.g. potato soup recipe) tasks.
- wtetzner 8mo ago> The compiler metaphor is simply incorrect If an LLM was analogous to a compiler, then we would be committing prompts to source control, not the output of the LLM (the "machine code").