6 ms·
What I find fascinating that there is so little substance in this article about the quality of produced code and the medium. Is the code documented and tested?
by eithed 4mo ago
What I find fascinating that there is so little substance in this article about the quality of produced code and the medium. Is the code documented and tested? Is it understandable and extendable? Is it secure? What language, framework, database was used? Author mentions judgement and taste - well, is the code tasteful? Will the model rearchitecture the entire thing if I ask it to add new functionality, spending another 9.5h in tokens? I assume that the research part is domain knowledge = how different types of travel translate to time making it presentable; how did the author verify this?
These questions are even not about AI: if I were to give money to a human agency and were given something they tell me works, I would ask the same questions. If I did not know how to evaluate, I would hire people that do. With LLMs the verification part is what bothers me the most.
- hypfer 4mo agoBeing the first to release an article gives you great SEO or whatever. Doing the things you've mentioned takes time.
- jstummbillig 4mo agoLess fascinating when you consider that this is a non-coders perspective.
- eithed 4mo agoFair enough, but enterpreunership should, I guess, ask questions if given Next Big Thing has substance behind it or is it just snake oil.
- munk-a 4mo agoAh, but billions of dollars depend on those questions not being asked in a genuine manner. Don't you want a slice of that or are you an... AI skeptic thunder clashes.
- nomel 4mo agoSo, the perspective of the one that gains the most, that will value this the most, and that will pay the most? ;)
- unholiness 4mo agoYeah, this made it basically clickbait for me, in terms of time I wasted with the wrong expectation. The lack of downvotes on posts on HN has always felt like more of a bug than a feature to me.
- CobrastanJorji 4mo agoIt's still fascinating, but for a different reason. The "Concord" tool that got created bills itself as "Instrument-grade measurement of qualitative text. Explore in minutes, publish with honest statistics." Instrument-grade! How wonderful! That presumably means its accuracy has been ensured, and it's been carefully calibrated, right? What, nobody's ever measured or even examined the code? Well, no matter, let's go ahead and publish it and advertise it as "honest" "instrument-grade measurements."
- reedlaw 4mo agoYeah, the README looks like slop to me.
- adamtaylor_13 4mo agoI'm becoming more convinced these are questions of the Before Times. Yes, yes—heresy, I know. Yet, I can't deny the reality that I observe working with LLMs every day. If this truly is a step-function (as some are sgguesting), then I have absolutely zero concern for the quality of the code.
- fwip 4mo agoKind of a circular argument, isn't it? "Some people are saying it's very good at coding. If that's true, I don't care if the code is good."
- adamtaylor_13 4mo agoI didn't say I don't care if the code is good. I said I had zero concern for the quality of the code. That is, I do not have concern that the quality of the code will be a concern in and of itself. It's a subtle, but IMO important difference. We only care about code quality so as it gives us stable, understandable systems. Historically that meant a human had to read and understand it. Suppose a future where that's no longer the case, then we may still end up with stable, understandable systems without understanding every minutiae of the substrate. It's the same way I don't really know if my compiler is correct, but the behavioral patterns of my code suggest it is without me understanding anything about its code quality.
- deleted 4mo ago[deleted]
- grafporno 4mo agoIt's an ad.
- cgearhart 4mo agoI’m starting to realize that LLMs are really good at building low-stakes projects. Your questions mostly presume that the stakes are higher. The software will last a long time; the requirements will evolve; we can’t tolerate mistakes; etc. The trick to getting good at using LLMs for software is to learn how to make _all_ projects low-stakes.
- acedTrex 4mo ago> The trick to getting good at using LLMs for software is to learn how to make _all_ projects low-stakes. this doesn't really work in the real world. There are many things that actually matter, engineering is fundamentally about handling them.
- qaq 4mo agoYou don't need LLM for that. You make _all_ projects low-stakes by working on green field project using (insert buzzword soup of the day) and leaving for a new green field opportunity (that requires experience with buzzword soup of the day) before the project ships.
- DrJokepu 4mo agoNo, what you’re describing still requires you to do some actual work, and also, while you work there, there is still some level of accountability. A much, much better grift is coaching. Like, an AI coaching session for executives at the yearly executive retreat. You show up, spend a few hours going through some nonsense slides ChatGPT put together for you, you charge an eye watering fee for it, HR or whoever organizes it will gladly pay for it because it will make them look all cutting edge in front of the CEO, by the next day everyone will forget about it. No accountability at all!
- majormajor 4mo agoIn the LLM world you never get a chance to get paid to work on those greenfield projects because the person with the idea is churning the prototyping and discovery work themselves. If you want to get paid to work on software, you get involved after its found success and the stakes get higher. (Which assumes there are still significant areas where economies of scale reward that vs everybody just having their own DIY version of everything.)
- coldtea 4mo ago>What I find fascinating that there is so little substance in this article about the quality of produced code and the medium. I clicked one of his examples intrigued "a snake game where the snake is self-aware and crazy things happen;". Played for 1-2 minutes, and it's the classic 1980s snake game. Am I missing something? What is "self-aware" about it? Some funny messages at the bottom of the screen? And what are the "crazy things"?
- vunderba 4mo agoI had the exact same thought. To me, it feels like they just took the fairly common “sentient video game character” trope and bolted it onto a very conventional snake game. I will say, the act of eating creates a "bulge distortion" that flows down the length of the snake is a nice touch though.
- starshadowx2 4mo agoIt sounds like you either didn't play enough or you are missing the new mechanics that get added over time. There's definitely more to it than just regular snake.
- kesor 4mo agoYou didn't play long enough. There are layers and layers and layers of features in that game if you play for 10 minutes or more.
- nozzlegear 4mo agoCan you spoil it for us?
- an0malous 4mo agoThese posts are never written by software engineers, it’s always some tech exec, retired engineer, or VC. This author is apparently a professor at the Wharton School of Management? None of these people have to ship or maintain real products, they’re just making side projects. The only decent software engineering perspective I’ve seen has been from Mitchell Hashimoto.
- jimbokun 4mo agoWell that’s kind of the point. They can just summon bespoke software out of the ether that only handles the use cases of themselves and a few of their collaborators. Making “side projects” was mot possible for non-developers before powerful LLMs. Now it is.
- shimman 4mo agoMaking side projects isn't a trillion dollar industry tho, adding to the fact that we are facing another global supply chain crisis due to the Iran War; the US is about to commit the biggest self-own ever in the history of empire.
- zelphirkalt 4mo agoThe US has been on a course of self-owns ever since Trump got into office. That they still are a dominant power on the globe shows how much they were one before Trump, but it seems to be changing. At every self-own they commit, China laughs and inches up a little closer. I think we will see the day, when they are evenly matched in our lifetimes. But which self-own exactly do you mean, of the many there are?
- Schmerika 4mo agoThere are actually quite a few trillion dollar industries that exist thanks to "side projects". Apple was Woz's side project, once upon a time. Adsense came from Google's 20% time. Social media started as a side project. Forests grow from trees. Trees grow from seeds. More potential seeds = more potential forests.
- jimbokun 4mo agoDoes it matter to the people requesting the software if it acts in the way they expect?
- eithed 4mo agoTrue, but you should say that about every thing. Does it matter to you how the car drives, as long as it takes you to your destination? Well, yes, it matters: how will it deal with a crash, and if it's possible to replace a part and if anybody can just open it if you leave it outside. I will be amazed if somebody shows me their home-printed car, but if they'll try to sell it to me like a new one...
- crystal_revenge 4mo agoWe've lived in a software bubble for so long, most software engineers have completely forgotten that the purpose of (most) software is to solve a problem. If that problem solves the problem well and reliably it doesn't matter the quality of the code. In fact, that's the entire reason we care about "quality code", because we assume that quality code is code that does what you expect well and consistently. I say this as someone who hand writes code pretty much every night for fun, just to experiment with computation. Which, oddly, is more fun than ever because I don't feel like there's any need to connect this type of programming with "real world software", and I can really enjoy code for it's own sake, meanwhile my job is mostly just running agent loops (which I quite like as well).
- SpicyLemonZest 4mo agoI haven't forgotten that, I affirmatively think it's false. High quality code is necessary to solve problems reliably. Perhaps some people call things code quality when they don't matter (I really don't care what most variables are named), but there have always been teams who try to increase velocity by disregarding code quality, and from what I've seen AI does not stop them from shipping outages constantly.
- munksbeer 4mo agoExactly. Quality of code is a programming invention to make it easier to write and maintain correctly functioning applications. That is the entire purpose of "quality of code". If the end user experiences a correctly performing application, now, and in the future, they don't care at all what the code looks like. AIs could resort to a single global array of primitives and forget all about functions, and just use gotos if it helped them (it probably doesn't).
- chickensong 4mo agoYou probably don't care about the ingredients or engineering of asphalt, only if the road does its job well or is filled with potholes. Outside of the software industry, nobody gives a shit about code or databases.
- eithed 4mo agoI agree. But if I'm paying for the road (even as a taxpayer) I get angry that after a year it's full of potholes and that there are unnecessary signs warning about penguin crossing, making it cost 2 times more than it should have (and dont get me started why this road is really a highway leading to my house). I'd want certain qualities. And this article is basically = you will get a road, built quickly But yes, you are right - I don't build roads and don't know what is a price to build a road and how to determine the quality of correctly built one, nor I will ever care or learn.
- aix1 4mo ago> And this article is basically = you will get a road, built quickly That's not how I am reading it. You will get a road built exactly to your spec, quickly. So no penguin crossings unless you ask for them. I am also not entirely sure how the pothole argument translates.
- eithed 4mo agoThe road will be built to some specs, including features nobody asked for. If the corpus was trained for roads built in Arctic, you will get penguin crossings.
- Tylerian 4mo agoThe ingredients and composition of the tarmac is the difference between having the road full of pot holes after a week of use
- fwip 4mo agoSure, but if there's a trillion dollar company saying that it's going to replace all our road workers or engineers - I'd want to listen to the opinion of an expert. Some reporter from CNN driving over it like "yeah seems good to me, good this" has approximately zero persuasive power to me.
- spicyusername 4mo agothe quality of produced code and the medium A thought I have been tossing around in my head as the models get better is that it really may not matter what the code looks like. If the observed behavior of the software is good, then the software is good. If a bug, of whatever kind, can be fixed by a model on a vibe-coded codebase, then that's a fixable bug. If there are no exploitable vulnerabilities, then the code is secure. If the performance is adequate, then the code is performant. It simply does not matter what the code looks like if, from the outside, it does what its supposed to, and, from the inside, a model can fix the issue if one is found. More than ever, software engineering is now really a job about making sure the code is doing what its supposed to. And even if it DOES matter what the code looks like, you can have a model fix that too.
- skydhash 4mo agoThe thing is that a lot of code rely on multiple layers of abstractions with their own correctness and failure states. And then you overlay the domain correctness and failure cases on top of that. But all of those correctness are imaginary. The hardware only enforce a few (and it may be buggy). The OS adds some more (and it’s buggy). The compiler/interpreter may have bugs (but that’s rarely a nuisance) and the libraries are often brittle. There are cracks everywhere in the tower of abstractions. The code has never mattered. What has always mattered is the knowledge of what is the model of correctness of the software (programming as a theory by NauR), so that you can discern where a program is wrong. The thing is a crash or some other immediate errors are actually nice to have. You get to react immediately and can have a core dump or a stacktrace that points you the error. What is truly a terror is silent corruption (wrong order of operations, wrong values for a comparison that has expanded the idea of correctness, security issues that has been backdoored for years,…). As Hoare said: There are two ways of constructing a software design: One way is to make it so simple that there are obviously no deficiencies and the other way is to make it so complicated that there are no obvious deficiencies. The first method is far more difficult. LLM are very much the second kind. You write a lot of complicated code, and then you can no longer reason about their correctness.
- gofreddygo 4mo ago
- soraminazuki 4mo agoWelcome to every LLM discussion in the past 2 years or so. When asked for anything of substance, we're faced with a barrage of "but humans aren't good at this too!" Very few quantifiable evidence and lots of pure rhetoric.
- skydhash 4mo agoI’ve seen this pattern again and again, and I don’t bother replying. There’s also the “strong statement, and when you contradict it, they point out some particular circumstances that no one cares about”.
- munksbeer 4mo agoI think a lot of us have stopped talking to each other about this. I see it the other way round to you. I see constant scepticism and doubt that LLMs can build anything useful, and whenever provided with examples, the goalposts just move. And at my own firm, I think every developer is generating most of their code using agentic coding. We're still sceptical enough that we are doing the usual heavy handed human review process, so we're not seeing a huge speed up in delivery times, but we are seeing a volume increase. That is because writing the changes and raising the PRs is much faster, but also a lot of boring admin and support work is now mostly done by LLMs. Reports of instability, vague client requests, etc? Throw the LLM at them and it usually figure it out why I continue to engineer. So I know, first hand, that these things are very good. I also know second and third hand that pretty much every fintech in the industry is as heavily using agentic coding as we are. And then I come to HN or reddit and I see people telling us that they cannot write decent production code, and this is just wrong. This isn't opinion wrong, it is objectively wrong. Any fintech that wants to keep up will tell you this. I can't speak for other industries but I can't imagine they're different. So, I'm not sure what to conclude from this. I don't want to be uncharitable, but when HN/reddit posts just don't match the reality I see for myself, I have no choice but to categorise them as being emotionally driven to stick to a particular narrative, and so I can dismiss them.
- skydhash 4mo ago
- markoloko 4mo agoSo would you be more comfortable if the user them just prompted the AI to use a specific language, framework and database. Aren't we all just going to reddit and finding out what all goes best with what? But also I don't trust nothing from it, even though I've seen it.
- andai 4mo agoThese days it's uneconomical for human to verify AI generated code. So we ask the AI to do it. Like when we asked the FBI to audit itself and they found no problems :)
- otabdeveloper4 4mo agoDon't harsh my vibes, man.
- sexylinux 4mo agoIt still does make errors, yes? Because it is not usable, if we need to verify everything. AI is only interesting if it can do things that humans can not do. If you can verify results because you can do it yourself, then why use AI? It will just bind highly skilled people to do verification work. Instead these people should do the actual work, results will come quicker. So AI is only interesting to you / your org / humans if it can do things that you can not achieve. But if it still does errors, how could we ever know that super-invention by AI is not wrong? If we can not rely on the correctness of the result, it is not usable at all. AI must create reliable and correct results always. That was a very fundamental requirement for computing. This problem has not been solved.
- fisf 4mo agoBy that measure, most software developers should be unemployed.
- danlugo92 4mo agoYou can either adapt or survive man, coping and negation dont help, AI is here to stay and yes it does require pilots but this map would have taken you weeks to do, the AI did it in 10 hours, you can still dedicate a week to refactor. Also this is easily solved by .md spec files, this whole "bad code" cope is just FUD'
- Anamon 4mo agoI don't think that putting a text file saying "don't make mistakes" is going to get LLM output to the point where it doesn't need professional input, guidance, review and refinement anymore. They don't make these systems more deterministic. There have even been study results showing spec files reducing prompt adherence.
- jknoepfler 4mo agoThere also isn't any meaningful articulation of why this is a "leap forward"... literally everything claimed in the article has been claimed in the same breathless tones in articles written a year prior. I get that there's little sense in arguing with the MBA hivemind, but... c'mon. I manage two teams of highly motivated, largely pro-AI engineers. Both teams have independently concluded that they needed to ramp down GenAI usage because of code quality / maintainability concerns. Both teams have suffered from protracted outages caused by LLM jank not being sufficiently fenced off and guarded against. Both teams have expressed concern that the code generated by LLMs is far too verbose, full of slop, and rapidly becomes an unmaintainable mess. These are teams that are building non-trivial LLM solutions (deep agentic data synthesis and multi-modal data tagging). They are using the technology creatively and pro-actively, not just vibe-coding slop and throwing their hands up when it fails. Both teams will continue using GenAI coding agents, don't get me wrong - but the gains are incremental, not transformative, and need careful fencing to make sustainable. Nothing in these articles resonates as real. People who work in reality don't agree. I don't understand why this shit keeps getting attention (or rather I do, but the reasons aren't good).
- kordlessagain 4mo ago[dead]