12 ms·
When AI promises speed but delivers debugging hell
- iamflimflam1 2y agoIt’s not mentioned anywhere in the post. But would be good to hear what the total time was including all the problems.
- andix 2y agoMy social feeds are full of tech bros who keep telling people AI codes everything for them. AI obviously has some impressive coding skills, but for me it never really worked well. So is this just an illusion they create, or is it really possible to build software with AI, at least at a mediocre level? I'm looking for open source projects that were built mostly with AI, but so far I couldn't find any bigger projects that were built with AI coding tools.
- HL33tibCe7 2y agoCurrent-gen AI can write obvious code well, but fails at anything that involves complexity or subtlety in my experience
- frodo8sam 2y agoFor me ai has been pretty useful. Difference is I'm not a software engineer, I just write scripts to help me do my job. If I wrote bigger applications I doubt llms could help me.
- andix 2y agoAI is awesome for small coding tasks, with a defined scope. I don't write shell (or powershell) scripts anymore, AI does it now for me. But once a project has more than 20 source files, most AI tools seem to be unable to grasp the context of the project. In my experience AI is really bad at multi threading code and distributed systems. It seems to be unable to build its "mental model" for those kind of problems.
- Earw0rm 2y agoThis. It's good at Cmake, and mostly good at dealing with COM boilerplate (although hallucinations are still a problem). But threading and asynchronous code are implicit - there's a lot going on that you can't see on the page, you need to think about what the system is actually doing rather than simply the words to make it do the thing.
- dahousecat 2y agoThey are useful for small tasks like refactoring a method however big the whole project is
- rco8786 2y agoCurrent gen AI can spit out some very, very basic web sites (I won't even elevate to the word "app") with some handholding and poking and prodding to get it to correct its own mistakes. There is no one out there building real, marketable production apps where AI "codes everything for them". At least not yet, but even in the future it seems infeasible because of context. I think even the most pro-AI people out there are vastly underestimating the amount of context that humans have and need to manage in order to build fully fledged software. It is pretty great as a ridealong pair programmer though. I've been using Cursor as my IDE and can't imagine going back to a non-AI coding experience.
- jgilias 2y agoI think it’s that AI unlocks the ability to code something up and test an idea for people who’re technical enough to get it working, but not really developers themselves. It’s not (yet at least) a substitute for a good dev team that knows what they’re doing. But this is still huge, and shouldn’t be disregarded.
- loveparade 2y agoIn my experience, AI is good at building stuff in two scenarios: - You have zero engineering background and you use an LLM to build an MVP from scratch. As long as the MVP is sufficiently simple there is plenty of training data for LLM to do well. E.g. some kind of React website with a simple REST API backend. This works as long as the app is simple enough, but it'll start breaking down as the app becomes more complex and requires domain-specific business knowledge or more sophisticated engineering techniques. Because you don't understand what the LLM is doing, you can't debug or extend any of it. - You are an experienced developer and know EXACTLY what you want. You then use an LLM to write all the boilerplate for you. I was surprised at how much of my daily engineering work is actually just boilerplate. Using an LLM has made me a significantly more productive. This only works if you know what you're doing, can spot mistakes immediately, and can describe in detail HOW an LLM should be doing the task. For use cases in middle, LLMs kind of suck. So I think the comparison to a (very) junior engineer is quit apt. If the task is simple you can just let them do it. If the task is hard or requires a lot of context, you need to give them step by step instructions on how to go about it, and that requires that you know how to do it yourself.
- bdangubic 2y agothese are exactly my experiences as well. senior devs on my team and rocking and rolling with the AI. my junior devs have all but given up using it even after numerous retros etc…
- wanderingbort 2y agoI think it’s selection bias. Marketers are going to post the proof-of-concept that it works (if only in a small isolated scenario), algorithms are going to emphasize the more amazing “toys” this produces, over the boring rebuttals. In the end, you will see hundreds of examples where it worked and not the the thousands where it produced buggy or dangerous code. That attention does not map well to the important, hard, and more valuable parts of development. Anecdotally, I still find it to be useful and it’s improving. I do think it’s going to be an huge impact in time. Hype is part of the industry and it can be distracting to users, developers, and investors BUT it can also be useful (and I don’t know how to replace it) so, we live with it.
- javier2 2y agoSame experience. It has become pretty good at writing creative SQL queries though. Its actually rather good at that. When I am working on something niche, it does not help either. I have tried to make it build modern UI applications for myself using modern Java, but it just can't. It hallucinates libs and functions that does not exists, and I cant really get it to produce what I want. I have had better experiences with languages that are simpler and more predictable (Go), and languages with huge amounts of learning material available (Typescript / React). But I have been trying to build open source UI apps in JavaFX and GTK, it just can not help me when I am stuck
- williamcotton 2y agohttps://github.com/williamcotton/webdsl https://github.com/williamcotton/webdsl Made almost entirely with Cursor and Claude 3.5 Sonnet. 11k lines of C and counting.
- andix 2y agoThanks, that's exactly what I'm looking for.
- thomasfromcdnjs 2y agoI used Windsurf mostly on a feature to build out user authentication and then another tool to generate the PR documentation entirely. https://github.com/jsonresume/jsonresume.org/pull/176 https://github.com/jsonresume/jsonresume.org/pull/176 Meets my good enough standards fo sure
- BeefWellington 2y agoIt's funny that this is MIT licensed, expecting credit for uncopyrightable work.
- williamcotton 2y agoI work in copyright law. Familiarize yourself with the AFC test concept and I'm willing to have a conversation about what would be copyrightable in this project of mine. https://en.wikipedia.org/wiki/Abstraction-Filtration-Comparison_test https://en.wikipedia.org/wiki/Abstraction-Filtration-Compari...
- BeefWellington 2y agoYou misunderstand, the repo lists an MIT license which requires attribution. You want people to give you credit if they use this LLM-generated code. LLMs which were trained on the works of thousands of other developers with similar licenses, who are offered no similar credit here. It also claims copyright of the code as though you have authored it, but you're claiming here to have used LLMs to generate it. Seems like trying to have it both ways.
- dagw 2y agoAI isn't great at creating software, but it is great at writing functions. I often ask AI to "write a function that takes A, looks up B in a SQL database, and returns C, or write a function that implements the FooBar algorithm in C++" and on the whole that works pretty well. Asking it to write documentations for those functions also works really well. Asking it to write unit tests for those functions works pretty well (although you have to be extra careful, because sometimes the tests are wrong). What you have to do, and what AI cannot do well, is to decide where in the codebase to put those functions, and decide how to structure the code around those functions. You have to decide how and when and why to call each of those functions.
- javier2 2y agoWhen I have to be that specific with it, it would be faster for me to just write it directly in my normal IDE with great auto complete
- dagw 2y agoit would be faster for me to just write it directly in my normal IDE Then you are a much better developer than me (which you may very well be). I'd like to think I'm pretty good, and I've many times spent hours trying to think through complex SQL queries or getting all the details right in some tricky equation or algorithm. Writing the same code with an AI often takes 2-20 minutes. If it's faster for me, it might not be faster for everybody, but it is probably faster for many people.
- jprete 2y agoThe way to get better is to do it a lot. Every time you dig through a problem to solve it, you're not just learning about that problem - you're learning about every problem near it in the problem space, including the meta-problems about how to get information, solve problems, and test solutions. In a sense you're slowly building the LLM in your head, but it's more valuable there because of the much-better idea evaluation, and lack of network/GPU overhead.
- 2y ago
- mft_ 2y agoI experimented with Cursor over christmas, with writing a simple-ish Swift/SwiftUI app on iOS as the challenge. I can code fairly well in Python, moderately in JS, and almost not at all in Swift. I was using Cursor on a Mac, in parallel to XCode. Basically, it worked, but not without issues: - The biggest issue was debugging: because the bugs appeared in XCode, not Cursor, it either meant laboriously describing/transcribing errors into Cursor, or manually fixing them. - The 'parallel' work between Cursor and XCode was clunky, especially when Cursor created new files. It took a while to figure out a halfway-decent workflow. - At one point something screwed up somewhere deep in the confusing depths of XCode, and the app refused to compile altogehter. Neither Cursor nor I could figure it out, but a new project with the files transferred over worked just fine. But... after a few short hours' chatting, learning, and fixing, I had a functional app. It wasn't free of frustrations, and it's pretty far from the level where a non-coder could do the same, but it impressed me that it's already at the level where it's a decent multiplier of someone's abilities.
- jstummbillig 2y agoaider (AI assistant, that will do coding for/with you, depending on how you use it) has one of the more illuminating pieces of information on this. Here is a graph of percentage contribution of aider to aider development itself over time: https://aider.chat/HISTORY.html https://aider.chat/HISTORY.html
- cejast 2y agoI don't think it is an illusion. It can remove a lot of barriers to entry for some people, and this is probably what you're seeing in the anecdata. For example, my brother. He is what I'd refer to as 'tech-aligned' - he can and has written code before, but does not do it for a living and only ever wrote basic Python scripts every now and then to help with his actual work. LLM's have enabled him to build out web apps in perhaps 1/5 of the time it would have taken him if he tried to learn and build them out from scratch. I don't think he would have even attempted it without an LLM. Now it doesn't 'code everything' - he still has to massage the output to get what he wants, and there is still a learning curve to climb. But the spring-board that LLM's can give people, particularly those who don't have much experience in software development, should not be underestimated.
- andix 2y agoThere is a big gap between being able to create a somehow working application and shipping a product to a customer. Those claims are about being able to create a profitable product with 10x efficiency.
- namaria 2y agoCoding is trying to order bytes into doing arbitrary stuff that is useful because of some transient conjunction of factors in the real world. We have developed programming languages because coding in machine language is horrible, and over the decades we've refined them into tools people can use fluently and just directly think in code when they have to make a computer system behave in a certain way. Only someone who has never built anything of significant complexity and utility can think that putting natural language encoding between you and the bytes is a net positive.
- deleted 2y ago[deleted]
- s_dev 2y agoAI is a tool like any other. Autocomplete on steroids -- markov chains taken to the extreme. We already put natural language between us and the bytes. Hence why most keywords and variable names (a hard part of computer science) are in simple English and it is considered a net positive.
- namaria 2y agoThe memory and compute requirements to develop and run these models make no sense if the marginal improvement in autocomplete is the big end result. They only make sense in a world where machine can derive intent from natural language and actually conform to what people mean when they ask for something. This is clearly a fantastical result that LLMs are very short of.
- nailer 2y agoIt’s interesting. I would’ve agreed that ‘driving intent from natural language is something that LLMs have fallen far short of’ maybe a month ago. Since then, I spent a week trying to get cursor to work, and after dealing dealing with all the bugs, and restarting the composer each time with a new prompt, was able to get what I would consider a quality output for a moderately complex app (a parimutuel betting market). The issue isn’t that LLMs are terrible, it’s the software like cursor is buggy and poorly written. It should know that I don’t want to use code from an old version of the library I am using because the new library I am using is already in my projects dependencies. It should let me set up preferences for different programming languages. And preferences for all programming languages. So when I give it a prompt, it looks at the dependencies and language rules I already have set up, adds those to the prompt and produces the quality output I’m seeing now without me having to manually specify all those things. Short version: LLMs rule the software is just shitty.
- helpfulContrib 2y ago[dead]
- stuaxo 2y agoOh look, a load of future work to fix these. Why is this just like the last cost cutting exercise where the cheapest people in India produced a lot of "interesting" code.
- nailer 2y agoYou can update requirements, educate developers, and fix bad code with an LLM many orders of magnitude faster than you can with Wipro.
- SunlitCat 2y agoWay back in the 2000 (or even before that, can't remember!) I wanted to get into winsock programming. I found a page where someone from India explained that with examples. The variables, functions and so on had names like: a aa aaa b bb bbb It helped me to grasp the basic concept, but was kinda hard to follow, tho. :D
- llm_trw 2y agoBecause ignoring a heroic effort from all the women in India the number of Indian developers does not double every 4 years. The number of flops a gpu can output on the other hand does.
- scarface_74 2y agoSee also: almost every bespoke internal app written in FoxPro, VB, Excel with VBScript, etc
- HL33tibCe7 2y agoThis. Great, AI can produce code. But it produces code without inducing understanding of the code in the person who wrote (or rather supervised the production of) it, which is half the point. At some point AI will probably be good enough that this won’t matter. But it feels like we’re still a long way off that.
- BoredPositron 2y agoCan anyone explain why everyone is so hyper focused on speed? 500 images per second, 100 minutes video in 30 minutes, thousand lines of loc per hour. Who is going to consume all that?
- osmsucks 2y agoOther machines.
- martin-t 2y agoMost of what generative models produce is shit so they have to produce a lot in hopes _some_ of it is OK-ish. It's also about responsiveness. LLMs produce junior-level quality of code at a rate of hundreds of lines per minute. I need it to produce enough to spot where it's completely wrong as quickly as possible to I can change the prompt. It's like a edit-compile-run cycle which you also need to be fast or you lose attention. I was tempted to say it's another _step_ in the edit-compile-run but often the code is so bad I don't even bother compiling.
- JTyQZSnP3cQGa8B 2y agoThe images are almost good but still in the uncanny valley. The code is almost good but full of bad practices and hidden bugs and undefined behavior. Since most AI grifters are neither coders nor artists, all they can do is produce more more more capitalism-style.
- deergomoo 2y agoI'm firmly of the belief that most software would benefit immensely from us all slowing the hell down and putting more thought into what we build. But it would appear stability and a focus on core strengths doesn't sell nearly as well as endless new features for the marketing sheet added as quickly as possible.
- geor9e 2y ago"Who is going to consume 1000 lines of code per hour?" he types into his mass-manufactured thinking machine running an advanced operating system, before clicking reply, sending it across a global mesh of said devices.
- SunlitCat 2y agoWell, still trying to get into nvrhi, I went on to ask ChatGPT to write me an example program using it. To make it short, it got better when I made a project, uploaded the headers and docs of it as project files and moved my chat into that project as well. That said, AI can help you but needs a lot support from you to do things somewhat right.
- zug_zug 2y agoThis roughly mirrors my experience so far. Mind you I'm an extremely qualified engineer who has worked at FAANG. Except I'd add that as one gets experience working with the AI I can only assume they'd get much better at making it go smoothly. For example, I wouldn't manually rewrite localhost, I'd tell the AI "Why is localhost everywhere? Will this worker if I deploy to a droplet?" and it will fix it for you. Also I just paste error-messages directly into the AI and it usually knows how to fix them. Sometimes it's net positive, sometimes it's net-negative due to creating a mess that's really hard to get out of or debug. But I imagine it's only a matter of time until the scopes in which it's cost-effective go up. I don't like that AI is a threat of huge monopolistic and job-reducing potential, but I don't think downplaying it is a long-term strategy to combat that.
- skydhash 2y ago> For example, I wouldn't manually rewrite localhost, I'd tell the AI "Why is localhost everywhere? Will this worker if I deploy to a droplet?" and it will fix it for you. The solution is multi occur (emacs), quickfix list (vim), or any editors that have whole project find and replace.
- MaKey 2y agoWhich will also be much faster because you don't have to worry about sanitizing your code before sending it to an LLM or that the LLM made a mistake somewhere along the way.
- GeoAtreides 2y ago> I'm an extremely qualified engineer who has worked at FAANG. > I just paste error-messages directly into the AI ...
- scarface_74 2y agoI find it funny that commenters on HN actually think their having past or current experience working at a FAANG is some sort of signal for two reasons. On HN especially, that’s really nothing novel, many of us have (including me) and the only thing that it takes to get into one as a software engineer is memorizing the solution to coding problems. When I’m hiring - mostly for green field initiatives - coming from BigTech is usually a negative signal for me.
- fcatalan 2y agoThis has been my experience with a recent try to guide the LLM to a complete implementation of a small internal tool. I had in an hour what would have taken me 4 or 5 to write. But after that, it was an endless loop of the LLM adding logging code to find some bug and failing to fix it, only to add more logging code and ineffectual changes and so on. The problem is that even after it's lost at sea, it's still answering in a completely confident and self assured tone, so when you decide to take matters in your hands you might be too far gone from sanity and have an unfixable mess in your hands. I guess I can go back to where it strayed and retake it from there, but by now the experiment seems to be a failure.
- morsecodist 2y agoAt least in my experience as soon as something goes a little wrong it just gets worse from there. The more of it's confusion and contradictory information are in the chat history the worse it gets. It also has to make changes to the code so you accumulate these spurious changes and the problem gets more confusing. I've had some luck starting over with a new chat asking what is wrong but if that doesn't work I just assume I'm on my own.
- diggan 2y agoI've found that quality degrades really quickly after just the first reply, for some reason. They all seem heavily biased towards one-shot correct answers, and as you say, they go down the wrong path really quickly if you even get the first message slightly wrong. I tend to restart chats from the beginning pretty much all the time, because of this.
- iamflimflam1 2y agoI’ve also found this to be the case. Starting a new chat or in Cursor composer session puts things back on the right track. Also, prompting is really important. A lot of people just seem to think they have some kind of oracle - “fix the bug” - how is anything supposed to work from that?
- 2y ago
- fredgrott 2y agothe real revolution will be when an AI tool can just be powered by our laptop to use our own codebase as the input.... Until then its just nonsense pretending to be something else...
- owenthejumper 2y agoI have had great experience with Claude for coding, but you really need to be a programmer yourself, to be able to divide the problems into manageable chunks.
- bboygravity 2y agoSame here, I really don't get all the "it's totally useless for programming" posts on here. It makes me think many people haven't taken the time to actually learn to use the tool. It just feels like they tried Copilot or ChatGPT for 5 minutes last year and concluded that all LLM's are useless and will be useless forever. It makes me wonder if those people know that Claude 3.5 sonnet projects and/or Cursor with Claude exist? Do they not appreciate some help to document their code? Do they never need to write or quickly understand scripts or code in one of the 100's of languages/stacks they're not too familiar with that they might encounter in the wild? How to get out of yet another git mess? Build a proof of concept in an hour that would've taken you days? A refresher on how to set up x toolchain to get started asap (the nr 1 hardest thing in programming :p) etc etc.
- chillingeffect 2y agoSame here. I see these tools as teaching me patiently and challenging me (unwittingly) in areas where i'm out of depth. When i'm lucky they will do simpler stuff for me, but for $40/month, I don't feel entitled to a SaaS-unicorn-terraformer.
- MaKey 2y ago> Do they not appreciate some help to document their code? How does an LLM help there? What the code does should be obvious by looking at it, WHY it was written that way is the interesting question. Answering it often requires more context and domain knowledge. > Do they never need to write or quickly understand scripts or code in one of the 100's of languages/stacks they're not too familiar with that they might encounter in the wild? I'd rather take the time to do it myself because if I'm not familiar with a language/stack I won't be able to spot mistakes made by the LLM as easily. > How to get out of yet another git mess? Learn to solve the git issue and apply the knowledge in the future so you don't rely on yet another tool. > Build a proof of concept in an hour that would've taken you days? I question the premise. > A refresher on how to set up x toolchain to get started asap (the nr 1 hardest thing in programming :p) etc etc. How often do you do that? I think it's worth spending the time to do it yourself so you get an understanding of what exactly you're doing there. When you're done you can document the process and come back to it next time.
- yodsanklai 2y agoMaybe AI will shine when working with strongly typed languages. Most errors can be caught at compile time avoiding debugging hell.
- bandrami 2y agoThere's not enough of a corpus out there for the LLMs to snarf up
- nowittyusername 2y agoAI allows for more people to be more productive an therefore code more and produce more lines of code. That alone means more debugging needs to be done. when more people are doing anything, within that realm of action there is more liability naturally simply because of a larger participation in those actions.
- keyle 2y agoIf you don't why it works when it works, you won't know why it doesn't work when it doesn't work. The key issues here were staying on top of the AI's help. Use AI wisely: as an assistant, not as a drunken lead developer.
- jbirer 2y agoAnyone who has ever worked with VCs or shareholders before knows that, if you tell them the reality and limitations of something, they will either fire you or ignore what you say. They have been desperate to remove the leverage programmers have due to their skill and replace us with AI that they don't have to pay salaries to. All we can do at this point is just take VC money promising them exactly what they want to hear, that they will be able to replace us with a NLP model. Sometimes you just can't save people but you can profit from their voluntary fall from the cliff?
- siva7 2y agoIt delivers debugging hell if you don't know what you do which is usually the case for inexperienced developers. It assists experienced developers very well who can sort through which parts of output are useful from the AI and which not so.
- wlindley 2y agoGarbage in, garbage out. Code spewed by a random generator that has not the slightest understanding of what it is doing, whacked at by a hammer until it seems to be working. What is this supposed to produce other than a mass of bugs and vulnerabilities? "A.I." is utter garbage and always will be, it is foolish to think otherwise.
- deleted 2y ago[deleted]
- joshstrange 2y agoAI tools are just that, tools. I’ve said this since the very beginning of LLMs. I’ve yet to see anything change my mind. Aider/Devin/Copilot/Cursor/etc, all the different flavors of LLM tools are great but if you don’t know what you’re doing they are going to get stuck in a loop/corner/bad-path. Sometimes it takes 2-6+ exchanges before you realize it’s lost the thread which is why I love Aider’s “auto git commit” feature (defaults to on). You can always jump back X steps if you realize the LLM is lost. You also have to get a good feel for when it’s best if you make a change vs the LLM. Aider doesn’t handle new files and moving around massive chunks super well. It can do it but if I want to rename someone everywhere or break out components/types/etc into different files then I know I should be doing that in my IDE myself. Same for little syntax errors when a diff the LLM makes isn’t quite right. I spent a few nights last week using LLMs to help build a chrome extension to match my Amazon transactions with my YNAB transactions for the purpose of updating the memo field in YNAB with the item names I bought from Amazon to speed up my categorization and serve as history of what I bought (previously I did this whole process manually). I think it really helped and made the whole process go much faster. It really excels (for me) in UI. I’d like to think I’m pretty competent at writing code/logic but I’m not great at UI. In many projects I get bogged down when it comes to UI. If I get stuck coming up with a UI or I don’t like how something looks I can lose motivation to continue forward on it. With Aider I can ask for UI and while it might be abhorrent to a designer I think it looks pretty damn good (better than what I could do) and lets me focus on the logic. Aider also lets me try radical changes knowing I can easy reset back a few steps if it doesn’t work out. I’ve said many times at work that a huge power of LLMs is taking something that would take 30-60min down to <5min, specifically around things like little scripts to investigate a problem or get more details. For example, I might have a log that I can see there is data in that I want to extract. I know I can write a chained/piped command of sed/awk/grep/cut/sort/uniq/etc but it’s going to take some trial and error as well as time. With an LLM I can bang out the full command in 1-3 exchanges. Same deal with visualizing some piece of data in the logs (note: yes, we use Prometheus/Grafana but not everything can go in there and for new bugs/issues in the field I’m normally dealing with something we haven’t seen before and thus haven’t setup monitoring/alerting on). I’ve had LLMs churn out simple HTML/JS/CSS files that I can feed data into “graph all instances of this happening if X > Y and time is between A and B, etc”. Again, I can write this stuff from scratch but often don’t do it in practice because the ROI isn’t guaranteed. In the middle of a production issue do I want to waste 10-30+ min writing the script to see if I can prove a theory? No, it’s not worth it if it doesn’t pan out, but if I’m using an LLM and it takes me less than five minutes then I can throw a lot more stuff at the wall to see if it sticks.
- emporas 2y ago> LLMs are useless if you don’t understand the context > AI can be worse than useless when you don't understand the underlying technologies I made a saying about this some weeks ago: "A.I. can make the road for you, but you have to know where you are going". In Greek it sounds a little bit better. Also code is the truth, but it is not the only truth. The underlying computer, the network infrastructure and other things have an effect on the code. So, there could be a saying in addition to the first: "A.I. can make the road for you, but you have to test the road".
- fenomas 2y agoI put it: "copilot doesn't save me much thinking, but it saves a ton of typing".
- CharlieDigital 2y agoI have a non-technical friend who in the last two months has bootstrapped a SaaS startup using nothing but AI. He's got just over a handful of paying customers at this point on a monthly subscription[0]. I asked him to show me his process[1] after trying my hand (20 year, principal) and noticed a big difference in how we used AI: I instruct the AI how to code, he asks the AI to fix problems. In other words, I have a tendency to look at the code and ask the AI to fix it in more specific and direct ways that I want it fixed. On the other hand, if something doesn't work, my friend will copy/paste the error to the AI directly out of the dev tools console and ask the AI to fix the error. The two approaches are totally different. My lesson here is that you're not meant to debug AI generated code; hand the error off to the AI and let it fix itself. I think if you're debugging AI generated code, you're doing AI generated code wrong. If you're an experienced dev picking up AI coding, I think you need to shift your mindset entirely. Ideally, someone out there will just create a closed loop where the AI can fix itself when it finds an error (integrate some browser and autonomous test loop into Cursor, for example, and let it fix its own errors). Conclusion: if you're going to use AI to code, commit to it and use AI to fix the errors as well. Use AI for every aspect of it. [0] Yes, I'm sure there are security holes and code issues galore, but those can always be fixed later when he's proven the business model. [1] Yes, I have told him that he should create a YT channel or stream on Twitch because the content itself is super interesting how well he's been able to use AI.
- Madmallard 2y agoYeah this works for crud apps with conventional methods for accounts email payments and not at all for anything complex especially if it isn’t a super commonly used language or framework. Try coding a single game with AI that isn’t something done 10000 times already. It actually is impossible.
- llamaimperative 2y agoThe vast majority of code in the world is the former though
- 2y ago
- rchowe 2y agoI played with OpenHands for a few days (using gpt-4o since I already had an OpenAI account). I found it to be decent at writing new code, but then it had a hard time making changes when there was a lot of repetitive code (in a TypeScript / React project that I had it create with vite). One of the interesting things about OpenHands is that you can see what the AI is doing in the terminal window where you launched it. Since it can't really load the whole codebase into its context window, it does a lot of greping files, showing 10 lines on either side of the match, and then doing a search and replace based on this. This is pretty similar to what a human might do: attempt to identify the relevant function and change it. I think I might have better luck with a simpler project, e.g. a Sinatra or Flask app where each route is relatively self-contained. I might give it or Cursor another try in the future when the tech has progressed a bit.
- macNchz 2y agoI’ve built and iterated a bunch of web applications with Claude in the past year—I think the author’s experience here was similar to some of my first tries, where I nearly just decided not to bother any further, but I’ve since come to see it as a massive accelerant as I’ve gotten used to the strengths and weaknesses. Quick thoughts on that: 1. It’s fun to use it to try unfamiliar languages and frameworks, but that exponentially increases the chance you get firmly stuck in a corner like OP’s deployment issue, where the AI can no longer figure it out and you find yourself needing to learn everything on the fly. I use a Django/Vue/Docker template repo that I’ve deployed many production apps from and know like the back of my hand, and I’m deeply familiar with each of the components of the stack. 2. Work in smaller chunks and keep it on a short leash. Agentic editors like Windsurf have a lot of promise but have the potential to make big sweeping messes in one go. I find the manual file context management of Aider to work pretty well. I think through the project structure I want and I ask it to implement it chunk by chunk—one or two moving pieces at a time. I work through it like I would pair programming with someone else at the keyboard: we take it step by step rather than giving a big upfront ask. This is still extremely fast because it’s less prone to big screwups. “Slow is smooth and smooth is fast.” 3. Don’t be afraid to undo everything it just did and re-prompt. 4. Use guidelines—I have had great success getting the AI to follow my desired patterns, e.g. how and where to make XHRs, by stubbing them in somewhere as an example or explicitly detailing them in a file. 5. Suggest the data structures and algorithms you want it to use. Design the software intentionally yourself. Tell it to make a module that does X with three classes that do A, B and C. 6. Let the AI do some gold plating: sometimes you gotta get in there and write the code yourself, but having an LLM assistant can help make it much more robust than I’d bother to in a PoC type project—thorough and friendly error handling, nice UI around data validation, extensive tests I’m less worried about maintaining, etc. There are lots of areas where I find myself able to do more and make better quality-oriented things even when I’m coding the core functionality myself. 7. Use frameworks and libraries the AI “knows” about. If your goal is speed, using something sufficiently mainstream that it has been trained on lots of examples helps a lot. That said, if something you’re using has had a major API change, you might struggle with it writing 1.0-style code even though you’re using 2.0. 8. Mix in other models. I’ve often had Claude back itself into a corner, only to loop in o1 via Aider’s architect mode and have it figure out the issue and tell Claude how to fix it. 9. Get a feel for what it’s good at in your domain—since I’m always ready to quickly roll back changes, I always go for the ambitious ask and see whether it can pull it off—sometimes it’s truly amazing in one shot! Other times it’s a mess and I undo it. Either way over time you get an intuition for when it will screw up. Just last week I was playing around with a project where I had a need to draw polygons over a photograph for debugging purposes. A nice to have on top of that was being able to add, delete, and drag to reshape them, but I never would have bothered coding it myself or pulling in a library just for that. I asked Claude for it, and got it in one shot.
- alwinaugustin 2y agoThe current state of so-called AI does not provide much meaningful assistance in software development beyond basic tasks such as explaining workflows, breaking down thought processes, and performing simple conversions. I believe that generative AI, in its current form, is not true artificial intelligence. Rather, it is a sophisticated prediction engine that lacks genuine reasoning or understanding. True AI should be capable of comprehending problems and devising its own solutions, rather than merely generating statistically likely outputs. Until AI reaches that level of cognitive ability, its applications in the real world remain limited, and much of what we see today is largely hype. Tokenization and embeddings merely help models predict the most probable next token, a process that is executed at scale using vast computational resources. This is not intelligence but large-scale probabilistic prediction. The terminology used in computer science, especially in recent years, can often be misleading.
- notjoemama 2y agoDo we really want this? As soon as possible, employers will fire software engineers and replace them with AI. I’m positive they will not care about what AI can do, only how many salaries they can eliminate and still achieve the same results. You and I will not be the inheritors of AI.
- lazide 2y ago“and still achieve the same results”. That is the part that won’t actually happen, at least pretty quickly.
- fellowmartian 2y agoYes, because employers will also be replaced by AI. Technology penetration won’t stop at some arbitrary boundary, it will go all the way through to logical conclusion. We have a chance at qualitatively better world, but we’ll need to act and push for new economic systems - when the time comes.
- pavel_lishin 2y agoMaybe I read too much science fiction, but my first thought when speaking about "true AI" isn't the worry that a lot of us will get fired, it's the worry that we'll have created an army of digital slaves.
- peterkelly 2y agoI dream of a world in which more investment is put into creating better programming languages and runtime environments than trying to use LLMs as a way of coping with the complexities of current systems.
- 3ptow 2y agoIf you drop all pretenses and use a photocopier to steal code directly instead of performing an elaborate laundering step, you will not have these issues.
- cushychicken 2y agoOne thing I think would have helped the author: write a spec first. Seriously. It seems stupid. But AI works a lot better with a written spec. The incredible thing is that the AI can actually be an excellent resource for writing the spec. And it will actually produce better code when you feed the spec back into said AI! The current generation of AI seems to have fooled a lot of people into thinking that somehow you can jump straight to coding. (Well, you can, and it will probably work if you want to make something small or limited in scope.) Not so! But, on the bright side, it’s just as good at design as code if you ask the right questions! I say this having used 4 and 4o extensively in this manner. Just started using sonnet3.5 in this way in the last month or so, and it is amazing at this.
- cduzz 2y agoPrograms are communication between 2 loosely coupled audiences -- the humans who have to maintain / modify the code and the computer that gets to run the code. Human language, used to convey ideas to other humans, is imprecise. It's fine that it's imprecise because the media (humans) have both good error correction and a reasonable set of global defaults. Computer languages require enormous precision because they're some mechanical translation to a set of machine code runtime. Perhaps you can train an LLM on lots of code, and it'll find semantic relationships between some clever code it's been trained on to and your specific request. Perhaps not, and it'll just give a dumb answer or an incorrect answer, (ideally some code copilot will actually try running the candidate answer code against your specific ask?) -- but once the answer gets complex you run into the "it's much harder to debug code than write it, so don't write code that's almost too complex for you to understand" problem. At work, I constantly have to remind people "don't use math data structures for identities" "but int is smaller" "Are you ever going to want the 95th percentile customerID?" "no that's silly" "then it isn't a number". Or I get to constantly remind people "a string with lots of curly braces and quotes isn't necessarily json; if you're not using a serialized API and just sending bytes to stdout someone else has to parse it" "but I'm using a logging library" "does anything else ever send stuff to stdout while your logging library is running?" "oh yes, we're going to open a ticket to debug that." So I'm not optimistic that running code written by a machine is long-term viable. That said -- there are situations where machine generated code works -- I think it's been a long time since anyone manually drew masks for etching dies when making CPUs.
- sega_sai 2y agoIt seems there is a battle of two opposite view-points. One is that LLMs are just dumb autocompletes with no ability to understand anything. Another is that LLMs can already right now be substitutes for programmers. I personally thing it is neither, but for experts who know what they are doing it is a massive time saver. I.e. in cases you know what code you want to write, but it's tedious, LLMs can do for you. Also LLMs are great in cases where you are less familiar with a new API, language, but have generally good understanding of programming. Despite my broadly positive view on usefulness of LLMs, I do not think they are good enough (yet) to build a full system from scratch without an expert supervisor. This should not IMO be used as a 'proof' they are dumb autocompleters.
- deadbabe 2y agoExperts who know what they are doing have long had alternatives beyond LLMs to make their work faster. They have open source libraries, stack overflow, tutorials, documentation, simple code generator tools and snippets. The speed up we’re seeing is from LLMs basically caching all those things into a huge mathematical model and retrieving information in summarized form ready for consumption. And while speed is always nice, LLMs are expensive, require maintenance themselves to maintain relevant context, are still error prone, and terrible at true innovation. In a few years we’ll be talking about the big “AI crash” and “what went wrong” when it has been obvious to experts all along. Winter is coming.
- sega_sai 2y agoI am sorry, but the comparison of 'stack overflow' and tutorials to LLMs is bizarre. The amount of time to get to the answer from LLMs is drastically shorter. And claiming that the they only 'cache thing' is just wrong. They are certainly capable of correctly answering things that were not directly in their training set.
- deadbabe 2y agoDo you have any examples of a question you could ask to an AI right now that you couldn’t find from a basic search on stack overflow and Google? Didn’t think so.
- yapyap 2y agoNever believe the snake oil sellers
- n_ary 2y agoThe issue with AI is, it generates what it is trained on. Most publicly available coding contents/examples are just docs or blogspam(geeksforgeeks/javapoint/whatever) where mostly surface level code is mostly peddled. Even, many OSS(small scale) do not have best practices or good code base, just enough to get whatever is needed to be done. Now when you train AI on such data, it’ll excel reproducing(statistically) the same thread of code. Once the quality of training data improves(somehow getting access to high quality codebase behind corporate walls by promoting these assistants and ingesting the codebase), the output improves. There was a popular saying, garbage in garbage out.
- ibloomt 2y agoHeh, after decades of functional programmers being the "well, actually..." crowd at every conference, turns out they were right all along. Just for the wrong reasons! The pitch: AI generates tons of plausible-looking garbage Static types catch garbage at compile time OCaml/F#/Haskell fans quietly sipping tea in the corner The irony? We spent years debating static vs dynamic typing for human developers. But the killer use case may ended up being catching AI hallucinations. Finally, a business case for monads that doesn't require a PhD! Time to dust off those Haskell books. Who knew safety could be so profitable? Plot twist: Category theory becomes a required interview question by 2025
- morsecodist 2y agoI appreciate posts that are about practical usage of AI and it's strengths/weaknesses and the kind of conversation it generates. Conversations about AI are tough for me to navigate because there are camps of people that seem very invested in AI being either omniscient or completely useless. I regularly see people saying that AI is at the level that it can replace engineers or build whole apps. When I try this with state of the art models, I am seeing results that are nowhere close. That said, I still use AI every day during my development and I have a flow I think makes me way more productive. I want more conversations like this about the mechanics of using AI as it currently, and honestly evaluating it's strengths and weaknesses without getting into hypothetical debates about the future or whether or not the AI "understands".
- senko 2y agoContent marketing for a new text editor thinly disguised as AI rage-bait. HN fell for it hard - 156 points, 180 comments (as of this writing). Well done Nick! :) And congrats on launching Codescribble! Hope to see a "how my post on AI grew my userbase" followup in a few weeks!
- avidphantasm 2y agoI was recently experimenting with local-only LLM coding assistant in JetBrains products. They did speed things up a bit, but I quickly realized that they were essentially automating the creating a copy-paste errors, resulting in time lost to debugging errors I never would have introduced myself, so I stopped using them.