10 ms·
I code with multiple LLMs every day and build products that use LLM tech under the hood. I dont think we're anywhere near LLMs being good at code design. Exis
by ssalazar 1y ago
I code with multiple LLMs every day and build products that use LLM tech under the hood.
I dont think we're anywhere near LLMs being good at code design.
Existing models make _tons_ of basic mistakes and require supervision even for relatively simple coding tasks in popular languages, and its worse for languages and frameworks that are less represented in public sources of training data.
I am _frequently_ having to tell Claude/ChatGPT to clean up basic architectural and design defects.
Theres no way I would trust this unsupervised.
Can you point to _any_ evidence to support that human software development abilities will be eclipsed by LLMs other than trying to predict which part of the S-curve we're on?
- fragmede 1y agohttps://chatgpt.com/c/681aa95f-fa80-8009-84db-79febce49562 https://chatgpt.com/c/681aa95f-fa80-8009-84db-79febce49562 it becomes a question of how much you believe it's all just training data, and how much you believe the LLM's got pieces that are composable. I've given the question on the link as an interview questions and had humans been unable to give as through an answer (which I chose to believe is due to specialization on elsewhere in the stack). So we're already at a place where some human software development abilities have been eclipsed on some questions. So then even if the underlying algorithms don't improve, and they just ingest more training data, then it doesn't seem like a total guess as to what part of the S-curve we're on - the number of questions for software development that LLMs are able to successfully answer will continue to increase.
- jackthetab 1y agoUnable to load conversation 681aa95f-fa80-8009-84db-79febce49562
- xyzzy123 1y agoI can't point to any evidence. Also I can't think of what direct evidence I could present that would be convincing, short of an actual demonstration? I would like to try to justify my intuition though: Seems like the key question is: should we expect AI programming performance to scale well as more compute and specialised training is thrown at it? I don't see why not, it seems an almost ideal problem domain? * Short and direct feedback loops * Relatively easy to "ground" the LLM by running code * Self-play / RL should be possible (it seems likely that you could also optimise for aesthetics of solutions based on common human preferences) * Obvious economic value (based on the multi-billion dollar valuations of vscode forks) All these things point to programming being "solved" much sooner than say, chemistry.
- cap4 1y agoThis is correct. No idea how people don't see this trend or consider it
- energy123 1y agoThis is my view. We've seen this before in other problems where there's an on-hand automatic verifier. The nature of the problem mirrors previously solved problems. The LLM skeptics need to point out what differs with code compared to Chess, DoTA, etc from a RL perspective. I don't believe they can. Until they can, I'm going to assume that LLMs will soon be better than any living human at writing good code.
- AnIrishDuck 1y ago> The LLM skeptics need to point out what differs with code compared to Chess, DoTA, etc from a RL perspective. An obviously correct automatable objective function? Programming can be generally described as converting a human-defined specification (often very, very rough and loose) into a bunch of precise text files. Sure, you can use proxies like compilation success / failure and unit tests for RL. But key gaps remain. I'm unaware of any objective function that can grade "do these tests match the intent behind this user request". Contrast with the automatically verifiable "is a player in checkmate on this board?"
- energy123 1y agoI'll hand it to you that only part of the problem is easily represented in automatic verification. It's not easy to design a good reward model for softer things like architectural choices, asking for feedback before starting a project, etc. The LLM will be trained to make the tests pass, and make the code take some inputs and produce desired outputs, and it will do that better than any human, but that is going to be slightly misaligned with what we actually want. So, it doesn't map cleanly onto previously solved problems, even though there's a decent amount of overlap. But I'd like to add a question to this discussion: - Can we design clever reward models that punish bad architectural choices, executing on unclear intent, etc? I'm sure there's scope beyond the naive "make code that maps input -> output", even if it requires heuristics or the like.
- ArthurStacks 1y agoI run a software development company with dozens of staff across multiple countries. Gemini has us to the point where we can actually stop hiring for certain roles and staff have been informed they must make use of these tools or they are surplus to requirements. At the current rate of improvement I believe we will be operating on far less staff in 2 years time.
- nnnnnande 1y ago[flagged]
- ArthurStacks 1y ago[flagged]
- nnnnnande 1y agoOh, just like every other business then! That's a nice strategic differentiator. Look, I'm sure focusing on inputs instead of outcomes (not even outputs) will work out great for you. Good luck!
- ArthurStacks 1y agoWeve done this since 1995 and it works perfectly well.
- DonHopkins 1y ago[flagged]
- ArthurStacks 1y ago[flagged]
- 1y ago
- Kinrany 1y agoWe're talking about predicting the future, so we can only extrapolate. Seeing the evidence you're thinking of would mean that LLMs will have solved software development by next month.
- ssalazar 1y agoIm saying, lets see some actual reasoning behind the extrapolation rather than "just trust me bro" or "sama said this in a TED talk". Many of the comments here and elsewhere have been in the latter categories.
- AIoverlord 1y ago[dead]
- thecupisblue 1y agoYou're using them in reverse. They are perfect for generating code according to your architectural and code design templete. Relying on them for architectural design is like picking your nose with a pair of scissors - yeah technically doable, but one slip and it all goes to hell.
- ssalazar 1y agoIm using them fine. Im refuting the grandparent's point that they will replace basically all programming activities (including architecture) in 5 years.
- piokoch 1y agoWell, I have asked LLM to fix some piece of Python Django code so it uses pagination for the list of entities. And LLM came up with the working solution, impressively complicated piece of Django ORM code, which was totally needles, as Django ORM has Paginator class that does all the job without manual fetching pages, etc. LLM sees pagination, it does pagination. After all LLM is an algorithm that calculates probability of the next word in a sequence of words, nothing less and nothing more. LLM does not think or feel, even though people believe in this saying thank you and using polite words like "please". LLM generates text on the base of what it was presented. That's why it will happily invent research that does not exist, create a review of a product that does not exist, invent a method that does not exist in a given programming language. And so on.
- cheema33 1y ago> I code with multiple LLMs every day and build products that use LLM tech under the hood. I dont think we're anywhere near LLMs being good at code design. I too use multiple LLMs every day to help with my development work. And I agree with this statement. But, I also recognize that just when we think that LLMs are hitting a ceiling, they turn around and surprise us. A lot of progress is being made on the LLMs, but also on tools like code editors. A very large number of very smart people are focused on this front and a lot of resources are being directed here. If the question is: Will the LLMs get good at code design in 5 years? I think the answer is: Very likely. I think we will still need software devs, but not as many as we do today.
- dubcanada 1y agoGood code design requires good input. And frankly humans suck at coding, so it will never get good input. You can’t just train a model on the 1000 github repos that are very well coded. Smart people or not, LLM require input. Or it’s garbage in garbage out.
- lytefm 1y ago> I think we will still need software devs, but not as many as we do today. I'm more of an optimist in that regard. Yes, if you're looking at a very specific feature set/product that needs to be maintained/develop, you'll need less devs for that. But we're going to see the Jevons Paradox with AI generated code, just as we've seen that in the field of web development where few people are writing raw HTML anymore. It's going to be fun when nontechnical people who'd maybe know a bit of excel start vibe coding a large amount of software, some of which will succeed and require maintenance. This maintenance might not involve a lot of direct coding either, but a good understanding of how software actually works.
- jpadkins 1y ago> I think we will still need software devs, but not as many as we do today. There is already another reply referencing Jevons Paradox, so I won't belabor that point. Instead, let me give an analogy. Imagine programmers today are like scribes and monks of 1000 years ago, and are considering the impact of the printing press. Only 5% of the population knew how to read & write, so the scribes and monks felt like they were going to be replaced. What happened is the "job" of writing language will mostly go away, but every job will require writing as a core skill. I believe the same will happen with programming. A thousand years from now, people will have a hard time imagining jobs that don't involve instructing computers in some form (just like today it's hard for us to imagine jobs that don't involve reading/writing).
- lallysingh 1y agoThe software tool takes a higher-level input to produce the executable. I'm waiting for LLMs to integrate directly into programming languages. The discussions sound a bit like the early days of when compilers started coming out, and people had been using direct assembler before. And then decades after, when people complained about compiler bugs and poor optimizers.
- pjmlp 1y agoExactly, I also see code generation to current languages as output only an intermediary step, like we had to have those -S switches, or equivalent, to convince developers during the first decades of compiler existence, until optmizing compilers took over. "Nova: Generative Language Models for Assembly Code with Hierarchical Attention and Contrastive Learning" https://arxiv.org/html/2311.13721v3 https://arxiv.org/html/2311.13721v3
- GTP 1y ago> I'm waiting for LLMs to integrate directly into programming languages. What do you mean? How would this look like in your view?
- coredog64 1y agoNot OP, but probably similar to how tool calling is managed: You write the docstring for the function you want, maybe include some specific constraints, and then that gets compiled down to byte code rather than human authored code.