5 ms·
I literally had the same experience when I asked the top code LLMs (Claude Code, GPT-4o) to rewrite the code from Erlang/Elixir codebase to Java. It got some th
by manca 1y ago
I literally had the same experience when I asked the top code LLMs (Claude Code, GPT-4o) to rewrite the code from Erlang/Elixir codebase to Java. It got some things right, but most things wrong and it required a lot of debugging to figure out what went wrong.
It's the absolute proof that they are still dumb prediction machines, fully relying on the type of content they've been trained on. They can't generalize (yet) and if you want to use them for novel things, they'll fail miserably.
- hammyhavoc 1y agoThey'll never be fit for purpose. They're a technological dead-end for anything like what people are usually throwing them at, IMO.
- zer00eyz 1y agoI will give you an example of where you are dead wrong, and one where the article is spot on (without diving into historic artifacts). I run HomeAssistant, I don't get to play/use it every day. Here, LLM's excel at filling in the (legion) of blanks in both the manual and end user devices. There is a large body of work for it to summarize and work against. I also play with SBC's. Many of these are "fringe" at best. LLM's are as you say "not fit for purpose". What kind of development you are using LLM's for will determine your experience with them. The tool may or may not live up to the hype depending how "common", well documented and "frequent" your issue is. Once you start hitting these "walls" you realize that no, real reason, leaps of inference and intelligence are still far away.
- SecuredMarvin 1y agoI also made this experience. As long as the public level of knowledge is high, LLMs are massively helpful. Otherwise not so much and still hallucinating. It does not matter if you think highly of this public knowledge. QFT, QED and Gravity are fine, AD emulation on SAMBA, or Atari Basic not so much. If I would program Atari Basic, after finishing my Atari Emulator on my C64, I would learn the environment and test my assumptions. Single shot LLMs questions won't do it. A strong agent loop could probably. I believe that LLMs are yanking the needle to 80%. This level is easy achievable for professionals of the trade and this level is beyond the ability of beginners. LLMs are really powerful tools here. But if you are trying for 90% LLMs are always trying to keep you down. And if you are trying for 100%, new, fringe or exotic LLMs are a disaster because they do not learn and do not understand, even while being inside the token window. We learn that knowledge, (power) and language proficiency are an indicator for crystalline but not fluid intelligence
- otabdeveloper4 1y ago> yanking the needle to 80% 80 percent of what, exactly? A software developer's job isn't to write code, it's understanding poorly-specified requirements. LLMs do nothing for that unless your requirements are already public on Stackoverflow and Github. (And in that case, do you really need an LLM to copy-paste for you?)
- zer00eyz 1y ago> fluid intelligence How about basic intelligence. Kids logic puzzles. https://daydreampuzzles.com/logic-puzzles/ https://daydreampuzzles.com/logic-puzzles/ LLM's whiffing hard on these sorts of puzzles is just amusing. It gets even better if you change the clues from innocent things like "driving tests" or "day care pickup" to things that it doesn't really want to speak about. War crimes, suicide, dictators and so on. Or just flat out make up words whole cloth to use as "activates" in the puzzles.
- motorest 1y ago> They'll never be fit for purpose. They're a technological dead-end for anything like what people are usually throwing them at, IMO. This comment is detached from reality. LLMs in general have been proven to be effective at even creating complete, fully working and fully featured projects from scratch. You need to provide the necessary context and use popular technologies with enough corpus to allow the LLM to know what to do. If one-shot approaches fail, a few iterations are all it takes to bridge the gap. I know that to be a fact because I do it on a daily basis.
- otabdeveloper4 1y ago> because I do it on a daily basis Cool. How many "complete, fully working" products have you released? Must be in the hundreds now, right?
- motorest 1y ago> Cool. How many "complete, fully working" products have you released? Fully featured? One, so far. I also worked on small backing services, and a GUI application to visualize the data provided by a backing service. I lost count of the number of API testing projects I vibe-coded. I have a few instruction files that help me vibecode API test suites from the OpenAPI specs. Postman collections work even better. And I'm far from an expert in the field. What point were you trying to make?
- otabdeveloper4 1y ago> What point were you trying to make? The point is that software developers can't evaluate their own work. (Especially the kind of n00b developers that use LLMs.) You initially made wild claims about insane productivity gains that turned out to be just one small product and a lot of wasted time under scrutiny. (Asking LLMs to write tests is a waste of time. LLMs can't evaluate risks, which is the only reason to write tests in the first place.)
- jeltz 1y agoIf you are far from an expert in the field maybe you should refrain from commenting so strongly because some people here actually are experts. So you have built a few small PoCs, does not tell us much.
- abrookewood 1y agoClearly the issue is that you are going from Erlang/Elixir to Java, rather than the other way around :) Jokes aside, they are pretty different languages. I imagine you'd have much better luck going from .Net to Java.
- nine_k 1y agoThis mostly means that LLMs are good at simpler forms of pattern matching, and have much harder time actually reasoning at a significant depth. (It's not easy even for human intellect, the finest we currently have.)
- tsimionescu 1y agoSure, it's easier to solve an easier problem, news at eleven. In particular, translating from C# to Java could probably be automated with some 90% accuracy using a decent sized bash script.
- mattmanser 1y agoI once redid a project from VB.Net to C# and pretty much did that. People misjudge many tasks as 'hard' when they are in fact easy but tedious. The problem is you need a high degree of accuracy, which you don't get with LLMs. The best you can do is set the LLM on a loop and try and Brute force it, which is the current vibe 'coding' trick. I sound pessimistic but I'm actually shocked at how effective it is.
- h4ck_th3_pl4n3t 1y agoI just wished the LLM model providers would realize this and instead would provide specialized LLMs for each programming language. The results likely would be better.
- chuckadams 1y agoThe local models JetBrains IDEs use for completion are specialized per-language. For more general problems, I’m not sure over-fitting to a single language is any better for a LLM than it is for a human.
- conception 1y agoI’m curious what your process was. If you just said “rewrite this in Java” I’d expect that to fail. If you treated the llm like a junior developer or an official project, worked with them to document the codebase, come up with a plan, tasks for each part of the code base and a solid workflow prompt- I would expect it to succeed.
- Marazan 1y agoYes, if you do all the difficult time consuming bits I bet it would work.
- shard972 1y ago[dead]
- FeepingCreature 1y agoIt's still a lot faster than doing it yourself ime. (Yes I've seen the study, it doesn't account for motivation.)
- conception 1y agoYou don’t do the work you just work with them to get the work done to plan it out.
- 4hg4ufxhy 1y agoThere is a reason to go the extra mile for juniors. They eventually learn and become seniors. With AI I'd rather just do it myself and be done with it.
- conception 1y agoBut you can just do it once with AI. It’s just a script process that you would set up for any project. It’s just an on boarding process. And I’ll know by when I say do it once I mean, obviously processes have to be it on to get exactly what you want out of them, but that’s just how process works . Once it’s working the way you want to just reuse it.
- nerdsniper 1y agoClaude Code / 4o struggle with this for me, but I had Claude Opus 4 rewrite a 2,500 line powershell script for embedded automation into Python and it did a pretty solid job. A few bugs, but cheaper models were able to clean those up. I still haven't found a great solution for general refactoring -- like I'd love to split it out into multiple Python modules but I rarely like how it decides to do that without me telling it specifically how to structure the modules.
- credit_guy 1y agoIf you try to ride a bicycle, do you expect to succeed at the first try? Getting AI code assistants to help you write high quality code takes time. Little by little you start having a feel for what prompts work, what don't, what type of tasks the LLMs are likely to perform well, which ones are likely to result in hallucinations. It's a learning curve. A lot of people try once or twice, get bad results, and conclude that LLMs are useless. But few people conclude that bicycles are useless if they can't ride them after trying once or twice.