7 ms·
This is the smoothest tom sawyer move I've ever seen IRL, I wonder how many people are now grinding out your GTK4 port with our favorite LLM/system to see if it
by snickell 2y ago
This is the smoothest tom sawyer move I've ever seen IRL, I wonder how many people are now grinding out your GTK4 port with our favorite LLM/system to see if it can. It'll be interesting to see if anyone gets something working with current-gen LLMs.
UPDATE: naive (just fed it your description verbatim) cline + claude 3.7 was a total wipeout. It looked like it was making progress, then freaked out, deleted 3/4 of its port, and never recovered.
- SV_BubbleTime 2y agoSmooth? Nah. Tom Sawyer? Yes.
- phkahler 2y ago>> This is the smoothest tom sawyer move I've ever seen IRL That made me laugh. True, but not really the motivation. I honestly don't think LLMs can code significant real-world things yet and I'm not sure how else to prove that since they can code some interesting things. All the talk about putting programmers out of work has me calling BS but also thinking "show me". This task seems like a good combination of simple requirements, not much documentation, real world existing problem, non-trivial code size, limited scope.
- snickell 2y agoYes, very much agree, an interesting benchmark. Particularly because it’s in a “tier 2” framework (gtkmm) in terms of amount of code available to train an LLM on. That tests the LLMs ability to plan and problem solve compared with, say, “convert to the latest version of react” where the LLM has access to tens of thousands (more?) of similar ports in its training dataset and more has to pattern match.
- phkahler 2y ago>> Particularly because it’s in a “tier 2” framework (gtkmm) in terms of amount of code available to train an LLM on. I asked GPT4 to write an empty GTK4 app in C++. I asked for a menu bar with File, Edit, View at the top and two GL drawing areas separated by a spacer. It produced what looked like usable code with a couple lines I suspected were out of place. I did not try to compile it so don't know if it was a hallucination, but it did seem to know about gtkmm 4.
- snickell 2y agoIt definitely knows what GTK4 is, when it freaked out on me and lost the code, it was using all gtkmm-4.0 headers, and had the compiler error count down to 10 (most likely with tons of logic errors, but who knows). But LLMs performance varies (and this is a huge critique!) not just on what they theoretically know, but how, erm, cross-linked it is with everything else, and that requires lots of training data in the topic. Metaphorically, I think this is a little like the difference for humans in math between being able to list+define techniques to solve integrals vs being able to fluidly apply them without error. I think a big and very valid critique of LLMs (compared to humans) is that they are stronger at "memory" than reasoning. They use their vast memory as a crutch to hide the weaknesses in their reasoning. This makes benchmarks like "convert from gtkmm3 to gtkmm4" both challenging AND very good benchmarks of what real programmers are able to do. I suspect if we gave it a similarly sized 2kloc conversion problem with a popular web framework in TS or JS, it would one-shot it. But again, its "cheating" to do this, its leveraging having read a zillion conversion by humans and what they did.
- cluckindan 2y agoI agree. I tried something similar: a conversion of a simple PHP library from one system to another. It was only like 500 loc but Gemini 2.5 completely failed around line 300, and even then its output contained straight up hallucinations, half-brained additions, wrong namespaces for dependencies, badly indented code and other PSR style violations. Worse, it also changed working code and broke it.
- blensor 2y agoDid you paste it into the chat or did you use it with a coding agent like Cline? I am majorly impressed with the combination VSCode + Cline + Gemini Today I had it duplicate an esp32 proram from UDP communication to TCP. It first copied the file ( funnily enough by writing it again instead of just straight cp ) Then it started to just change all the headers and declarations Then in a third step it changed one bigger function And in the last step it changed some smaller functions And it reasoned exactly that way "Let's start with this first ... Let's now do this .... " until is was done
- ionwake 2y agoI’ve just moved from expensive claudecode to cursor and Gemini - what are you thoughts on cursor vs cline? Thank you
- stavros 2y agoTry asking it to generate a high-level plan of how it's going to do the conversion first, then to generate function definitions for the new functions, then have it generate tests for the new functions, then actually write them, while giving it the output of the tests. It's not like people just one-shot a whole module of code, why would LLMs?
- SpaceNoodled 2y agoOnly 500 lines? That's miniscule.
- semi-extrinsic 2y ago
- nico 2y ago> I honestly don't think LLMs can code significant real-world things yet and I'm not sure how else to prove that since they can code some interesting things In my experience it seems like it depends on what they’ve been trained on They can do some pretty amazing stuff in python, but fail even at the most basic things in arm64 assembly These models have probably not seen a lot of GTK3/4 code and maybe not even a single example of porting between the two versions I wonder if finetuning could help with that
- ksec 2y ago>All the talk about putting programmers out of work I keep thinking may be specifically Web programmers. Given a lot of the web essentially CRUD / have the same function.
- codebra 1y agoProgrammers who code interesting things likely shouldn’t worry. The legions who code voluminous but shallow corporate apps and glue might be more concerned.