4 ms·
LLMs still seem pretty bad at writing large scale software like web browsers. Though it’s probably a problem of managing large context windows more than anythin
by josephg 15d ago
LLMs still seem pretty bad at writing large scale software like web browsers. Though it’s probably a problem of managing large context windows more than anything. Not sure if larger models will magically overcome that.
- cyanydeez 15d agoThe LLM by itself will never create software of any nontrivial (training) complexity. The harness though will improve while parameter count stagnates. The Qwen3.8 models are strong enough when given proper context.
- josephg 15d ago> The LLM by itself will never create software of any nontrivial (training) complexity. Huh? I'm not sure what the word "training" does in that sentence. But "never" my arse. Frontier models can make nontrivial software already. For example, the other day I asked fable to reverse engineer the satisfactory blueprint file format. Then write a program to read the logistic flow graph in a blueprint. Then make an auditing tool that can analyse the graph to find problems. Well, it totally knocked it out of the park: https://seph.au/blueprints/#bp=0%3Aalumina.sbp https://seph.au/blueprints/#bp=0%3Aalumina.sbp This is a relatively small program, but it's not trivial. I'd consider a trivial program to be something I could code up in 20 minutes. It would have taken me a couple weeks to make this blueprint auditing tool, including reverse engineering the file format, writing the analysis code, making the website, scraping all the in-game data on available recipes and icons and so on. I've got a lot of mixed feelings about LLMs. But it seems very silly to lie about what they're capable of.
- cyanydeez 15d agotraining is in there because it's "trained" to do trivial apps like TODO lists, etc. I'm well aware it can build apps. But if you arn't tracking what's going on, they're basically creating deterministic gates and tools to get it to do anything. There's no lie here, it's simply about what you think is _LLM_ and what is the rest of the software that's making it go. I use opencode consistently to build non-trivial apps with it, but it's not doing it with zero guidance, and it's routinely wrong about it's assumptions, and the rest. The thing keeping it on track is opencode, not the LLM's training.
- josephg 14d ago> But if you arn't tracking what's going on, they're basically creating deterministic gates and tools to get it to do anything. So, the same as skilled humans then?
- ohyes 14d agoThat’s definitely how I like to ensure my code isn’t garbage. But now I use the tools I would have made to make the LLM & harness more useful and I might make more with the LLM because it costs me slightly less mentally.