6 ms·
Which also fits with how it performs at software engineering (in my experience). Great at boilerplate code, tests, simple tutorials, common puzzles but bad at n
by fire_lake 1y ago
Which also fits with how it performs at software engineering (in my experience). Great at boilerplate code, tests, simple tutorials, common puzzles but bad at novel and complex things.
- brundolf 1y agoYep. But wonderful at aggregating details from twelve different man pages to write a shell script I didn't even know was possible to write using the system utils
- fundingshovel 1y agoI use it for this a lot.
- genewitch 1y ago[flagged]
- HenryBemis 1y agoIs it 'only' "aggregating details from twelve different man pages" or has it 'studied' (scraped) all (accessible) code in GitHub/GitLab/Stachexchange/etc. and any other publicly available coding repositories on the web (and for the case of MS the Git it owns)? Together with descriptions of what is right and what is wrong.. I use it for code, and I only do fine tuning. When I want something that is clearly never done before, I 'talk' to it and train it on which method to use, and for a human brain some suggestions/instructions are clearly obvious (use an Integer and not a Double, or use Color not Weight). So I do 'teach' it as well when I use it. Now, I imagine that when 1 million people use LLMs to write code and fine tune it (the code), then we are inherently training the LLMs on how to write even better code. So it's not just "..different man pages.." but "the finest coding brains (excluding mine) to tweak and train it".
- jdiff 1y agoDefinitely matches my experience as well. I've been working away on a very quirky, non-idiomatic 3D codebase, and LLMs are a mixed bag there. Y is down, there's no perspective distortion or Z buffer, there are no meshes, it's a weird place. It's still useful to save me from writing 12 variations of x1 = sin(r2) - cos(r1) while implementing some geometric formula, but absolutely awful at understanding how those fit into a deeply atypical environment. Also have to put blinders on it. Giving it too much context just throws it back in that typical 3D rut and has it trying to slip in perspective distortion again.
- westmeal 1y agoI gotta ask what are you actually doing because it sure sounds funky
- jdiff 1y agoWorking on extending the [Zdog](https://zzz.dog https://zzz.dog) library, adding some new types and tooling, patching bugs I run into on the way. All the quirks inherit from it being based on (and rendering to) SVG. SVG is Y-down, Zdog only adds Z-forward. SVG only has layering, so Zdog only z-sorts shapes as wholes. Perspective distortion needs more than dead-simple affine transforms to properly render beziers, so Zdog doesn't bother. The thing that really throws LLMs is the rendering. Parallel projection allows for optical 2D treachery, and Zdog makes heavy use of it. Spheres are rendered as simple 2D circles, a torus can be replicated with a stroked ellipse, a cylinder is just two ellipses and a line with a stroke width of $radius. LLMs struggle to even make small tweaks to existing objects/renderers.
- josephg 1y agoYeah I have the same experience. I’ve done some work on novel realtime text collaboration algorithms. For optimisation, I use some somewhat bespoke data structures. (Eg I’m using an order-statistic tree storing substring lengths with internal run-length encoding in the leaf nodes). ChatGPT is pretty useless with this kind of code. I got it to help translate a run length encoded b-tree from rust to typescript. Even with a reference, it still introduced a bunch of new bugs. Some were very subtle.
- jitl 1y agoIt’s just not there yet but I think it will get there for translation kind of tasks quite capably in the next 12 months, especially if asked to translate a single file or a selection in a file line by line. Right now it’s quite bad which I find surprising. I have less confidence we’ll see whole-codebase or even module level understanding for novel topics in the next 24 months. There’s also a question of quality of source data. At least in TypeScript/JavaScript land, the vast majority of code appears to be low quality and buggy or ignores important edge cases and so even when working on “boilerplate” it can produce code that appears to work but will fall over in production for 20% of users (for example string handling code that will tear Unicode graphemes like emoji).
- imatworkyo 1y agohow often are we truly writing actual novel programs that are complex in a way AI does not excel at? There are many types of complex, and many times complex for a human coder, are trivial for AI and its skillset.
- gf000 1y agoDepends on the field of development you do. CRUD backend app for a business in a common sector? It's mostly just connecting stuff together (though I would argue that an experienced dev with a good stack takes less time to write it as is than painstakingly explaining it to an LLM in an inexact human language). Some R&D stuff, or even debugging any kind of code? It's almost useless, as it would require deep reasoning, where these models absolutely break down.
- simonw 1y agoHave you tried debugging using the new "reasoning" models yet? I have been extremely impressed with o1, o3, o4-mini and Gemini 2.5 as debugging aids. The combination of long context input and their chain-of-thought means they can frequently help me figure out bugs that span several different layers of code. I wrote about an early experiment with that here: https://simonwillison.net/2024/Sep/25/o1-preview-llm/ https://simonwillison.net/2024/Sep/25/o1-preview-llm/ Here's a Gemini 2.5 Pro transcript from this afternoon where I'm trying to figure out a very tricky bug: https://gist.github.com/simonw/4e208ab9edb5e6a814d3d23d7570db58 https://gist.github.com/simonw/4e208ab9edb5e6a814d3d23d7570d...
- bla3 1y agoIn my experience they're not great with mathy code for example. I had a function that did subdivision of certain splines and had some of the coefficients wrong. I pasted my function into these reasoning models and asked "does this look right?" and they all had a whole bunch of math formulas in their reasoning and said "this is correct" (which it wasn't).
- expensive_news 1y ago
- jeswin 1y ago> novel and complex things a) What's an example? b) Is 90% (or more) of programming mundane, and not really novel?
- deleted 1y ago[deleted]
- nurettin 1y agoIf you'd like a creative waste of time, make it implement any novel algorithm that mixes the idea of X with Y. It will fail miserably, double down on the failure and hard troll you, run out of context and leave you questioning why you even pay for this thing. And it is not something that can be fixed with more specific training.
- AlexCoventry 1y agoCan you give an example? Have you tried it recently with the higher-end models?
- nurettin 1y agoMy favorite example is implementing NEAT with keras dense layers instead of graphs. Last time I tried with claude 3.7, it wrote code to mutate the output layer (??). I tried to prevent that a few times and gave up.
- AlexCoventry 1y agoThis NEAT? https://web.archive.org/web/20231205130538/http://www.cs.ucf.edu/~kstanley/neat.html https://web.archive.org/web/20231205130538/http://www.cs.ucf... Is the idea to use a keras dense layer to represent a weighted graph by identifying the input nodes with the corresponding outputs?
- nurettin 1y agoThe idea is to evolve the multi layer dnn using ga
- spaceman_2020 1y agoThis is also why I buy the apocalyptic headlines about AI replacing white collar labor - most white collar employment is mostly creating the same things (a CRUD app, a landing page, a business plan) with a few custom changes Not a lot of labor is actually engaged in creating novel things. The marketing plan for your small business is going to be the same as the marketing plan for every other small business with some changes based on your current situation. There’s no “novel” element in 95% of cases.
- econ 1y agoI wonder what the impact will be when replicating the same thing becomes machine readable with near 100% accuracy.
- coffeebeqn 1y agoI don’t know if most software engineers build toy CRUD apps all day? I have found the state of the art models to be almost completely useless in a real large codebase. Tried Claude and Gemini latest since the company provides them but they couldn’t even write tests that pass after over a day of trying
- fl0id 1y agoSame. Like Claude code for example will write some tests. But what they are testing is often incorrect
- the_duke 1y agoAgreed in general, the models are getting pretty good at dumping out new code, but for maintaining or augmenting existing code produces pretty bad results, except for short local autocomplete. BUT it's noteworthy that how much context the models get makes a huge difference. Feeding in a lot of the existing code in the input improves the results significantly.
- nradov 1y agoThis might be an argument in favor of a microservices architecture with the code split across many repos rather than a monolithic application with all the code in a single repo. It's not that microservices are necessarily technically better but they could allow you to get more leverage out of LLMs due to context window limitations.