3 ms·
If you said this in 2025 I would've 100% understood, but to be honest getting AI models to do a pretty good job on day-to-day ticket work has become so boring t
by jchw 2mo ago
If you said this in 2025 I would've 100% understood, but to be honest getting AI models to do a pretty good job on day-to-day ticket work has become so boring that we don't even bother using the top tier models and higher effort slots for that anymore. I personally wind up tweaking the results a lot and recursively having fresh agents review the diff, but that's just because I'm picky; in a lot of cases the first diff is actually pretty damn decent.
Compared to what I am doing at home experimentally, I feel like day-to-day work is absolutely nothing. Not only am I also working with existing codebases in my experimental prototyping, but I am also doing things vastly more complex with vastly harder constraints.
- com2kid 2mo agoAll non-trivial code terra has generated for me has had at least one serious bug in it. Typically caught by a review from myself or Sol. But I wouldn't trust lower tier models for end to end solutions.
- jchw 2mo agoPersonally I wouldn't want bots running autonomously on a repo, even if there were other bots cross checking them. But that having been said, I'd also say that my experience was similar with human code: it is rare to not find at least some issue worth at least pointing out. The only real difference is that the LLMs have vastly different holes than people do, making the real challenge trying to make sure you're covering them. In most cases for me the best solution seems to be just giving them a way to test and attempt to prove things out in a realistic environment. But clearly, we haven't really left the era of having humans in the loop. Fully vibe-coded codebases clearly suffer from a myriad of issues.
- com2kid 2mo agoMultiple security holes. Issues marked fixed that aren't really. Giant holes left in solutions. Sol over engineers now and then (hey please don't factor that function out into its own file....) but it doesn't do the same level of stupid terra does. That said, plan with Sol, implement with terra, have Sol fix all the mistakes, then I go over the code and make recommendations for the architecture to fix Sol's foolishness.
- hatthew 2mo agoMy experience is that current models are pretty good at making functional changes without too many more bugs than a human would make, but are still bad at making good high quality changes. I have a degree in software engineering specifically, so perhaps I am overly sensitive to design issues.