Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mrothroc
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
mrothroc
6mo ago
Non-coders often think all engineers do is write code. They don't realize how much more hours are spent on making sure the code we write is correct, from many angles. Functional bugs? Easy to maintain? Cost optimized? Meets user expect
32.
▲
by
mrothroc
6mo ago
I see this has been updated by the user showing it is their own tool doing the damage. These things happen. They happened before coding agents, they happen now. I've done plenty of damage with my own ten fingers on the keyboard without
33.
▲
by
mrothroc
7mo ago
It's the whole tool that's important, not so much how you get screenshots. That's what I'm saying: this is headed in the right direction, it just falls a little short of what I do, where I get tons of value over and abov
34.
▲
by
mrothroc
7mo ago
Everyone is comparing this to Playwright but it's solving a different problem. Playwright checks structural properties, like does element X exist, is it visible, etc. That's useful but it can't tell you whether the page actua
35.
▲
by
mrothroc
7mo ago
I do this across several different codebases, all event-driven microservices. Some greenfield some brownfield. My strategy is to have a central spec, typically protobuf or openapi, and every service has a make target to generate code from t
36.
▲
by
mrothroc
7mo ago
Everyone keeps saying 80/20 but that undersells what's going on. The last 20% isn't just hard. It's hard because of what happened during the first 80%. When an agent takes a shortcut early on, the next step doesn't
37.
▲
by
mrothroc
7mo ago
I use prompt templates, so in the first version of my analysis script on my own logs I looked for those. However, to make it generic, I switched to using gemini as a classifier. That's what's in the repo.
38.
▲
by
mrothroc
7mo ago
Nice, I've been working on the same problem from a different direction. Instead of analyzing sessions after the fact, I built a pipeline that structures them. Stages (plan, design, code, review, same as you'd have with humans) wit
39.
▲
by
mrothroc
7mo ago
Senior review can definitely help, regardless if the code comes from a junior or an LLM. We've done this since the dawn of time. However, it doesn't scale and since LLM volume far exceeds what juniors can do, you end up overwhelmi
40.
▲
by
mrothroc
7mo ago
The disposition problem you describe maps to something I keep running into. I've been running fully autonomous software development agents in my own harness and there's real tension between "check everything" and "a
41.
▲
by
mrothroc
7mo ago
Everyone is circling around this. We are shifting to "code factories" that take user intent in at one end and crank out code at the other end. The big question: can you trust it? We're building our tooling around it (thanks,
42.
▲
by
mrothroc
7mo ago
Yeah, this is what happens when there's nothing between "the agent decided to do this" and "it happened." The agent followed the state file logically. It wasn't wrong. It just wasn't checked. His post-mort
43.
▲
by
mrothroc
7mo ago
I've been running a multi-agent setup for quite a while to do software development. I set up a workflow with agents at each stage, spec->plan->design->code->review. The key thing I learned was that the arrangement of the ch
44.
▲
by
mrothroc
7mo ago
I ended up building my own for this. SQLite backend, breaks work into stages with gates between them. A gate checks each handoff before the next stage starts. Does the code match the spec, did the tests pass, that kind of thing. I've b
45.
▲
by
mrothroc
7mo ago
The checkpoint pattern you describe is exactly right. I've been dealing with this as well. Instead of vibe coding, it's vibe system engineering and I don't care for it. So I thought about it and came up with a framework to de
46.
▲
by
mrothroc
7mo ago
I'm old enough to remember that engineers researching distributed systems had the same challenge. Everyone was trying to build 100% reliable nodes, which is impossible. Then Lamport came along and showed you could actually achieve your
47.
▲
by
mrothroc
7mo ago
My experience is similar to yours: LLMs can write excellent code, though you really have to drive them the right way. I use a harness to drive long-run autonomous agents to create production code. (Not open source, but it is an actual produ