Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dhorthy
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
dhorthy
2mo ago
i did a write up on fable while it was out - it can do big refactors, but it does not know what to change without human steering. For that, you need humans to know what to ask for. lights off is still out for me https://x.com
32.
▲
by
dhorthy
2mo ago
its the optimizing utilization instead of overall throughput all over again. eli goldratt talked about this in the 1970s[1]. we still haven't learned 1 - https://en.wikipedia.org/wiki/The_Goal_(novel)
33.
▲
by
dhorthy
2mo ago
exactly. a great pr is a joy to review. we've found some success in agents generating static HTML walkthroughs that order the diffs in something other than GitHub's default alphanumeric ordering, but it can only go so far
34.
▲
by
dhorthy
2mo ago
i guess to clarify my contention: generic prompting like "review this code" or "make the architecture better" will raise the floor but cannot come close to human-quality code without humans understanding what code exists
35.
▲
by
dhorthy
2mo ago
(op here btw) - the question isn't "can models make code better" - its "left fully unattended, will they turn your codebase to slop over time" from the footnotes (sorry if this got a little buried) > yes of cours
36.
▲
by
dhorthy
2mo ago
the thesis of the post is that this is not true. fable can solve more problems but for complex systems it is not sufficiently diligent that you can turn the lights off.
37.
▲
by
dhorthy
2mo ago
that's probably part of it but every enterprise in the world (even teams as small as 10-20 engineers) are paying per token, not with subscriptions. Claude code did make it into those orgs because people played with it at home with subs
38.
▲
by
dhorthy
2mo ago
awesome - i have updated the post with a link to this thread!
39.
▲
by
dhorthy
2mo ago
yeah my best articulation of taste is something i got from Jake Nations[1] while he was still at netflix - "you know a bad pattern when you see it because at some point you were up at 2am debugging it" taste is the hard-earned int
40.
▲
by
dhorthy
3mo ago
we out here trying
41.
▲
by
dhorthy
3mo ago
it is absolutely wild to me that this keeps floating to top comment thread when the guy very clearly did not read even the first half
42.
▲
by
dhorthy
3mo ago
> I spent a week or so and like a billion+ tokens trying to refactor and save it. It just wasn't worth it. this is exactly what happened to github.com/humanlayer/humanlayer - it was overslopped and we reset from scratch to
43.
▲
by
dhorthy
3mo ago
one thing I probably didn't mention is we do the program design having already done an in-depth codebase research, with current patterns and architecture surfaced - that actually seeds every step of the flow including even the product
44.
▲
by
dhorthy
3mo ago
and build it incrementally! You don't have to build the entire software factory at once. You don't have to mastermind the whole future system, instead you're actually stacking and layering these small, isolated problems. I th
45.
▲
by
dhorthy
3mo ago
I have had the phrase "programming is building a theory" spinning in my head for days, especially since watching this pragmatic engineer pod with Kent Beck https://www.youtube.com/watch?v=ddHQQtjIOpw
46.
▲
by
dhorthy
3mo ago
> Historically I know that the majority maintenance problems occur from slow continuous evolution of a system that it initially was never designed for. And the only way to address this was continuous system design. yes exactly - this is
47.
▲
by
dhorthy
3mo ago
appreciate that context! I definitely did not mean to come out and say "its definitely not working" or anything, but would love to hear from y'all a retrospective on the ~5-6 month anniversary - what was right, what did we ge
48.
▲
by
dhorthy
3mo ago
this is 100% right. you have to guard the codebase patterns with your life. because the codebase is part of the prompt.
49.
▲
by
dhorthy
3mo ago
> In order for coding with LLMs to go well, there has to be more rigor, more discipline, more good engineering hard-assedness. To reiterate, the teams seeing the best results with AI were already high-discipline and high-hygiene. hard ag
50.
▲
by
dhorthy
3mo ago
yes exactly
51.
▲
by
dhorthy
3mo ago
i can't tell if this is a compliment or not
52.
▲
by
dhorthy
3mo ago
yeah i like this and others in the thread mentioned that understanding RL and RLHF and the shape of the data is really important (at least the fundamentals, I'm sure there's quite complex industrialization of RL inside labs as Nat
53.
▲
by
dhorthy
3mo ago
blake smith has a really good post on this - that mental alignment among the team is the primary purpose of code review - https://blakesmith.me/2015/02/09/code-review-essentials-for-...
54.
▲
by
dhorthy
3mo ago
yeah that was another thing i hoped would pour through here - that deterministic systems are much better for evaluating quality (test, linters, cyclomatic complexity, etc) - but that we don't have such a system for code maintainability
55.
▲
by
dhorthy
3mo ago
fair point, this is the thing I struggled most to extract out while writing it - if you can propose an RL environment that penalizes a model for bad design, then I'm all ears - right now there's no fast oracle/verifier for th
56.
▲
by
dhorthy
3mo ago
yeah I 100% agree - and I think the most popular coding agent workflows / skill kits are designed to pull those insights and intuition out of humans in a way that optimizes for the developer's experience building the plans or bu
57.
▲
by
dhorthy
3mo ago
normative specifications can help, but the thesis here is that specs that define behavior of the product or even architecture are helpful but there's MORE that can be done and even though "program design" feels too in the wee
58.
▲
by
dhorthy
3mo ago
interesting - i'd say my main goal is to put the current "agentic software factory" hype in the historical context of "we've actually been rube-goldberging software deploys for a while now"
59.
▲
by
dhorthy
3mo ago
yeah this is along the lines of what some friends of mine call "core vs. pragmatic modules" or even s/modules/codebase zones/ the idea that if you have a solid core and decoupled modules, you can have "zones&qu
60.
▲
by
dhorthy
3mo ago
my perhaps controversial take is that opus 4.1 was smarter than 4.5 for complex engineering work, but 4.5 was faster and "squishier" - it responded better to simpler prompts, it read between the lines of user input better, and tha
More ›