4 ms·
There's a reason I didn't: the article itself did not bother. They just put things out there, "thought-leader to thought-leader just like that", to quote a rece
by perching_aix 14d ago
There's a reason I didn't: the article itself did not bother. They just put things out there, "thought-leader to thought-leader just like that", to quote a recent HN comment I found hysterical in an unrelated thread. So to elaborate felt like an isolated demand for rigor on my part.
Nevertheless, point by point then:
1. price competitiveness and autonomy:
They falsely assert that "current frontier models need laborious oversight and guardrails on even the simplest tasks" - this is simply blatantly untrue. We have countless workloads at work that are easily doable for an agent, and they're not really trivial: bootstrapping entire new environments, migrating existing ones, investigating incidents and crafting RCAs, and so on. I experience this time and time again, and have been for half a year now. There are lots of colleagues of mine who only do such tasks. They do a worse job at them (sometimes way more), and take longer to do so. Using guardrails is basically unneeded, and the oversight burden is also far from laborious.
I also work with those "crappy engineers" they speak of, holding titles inflated way beyond "cracked interns", and I wish it was only a bottom quartile. Haiku outperforms them; they need laborious handholding. Handholding that they basically do not comprehend, because their grasp of the English language matches their general technical expertise (i.e. you could hire a random guy off the street and they'd be better). Mind you, I do also keep running the numbers, and they're already not worth it regardless of how kindly I measure. This entire section was just blatantly wrong, certainly at least in my sector (cloud operations and devops). A care for the anecdotal and the subjective, that the author generally doesn't ever give a hoot about, at least in their article.
2. jagged intelligence / the models are idiot savants
Yes, they are. Their intelligence rises with the amount of parameters, and remains mostly local to their "area of expertise", so they clearly scale that way. This has been discussed to hell and back, and should be more than familiar enough to anyone reading this forum; it's trite.
3. specification burden
It's the exact same as with crappy people. Literally the exact same. If you find yourself specifying things so hard, the model is simply not good enough for that workload. You can indeed absolutely dig yourself into a hole and make things not worth it, this is also extremely trite. It's been an adage with automating anything since forever. "You spend 5 minutes doing something that annoys you, or spend 30 minutes automating it." Unless you never heard this adage before somehow, this won't be new.
4. same as the previous one
5. navier-stokes and proof hacking
All of these caveats are well understood by the relevant community, and are well accessible to those still within tech but outside of math, too: https://news.ycombinator.com/item?id=49672339 https://news.ycombinator.com/item?id=49672339
Separately, I think there's a quite a bit bigger issue with this whole thing, that is going to be a lot more salient angle in this context, specifically given the discussion of price competitiveness: https://news.ycombinator.com/item?id=49737622 https://news.ycombinator.com/item?id=49737622
6. human review
It is not an alternative, it is inescapable. See my previous link in the first passage of point five. We're already a bottleneck, which is again obvious to anyone who uses these things on the daily. Can kick the rocks down the road as much as we want, an outsourced concern is outsourced.
7. (they start counting from 1. again) cheap iteration is key
Yes, just like with people. And again, this is just readily apparent if you ever tried to guardrail an agent. It's an uphill battle. More cheaper failures basically always beat fewer expensive failures, this is nothing new, and nothing necessarily specific to this or unintuitive.
It applies to people for different reasons (unfamiliarity is already anxiety inducing, a potential to do severe damage is super not helping that). Gotta let people experiment for them to actually learn and be productive. Either way, the underlying support methodology is the same: de-impact the risk first, de-risk the thing second.
8. repetitive work is more easily performed
Yes... just like with people. That's how the whole manufacturing line and the various specialized stations came to be. Businesses are built around continuously ossifying their processes into more standard, more textbook ones; they're like JIT compilers. This is exactly why agents are eating the bottom layers, and why people here keep saying "the bar is rising".
9. doesn't check out cost wise, what does is already code-automated
No, there is absolutely a space between the unautomated and the code-automated where these agents slot in fine, because point 1 was false.
That's all their points, the rest is just a conclusion. These are very basic experiences anyone can identify, littered with some tropes everyone sees here every day, delivered in a quirky format that people are hyping the snot out of. It's also full opinion, just like this comment: that is to say, it literally could have been just a random comment in a thread for me to scroll past.
I'm hating on this thing big time, but it's not even that I disagree with it for the most part, it's everything else. The style is obnoxious, it's basically just the lukewarm opinion of some random dude, asserting itself, upvoted to high heavens. So to now be presented with the notion that this is some legendary piece of writing is just kind of painful. I appreciate if you found value in it, that's great, there are parts of it that are fairly impossible to argue and will be actually informative if you were unfamiliar with them (the parts that for me registered as trite instead). Just worried that since it's all presented tonally the same, including the parts not like that, it will cause a lot of people to repeat those other parts too, uncritically. Being told it is essentially some absolute must read piece of writing is already not looking good in that regard.