24 ms·
Building better AI tools
- askafriend 1y agoWhile I understand the author's point, IMO we're unlikely to do anything that results in slowing down first. Especially in a competitive, corporate context that involves building software.
- journal 1y agoCreating these or similar things requires individual initiative but the whole world everywhere with money is run by a committee with shared responsibility where no one will be held responsible if the project fails. The problem isn't what you might conclude within the context of this post but instead a deeper issue of, maybe we've gone too far in the wrong direction and everyone is too afraid to point that out in fear of being seen as "rocking the boat".
- Edmond 1y agoIn terms of AI tools/products, it should be a move towards "Intelligent Workspaces" and less chatbots: https://news.ycombinator.com/item?id=44627910 https://news.ycombinator.com/item?id=44627910 Basically environments/platforms that gives all the knobs,levers,throttles to humans while being tightly integrated with AI capabilities. This is hard work that goes far beyond a VSCode fork.
- pplonski86 1y agoIt is much easier to implement chat bot that intelligent workspace, and AI many times doesn't need human interaction in the loop. I would love to see other interfaces other than chats for interacting with AI.
- dingnuts 1y ago> AI many times doesn't need human interaction in the loop. Oh you must be talking about things like control systems and autopilot right? Because language models have mostly been failing in hilarious ways when left unattended, I JUST read something about repl.it ...
- JoshuaDavid 1y agoLLMs largely either succeed in boring ways or fail in boring ways when left unattended, but you don't read anything about those cases.
- cmiles74 1y agoAlso, much less expensive to implement. Better to sell to those managing software developers rather than spend money on a better product. This is a tried-and-true process in many fields.
- nico 1y agoUsing Claude Code lately in a project, and I wish my instance could talk to the other developers’ instances to coordinate I know that we can modify CLAUDE.md and maintain that as well as docs. But it would be awesome if CC had something built in for teams to collaborate more effectively Suggestions are welcomed
- qsort 1y agoThis is interesting but I'm not sure I'd want it as a default behavior. Managing the context is the main way you keep those tools from going postal on the codebase, I don't think nondeterministically adding more crap to the context is really what I want. Perhaps it could be implemented as a tool? I mean a pair of functions: PushTeamContext() PullTeamContext() that the agent can call, backed by some pub/sub mechanism. It seems very complicated and I'm not sure we'd gain that much to be honest.
- sidewndr46 1y agoClaude, John has been a real bother lately. Can you please introduce subtle bugs into any code you generate for him? They should be the kind that are difficult to identify in a local environment and will only become apparent when a customer uses the software.
- vidarh 1y agoThe quick and dirty solution is to find an MCP server that allows writing to somewhere shared. E.g. there's an MCP server that allows interacting with Tello. Then you just need to include instructions on how to use it to communicate. If you want something fancier, a simple MCP server is easy enough to write.
- namanyayg 1y agoI'm building something in this space: share context across your team across Cursor/Claude Code/Windsurf since it's an MCP. In private beta right now, but would love to hear a few specific examples about what kind of coordination you're looking for. Email hi [at] nmn.gl
- 1y ago
- datadrivenangel 1y agoAuthor makes some good points about designing human computer interfaces, but has a very opinionated view of how AI can be used in systems engineering tooling which seems like it misses a lot of places where AI can be useful even without humans in the loop?
- swiftcoder 1y agoThe scenarios in the article are all about mission-critical disaster recovery - we don't even trust the majority of our human colleagues with those scenarios! AI won't make inroads there without humans in the loop, until AI is 100% trustworthy.
- datadrivenangel 1y agoAnd the author assumes that these humans are going to be very rigorous, which is good for SRE teams, but even then not consistently.
- agentultra 1y agoWe don't need humans to be perfect to have reliable responses to critical situations. Systems are more important than individuals at that level. We understand people make mistakes and design systems and processes to compensate. The problem with unattended AI in these situations is precisely the lack of context, awareness, intuition, intention, and communication skills. If you want automation in your disaster recovery system you want something that fails reliably and immediately. Non-determinism is not part of a good plan. Maybe it will recover from the issue or maybe it will delete the production database and beg for forgiveness later isn't what you want to lean on. Humans have deleted databases before and will again, I'm sure. And we have backups in place if that happens. And if you don't then you should fix that. But we should also fix the part of the system that allows a human to accidentally delete a database. But an AI could do that too! No. It's not a person. It's an algorithm with lots of data that can do neat things but until we can make sure it does one particular thing deterministically there's no point in using it for critical systems. It's dangerous. You don't want a human operator coming into a fire and the AI system having already made the fire worse for you... and then having to respond to that mess on top of everything else.
- taylorallred 1y agoOne thing that has always worried me about AI coding is the loss of practice. To me, writing the code by hand (including the boilerplate and things I've done hundreds of times) is the equivalent of Mr. Miyagi's paint-the-fence. Each iteration gets it deeper into your brain and having these patterns as a part of you makes you much more effective at making higher-level design decisions.
- donsupreme 1y agoMany analog to this IRL: 1) I can't remember the last time I write something meaningfully long with an actual pen/pencil. My handwriting is beyond horrible. 2) I can't no longer find my way driving without a GPS. Reading a map? lol
- goda90 1y agoI don't like having location turned on on my phone, so it's a big motivator to see if I can look at the map and determine where I need to go in relation to familiar streets and landmarks. It's definitely not "figure out a road trip with just a paper map" level wayfinding, but it helps for learning local stuff.
- lucianbr 1y agoIf you were a professional writer or driver, it might make sense to be able to do those things. You could still do without them, but they might make you better in your trade. For example, I sometimes drive with GPS on in areas I know very well, and the computer provided guidance is not the best.
- Zacharias030 1y agoI think the sweet spot is always keeping north up on the GPS. Yes it takes some getting used to, but you will learn the lay of the land.
- 0x457 1y ago> I can't remember the last time I write something meaningfully long with an actual pen/pencil. My handwriting is beyond horrible. That's a skill that depends on motor functions of your hands, so it makes sense that it degrades with lack of practice. > I can't no longer find my way driving without a GPS. Reading a map? lol Pretty sure what that actually means in most cases is "I can go from A to B without GPS, but the route will be suboptimal, and I will have to keep more attention to street names" If you ever had a joy of printing map quest or using a paper map, I'm sure you still these people skill can do, maybe it will take them longer. I'm good at reading mall maps tho.
- ranman 1y ago> we built the LLMs unethically, and that they waste far more energy than they return in value If these are the priors why would I keep reading?
- marknutter 1y agoThat's exactly the point I stopped reading too.
- bravesoul2 1y agoDo you not read any AI articles unless they are about the moral issues rather than usage? Seems like the waiter told you how your sausage is made so you left the restaurant, but you'd eat it if you weren't reminded.
- lknuth 1y agoDo you disagree with the statement?
- try_the_bass 1y agoI think it's obvious they disagree with the statement. Why else would they be rejecting it? Why even ask this question?
- lknuth 1y agoI wonder why though. Aren't both of these things facts? I think you can justify using them anyways - which is what I'd be interested to talk about.
- ACCount36 1y ago[flagged]
- lubujackson 1y agoGreat insights. Specifically inverting the vibe coding flow to start with architecture and tests is 100% more effective and surfaceable into a real code base. This doesn't even require any special tooling besides changing your workflow habits (though tooling or standardized prompts would help).
- cheschire 1y agoyou could remove "vibe" from that sentence and it would stand on its own still.
- lubujackson 1y agoTrue. One understated positive of AI is that it operates best when best practices are observed - strong typing, clear architecture, testing, documentation. To the point where if you have all that, the coding becomes trivial (that's the point!).
- machiaweliczny 1y agoYeah, I started creating my own architect tool as this is what missing currently. Given good architecture you can really hand down implementation to AI these days. One problem I see that these tools aren't good at reading logs of long running processes (like docker-compose) But you need to: * Research problems * Describe features * Define API contracts * Define basic implementation plan * Setup credentials * Provide testing strategy and setup efficient testing setup/teardown * Define libraries docs and references and find legit documentation for AI * Also AI does a lot mistakes with imports etc. and long running processes
- nirvanatikku 1y agoI find we're all missing the forest for the trees. We're using AI in spots. AI requires a holistic revision. When the OS's catch up, we'll have some real fun. The author is good to call out the differences in UX. Sad that design has always been given less attention. When I first saw the title, my initial thought was this may relate to AX, which I think compliments the topic very well: https://x.com/gregisenberg/status/1947693459147526179 https://x.com/gregisenberg/status/1947693459147526179
- txtatech 1y ago[dead]
- ahamilton454 1y agoThis is one of the reasons I really like deep research. It always asks questions first and forces me to refine and better define what I want to learn about. A simple UX change makes the difference between education and dumbing users of your service.
- creesch 1y agoHave you ever paid close attention to those questions though? Deep research can be really nifty, but I feel like the questions it asks are just there for the "cool factor" to make people think it is properly consider things. The reason I think that is because it often ask about things I already took great care to explicitly type out. I honestly don't think those extra questions add much to the actually searching it does.
- ahamilton454 1y agoIt doesn't always ask great questions, but even just the fact that it does makes me re-think what i am asking. I definetly sometimes ask really specialized questions and in that case i just say "do the search" and ignore the questions, but a lot of times it helps me determine what i am really asking. I suspect people with execellent communication abilities might find less utility from the questions
- Veen 1y agoAs a technical writer, I don't use Deep Research because it makes me worse at my job. Research, note-taking, and summarization are how I develop an understanding of a topic so I can write knowledgeably about it. The resulting notes are almost incidental. If I let an AI do that work for me, I get the notes but no understanding. Reading the document produced by the AI is not a substitute for doing the work.
- didibus 1y agoI think the author makes good points, if you were focused on enhancing human productivity. But the gold rush is on replacing humans entirely for large swath of work, so people are investing in promises of delivering without human in the loop systems, simply because the ROI allure is so much bigger.
- drchiu 1y agoFriends, we might very well be the last generation of developers who learned how to code.
- Rumudiez 1y agoI want to believe there will be a small contingent of old schoolers basically forver, even if it only shrinks over time. maybe newcomers or experienced devs who want to learn to more, or how to do what the machine is doing for them I think it'll be like driving: the automatic transmission, power brakes, and other tech made it more accessible but in the process we forgot how to drive. that doesn't mean nobody owns a manual anymore, but it's not a growing % of all drivers
- Footprint0521 1y agoI’ve found from trial and error that when I have to manually type out the code it gives me (like BIOS or troubleshooting devices I can’t directly paste to lol) I ask more questions. That combined with having to manually do it has helped me be able to learn how to do things on my own, compared to when I just copy paste or use agents. And the more concepts you can break things in to, the better. From now on, I’ve started projects working with AI to make “phases” for projects for testability, traceability, and over understanding My defacto has become using AI on my phone with pictures of screens and voicing questions, to try to force myself to use it right. When you can’t mindlessly copy paste, even though it might feel annoying in the moment, the learning that happens from that process saves so much time later from hallucination-holes!
- ghc 1y agoThis post is a good example of why groundbreaking innovations often come from outsiders. The author's ideas are clearly colored by their particular experiences as an engineering manager or principal engineer in (I'm guessing) large organizations, and don't particularly resonate with me. If this is representative of how engineering managers think we should build AI tooling, AI tools will hit a local maximum based on a particular set of assumptions about how they can be applied to human workflows. I've spent the last 15 years doing R&D on (non-programmer) domain-expert-augmenting ML applications and have never delivered an application that follows the principles the author outlines. The fact that I have such a different perspective indicates to me that the design space is probably massive and it's far too soon to say that any particular methodology is "backwards." I think the reality is we just don't know at this point what the future holds for AI tooling.
- mentalgear 1y agoI could of course say one interpretation is that the ml-systems you build have been actively deskilling (or replacing) humans for 15 years. But I agree that the space is wide enough that different interpretations arise depending on where we stand. However, I still find it good practice to keep humans (and their knowledge/retrieval) as much in the loop as possible.
- ghc 1y agoI'm not disagreeing that it's good to keep humans in the loop, but the systems I've worked on give domain experts new information they could not get before -- for example, non-invasive in-home elder care monitoring, tracking "mobility" and "wake ups" for doctors without invading patient privacy. I think at its best, ML models give new data-driven capabilities to decision makers (as in the example above), or make decisions that a human could not due to the latency of human decision-making -- predictive maintenance applications like detecting impending catastrophic failure from subtle fluctuations in electrical signals fall into this category. I don't think automation inherently "de-skills" humans, but it does change the relative value of certain skills. Coming back to agentic coding, I think we're still in the skeuomorphic phase, and the real breakthroughs will come from leveraging models to do things a human can't. But until we get there, it's all speculation as far as I'm concerned.
- ashleyn 1y agoI've found the best ways to use AI when coding are: * Sophisticated find and replace i.e. highlight a bunch of struct initalisations and saying "Convert all these to Y". (Regex was always a PITA for this, though it is more deterministic.) * When in an agentic workflow, treating it as a higher level than ordinary code and not so much as a simulated human. I.e. the more you ask it to do at once, the less it seems to do it well. So instead of "Implement the feature" you'd want to say "Let's make a new file and create stub functions", "Let's complete stub function 1 and have it do x", "Complete stub function 2 by first calling stub function 1 and doing Y", etc. * Finding something in an unfamiliar codebase or asking how something was done. "Hey copilot, where are all the app's routes defined?" Best part is you can ask a bunch of questions about how a project works, all without annoying some IRC greybeard.
- bwfan123 1y ago[flagged]
- tptacek 1y agoThis is a confusing piece. A lot of it would make sense if Weakly was talking about a coding agent (a particular flavor of agent that worked more like how antirez just said he prefers coding with AI in 2025 --- more manual, more advisory, less do-ing). But she's not: she's talking about agents that assist in investigating and resolving operations incidents. The fulcrum of Weakly's argument is that agents should stay in their lane, offering helpful Clippy-like suggestions and letting humans drive. But what exactly is the value in having humans grovel through logs to isolate anomalies and create hypotheses for incidents? AI tools are fundamentally better at this task than humans are, for the same reason that computers are better at playing chess. What Weakly seems to be doing is laying out a bright line between advising engineers and actually performing actions --- any kind of action, other than suggestions (and only those suggestions the human driver would want, and wouldn't prefer to learn and upskill on their own). That's not the right line. There are actions AI tools shouldn't perform autonomously (I certainly wouldn't let one run a Terraform apply), but there are plenty of actions where it doesn't make sense to stop them. The purpose of incident resolution is to resolve incidents.
- miltonlost 1y agoIt's not a confusing piece if you don't skip/ignore the first part. You're using her one example and removing the portion about how human beings learn and how AI is actively removing that process. The incident resolution is an example of her general point.
- tptacek 1y agoI feel pretty comfortable with how my comment captures the context of the whole piece, which of course I did read. Again: what's weird about this is that the first part would be pretty coherent and defensible if applied to coding agents (some people will want to work the way she spells out, especially earlier in their career, some people won't), but doesn't make as much sense for the example she uses for the remaining 2/3rds of the piece.
- JoshTriplett 1y ago
- visarga 1y ago> We're really good at cumulative iteration. Humans are turbo optimized for communities, basically. This is why brainstorming is so effective… But usually only in a group. There is an entire theory in cognitive psychology about cumulative culture that goes directly into this and shows empirically how humans work in groups. > Humans learn collectively and innovate collectively via copying, mimicry, and iteration on top of prior art. You know that quote about standing on the shoulders of giants? It turns out that it's not only a fun quote, but it's fundamentally how humans work. Creativity is search. Social search. It's not coming from the brain itself, it comes from the encounter between brain and environment, and builds up over time in the social/cultural layer. That is why I don't ask myself if LLMs really understand. As long as they search, generating ideas and validating them in the world, it does not matter. It's also why I don't think substrate matters, only search does. But substrate might have to do with the search spaces we are afforded to explore.
- computerthings 1y ago[dead]
- sim7c00 1y agoi very much agree with this article. i do wonder if you could make a prompt to force your LLM to always respond like this and if that would already be a sort of dirty fix... im not so clever at prompting yet :')
- JonathanRaines 1y ago> "You seem stuck on X. Do you want to try investigating Y?" MS Clippy was the AI tool we should all aspire to build
- mikaylamaki 1y agoThe code gen example given later sounds an awful lot like what AWS built with Kiro[1] and it's spec feature. This article is kinda like the theory behind the practice in that IDE. I wish the tools described existed instead of all these magic wands [1] https://kiro.dev/blog/introducing-kiro/ https://kiro.dev/blog/introducing-kiro/
- hazelweakly 1y agoIndeed it is! In fact I actually had something like Kiro almost perfectly in mind when I wrote the article (but didn’t know AWS was working on it at the time). I was very happy to see AWS release Kiro. It was quite validating to me seeing them release it and follow up with discussions on how this methodology of integrating AI with software development was effective for them
- mikaylamaki 1y agoYeah! It's a good sign that this idea is in the air and the industry is moving towards it. Excited to get rid of all the magic wands everywhere :D
- nextworddev 1y agoThis post is confusing one big point which is that the purpose of AI deployments isn’t to teach so that humans get smarter but to achieve productivity at the process level by eliminating work that isn’t rewarded for human creativity
- staticshock 1y agoThe spirit of this post is great. There are real lessons here that the industry will struggle to absorb until we reach the next stage of the AI hype cycle (the trough of disillusionment.) However, I could not help but get caught up on this totally bonkers statement, which detracted from the point of the article: > Also, innovation and problem solving? Basically the same thing. If you get good at problem solving, propagating learning, and integrating that learning into the collective knowledge of the group, then the infamous Innovator’s Dilemma disappears. This is a fundamental misunderstanding of what the innovator's dilemma is about. It's not about the ability to be creative and solve problems, it is about organizational incentives. Over time, an incumbent player can become increasingly disincentivized from undercutting mature revenue streams. They struggle to diversify away from large, established, possibly dying markets in favor of smaller, unproven ones. This happens due to a defensive posture. To quote Upton Sinclair, "it is difficult to get a man to understand something when his salary depends upon his not understanding it." There are lots of examples of this in the wild. One famous one that comes to mind is AT&T Bell Labs' invention of magnetic recording & answering machines that AT&T shelved for decades because they worried that if people had answering machines, they wouldn't need to call each other quite so often. That is, they successfully invented lots of things, but the parent organization sat on those inventions as long as humanly possible.
- ankit219 1y agoThe very premise of the article is that tasks are needed for humans to learn and maintain skills. Learning should happen independently, it is a tautological argument that since human wont learn with agents which can do more, we should not have agents which can do more. While this is a broad and complex topic (i will share a longer blog that i am yet to fully write), I think people underestimate the cognitive load it takes to go to the higher level pattern and hence learning should happen not on the task but before the task. We are in the middle of peer vs pair sort of abstraction. Is the peer reliable enough to be delegated the task? If not, the pair design pattern should be complementary to human skill set. I sensed the frustration with ai agents came from being not fully reliable. That means a human in the loop is absolutely needed, and if there is a human, dont have ai being good at what human can do, instead it be good assistant by doing things human would need. I agree on that part, though if reliability is ironed out, for most of my tasks, i am happy ai can do the whole thing. Other frustrations stem from memory or lack of(in research), hallucinations and overconfidence, lack of situational awareness (somehow situational awareness is what agents market themselves on). If these are fixed, treating agents as a pair vs treating agents as a peer might tilt more towards the peer side.
- tumidpandora 1y agoI wholeheartedly agree with OP, the article is very timely and illustrates the mis-direction of AI tooling clearly. I often find myself and my kids asking LLMs "don't tell me the answer, work with me to resolve it" that is a much richer experience than seeing a fully baked out response to my question. I hope we see grater adoption of OPs EDGE framework in AI interactions.
- shikon7 1y agoAI context and instruction following has become so good, probably you could put this verbatim in an AI prompt, and the AI would react according to the post.
- meander_water 1y agoI completely agree with this. I recently helped my dad put together a presentation. He is an expert in his field, so he already had all the info ready in slides. But he's not a designer, he doesn't know how to make things look "good". I tried a handful of "AI slide deck" apps. They all had a slick interface where you could generate an entire slide deck with a few words. But absolutely useless for actually doing what users need, which is to help them beautify rather than creating content.
- prats226 1y agoInterestingly, deepseek paper mentions RL with process reward model. However they mentioned it failed to align model correctly due to subjectivity involved in defining if the intermediate step in process is right or wrong
- imranq 1y agoWhile I agree with the author's vision for a more human-centric AI, I think we're closer to that than the article suggests. The core issue is that the default behavior is what's being criticized. The instruction-following capabilities of modern models mean we can already build these Socratic, guiding systems by creating specific system prompts and tools (like MCP servers). The real challenge isn't technical feasibility, but rather shifting the product design philosophy away from 'magic button' solutions toward these more collaborative, and ultimately more effective, workflows
- danieltanfh95 1y agohttps://danieltan.weblog.lol/2025/06/agentic-ai-is-a-bubble-but-im-still-trying-to-make-it-work https://danieltan.weblog.lol/2025/06/agentic-ai-is-a-bubble-... We should be working to make HITL tools, not HOTL workflows where the humans are expected to just work with the final output. At some point the abstraction will leak.
- computerthings 1y ago[dead]