18 ms·
The current hype around autonomous agents, and what actually works in production
- Retr0id 1y ago> Each new interaction requires processing ALL previous context I was under the impression that some kind of caching mechanism existed to mitigate this
- _heimdall 1y agoCaching would only help to keep the context around, but caching would only be needed if it still ultimately needs to read and process that cached context again.
- Retr0id 1y agoYou can cache the whole inference state, no? They don't go into implementation details but Gemini docs say you get a 75% discount if there's a context-cache hit: https://cloud.google.com/vertex-ai/generative-ai/docs/context-cache/context-cache-overview https://cloud.google.com/vertex-ai/generative-ai/docs/contex...
- _heimdall 1y agoIt that just avoids having to send the full context for follow-up requests, right? My understanding is that caching helps to keep the context around but can't avoid the need to process that context over and over during inference.
- deleted 1y ago[deleted]
- bakugo 1y agoThe initial context processing is also cached, which is why there's a significant discount on the input token cost.
- _heimdall 1y agoWhat exactly is cached though? Each loop of token inference is effectively a recursive loop that takes in all context plus all previously inferred tokens, right? Are they somehow caching the previously inferred state and able to use that more efficiently than if they just cache the context then run it all through inference again?
- csomar 1y agoMy understanding is that caching reduce computation but the whole input is still processed. I don’t think is fully disclosing how their cache works. LLMs degrade with long input regardless of caching.
- blackbear_ 1y agoYou have to compute attention between all pairs of tokens at each step, making the naive implementation O(N^3). This is optimized by caching the previous attention values, so that for each step you only need to compute attention between the new token and all previous ones. That's much better but still O(N^2) to generate a sequence of N tokens.
- stpedgwdgfhgdd 1y agoCompact the conversation (CC)
- ilaksh 1y agoYes, prompt caching helps a lot with the cost. It still adds up if you have some tool outputs with long text. I have found that breaking those out into subtasks makes the overall cost much more reasonable.
- Too 1y agoWhen inference requires maxing out the memory of a gpu, where are you planning to keep this cache? Unless there is a way to compress the context into a more manageable snapshot, the cloud provider surely won’t keep a gpu idling just for holding a conversation in memory.
- dmezzetti 1y agoIt's clear that what we currently call AI is best suited for augmentation not automation. There are a lot of productivity gains available if you're willing to accept that.
- andrekandre 1y ago> AI is best suited for augmentation not automation. i agree with this sentiment, but with the caveat of "when its not lying to you". the most frustrating part of these interactive ai assistants is when it sends me down a rabbit hole of an api that doesn't exist (but looks almost right)
- vntok 1y ago> Production systems need 99.9%+ reliability This is not remotely true. Think of any business process around your company. 99.9% availability would mean only 1min26 per day allowed for instability/errors/downtime. Surely your human collaborators aren't hitting this SLA. A single coffee break immediately breaks this (per collaborator!). Business Process Automation via AI doesn't need to be perfect. It simply needs to be sufficiently better than the status quo to pay for itself.
- hansmayer 1y agoThis may not be about internal business processes. In e-commerce 90 sec can be a lot of revenue lost, and mission-critical applications such as telecommunications or air control, it would be downright a disaster (ever heard of five nines availability)?
- lexicality 1y agoCurrently I'm thinking about how furious the developers get any time Jenkins has any kind of hiccough, even if the solution is just "re-run the workflow" - and that's just network timeouts! I don't want to imagine the tickets if the CI system started spitting out hallucinations...
- Pasorrijer 1y agoI think you're crossing reliability and availability. Reliability means 99.9% of the time when I hand something off to someone else it's what they want. Availability means I'm at my desk and not at the coffee machine. Humans very much are 99.9% accurate, and my deliverable even comes with a list of things I'm not confident about
- vntok 1y ago> Humans very much are 99.9% accurate This is an extraordinary claim, which would require extraordinary evidence to prove. Meanwhile, anyone who spends a few hours with colleagues in a predominantly typing/data entry/data manipulation service (accounting, invoicing, presales, etc.) KNOWS the rate of minor errors is humongous.
- KoolKat23 1y agoHuman multi-step workflows tend to have checkpoints where the work is validated before proceeding further, as humans generally aren't 99%+ accurate either. I'd imagine future agents will include training to design these checks into any output, validating against the checks before proceeding further. They may even include some minor risk assessment beforehand, such as "this aspect is crucial and needs to be 99% correct before proceeding further".
- a_bonobo 1y agoThat's what Claude Code does - it constantly stops and asks you whether you want to proceed, including showing you the suggested changes before they're implemented. Helps with avoiding token waste and 'bad' work.
- KoolKat23 1y agothats good to hear, theyre on their way there! on a personal note, I'm happy to hear that. I've been apprehensive and haven't tried it, purely due to my fear of the cost.
- queenkjuul 1y agoMy work has a corporate subscription and on the one hand it's very impressive and on the other i don't actually find it useful.
- Filligree 1y agoIt’s best at small to medium projects written in a consistent style. So. It’s a potential superpower for personal projects, yet I don’t see it being very useful in a corporate setting. I used Claude Code to make this little thing: https://github.com/Baughn/ScriptView https://github.com/Baughn/ScriptView …took me thirty minutes. It wouldn’t have existed otherwise.
- iwontberude 1y ago
- infecto 1y agoLink does not work for me but as someone who does a lot of work with LLMs I am also betting against agents. Agents have captivated the minds of groups of people in each large engineering org. I have no idea what their goal is other then they work on “GenAI”. For over a year now they have been working on agents with the promise that the next framework that MSFT or Alphabet publishes will solve their woes. They don’t actually know what they are solving for except everything involves agents. I have yet to see agents solve anything but for some reason this idea that having an agent that you can send anything and everything will solve all problems for the company. LLMs have a ton of interesting applications but agents have yet to grasp me as interesting, I also don’t understand why so many large companies have focused time around it. They are not going to be cracking the code ahead of a commercial tool or open source project. In the time spent toying around with agents there are a lot of interesting applications that could have built, some of which may be technically an agent but without so much focus and effort on trying to solve for all use cases. Edit: after rereading my post wanted to clarify that I do think there is a place for tool call chains and the like but so many folks I have talked to first hand are trying to create something that works for everything and anything.
- JKCalhoun 1y agoLink is working for me — perhaps it was not 30 minutes ago? (Safari, MacOS)
- johnisgood 1y agoI have no idea what agents are for, could be my own ignorance. That said, I have been using LLMs for a while now with great benefit. I did not notice anything missing, and I am not sure what agents bring to the table. Do you know?
- mhog_hn 1y agoAn agent is an LLM + a tool call loop - it is quite a step up in terms of value in my experience
- 1y ago
- danieltanfh95 1y agoSame. https://danieltan.weblog.lol/2025/06/agentic-ai-is-a-bubble-but-im-still-trying-to-make-it-work https://danieltan.weblog.lol/2025/06/agentic-ai-is-a-bubble-... The fundamental difference is we need HITL to reduce errors instead of HOTL which leads to the errors you mentioned
- Xmd5a 1y ago>A database query might return 10,000 rows, but the agent only needs to know "query succeeded, 10k results, here are the first 5." Designing these abstractions is an art. It seems the author never used prompt/workflow optimization techniques. LLM-AutoDiff: Auto-Differentiate Any LLM Workflow https://arxiv.org/pdf/2501.16673 https://arxiv.org/pdf/2501.16673
- constantcrying 1y agoNo, it is not "mathematically impossible". It is empirically implausible. There is no statement in mathematics that says that agents can not have a 99.999% reliability rate. Also, if you look at any human process you will realize that none of them have a 100% reliability rate. Yet, even without that we can manufacture e.g. a plane, something which takes millions of steps, each without a 100% success rate. I actually think the article makes some good points, but especially when you are making good points it is unnecessary to stretch credibility with exaggerating your arguments.
- macleginn 1y agoThis is a good point, but it seems, empirically, that most parts of a standard passenger airplane have reliability approximating 100% in a predefined time window with proper inspection and maintenance, otherwise passenger transit would be impossible. When the system does start to degrade, e.g. because replacement parts and maintenance becomes unavailable or too costly (cf. the use of imported planes by Russian airlines after the sanctions hit), incidents quickly start piling up.
- constantcrying 1y agoIt's about what you do with errors. If you let them compound they lead to destruction, if instead you inspect, maintain, reinspect, replace, etc. you can manage them. My point was that something extremely complex, like a plane, works, because the system tries hard to prevent compounding errors.
- sarchertech 1y agoThat works because each plane is (nearly) exactly the same as the one before it and we have exact specifications for the plane. You can do maintenance, inspections, and replacement because of those specifications. In software the equivalent of blueprints is code. The room for variation outside software “specifications” is infinite. Human reliability when comes to assembling planes is also much higher than 99%, and LLM reliability creating code is much, much lower than 99%.
- deadbabe 1y agoI just want someone to give me one legit use case where an AI Agent now enables them to do something that couldn’t be done before, and actually makes an impact on overall profit.
- stavros 1y agoI can write code I wouldn't have been bothered to before, and make money from it.
- digitcatphd 1y agoI’m sure most of the problems cited in this article will be easily solved within the next five years or so, waiting for perfection and doing nothing won’t pay dividends
- snappr021 1y agoThe alternative is building Functional Intelligence process flows from the ground up on a foundation of established truth? If 50% of training data is not factually accurate, this needs to be weeded out. Some industries require a first principles approach, and there are optimal process flows that lead to accurate and predictable results. These need research and implementation by man and machine.
- mritchie712 1y ago> I've built 12+ production AI agent systems across development, DevOps, and data operations It's hard to make *one* good product (see startup failure rates). You couldn't make 12 (as seemingly a solo dev?) and you're surprised? we've been working on Definite[0] for 2 years with a small team and it only started getting really good in the past 6 months. 0 - data stack + AI agent: https://www.definite.app/ https://www.definite.app/
- AstroBen 1y agoThey've built 12+ products with a full time job for the last 3 years Something seems off about that...
- senko 1y agoHis full time job is building AI systems for others (and the article is a well written promo piece). If most of these are one-shot deterministic workflows (as opposed of input-llm-tool loop usually meant by the current use of the term "ai agent"), it's not hard to assume you can build, test and deploy one in a month on average.
- Rexxar 1y agoHe didn't say he made 12 independent saleable products, he says he built 12 tools that fill a need at his job and are used in production. They are probably quite simple and do a very specific task as the whole article is telling us that we have to keep it simple to have something useable.
- mritchie712 1y agothat's my point. He's "Betting Against AI Agents" without having taken a serious attempt at building one. > agents that technically make successful API calls but can't actually accomplish complex workflows because they don't understand what happened. It takes a long time to get these things right.
- RamblingCTO 1y agoI also build agents/ai automation for a living. Coding agents or anything open-ended is just a stupid idea. It's best to have human validated checkpoints, small search spaces and very specific questions/prompts (does this email contain an invoice? YES/NO). Just because we'd love to have fully intelligent, automatic agents, doesn't mean the tech is here. I don't work on anything that generates content (text, images, code). It's just slob and will bite you in the ass in the long run anyhow.
- la_fayette 1y agoIn general I would agree, however the resulting systems of such an approach tend to be "just" expensive workflow systems, which could be done with old tech as well... Where is the real need for anything LLM here?
- barbazoo 1y agoExtracting structured data from unstructured text comes to mind. We’ve built workflows that we couldn’t before by bridging a non deterministic gap. It’s a business SaaS but the folks using our software seem to be really happy with the result.
- anon191928 1y agoit would take months with old tech to create a bot that can check multiple websites for specific data or information? so LLM reduces the time a lot? am I wrong?
- dlisboa 1y agoMonths? Scraping wasn’t a hard problem then. Classifying information is a different and more complex thing, which is what these models are very good at. Then again we had other means of classification before LLMs without having to go through chat bots.
- RamblingCTO 1y agoI think classical ML should still be compared when you use an LLM as a classifier. some problems are so well defined you can just drop an SVM on it and be done with it. the biggest benefit of LLMs I think is you can get results that are ok quite fast. no need to split or clean data, compare metrics etc. etc.
- rco8786 1y agoI still don’t even know what an agent is. Everyone seems to have their own definition. And invariably it’s generic vagaries about architecture, responsibilities of the LLM, sub-agents, comparisons to workflows, etc. But still not once have I seen an actual agent in the wild doing concrete work. A “No True Agent” problem if you will.
- iamjackg 1y agoTechnically speaking, Claude Code is an agent, for example. It's just a fancy term for an LLM that can call tools in a loop until it thinks it's done with whatever it was tasked to do. ChatGPT's Deep Research mode is also an agent: it will keep crawling the web and refining things until it feels it has enough material to write a good response.
- neom 1y ago"The real challenge isn't AI capabilities, it's designing tools and feedback systems that agents can actually use effectively." - this part I agree with - I'd been sitting the AI stuff out because I was unclear where I thought the dust would settle or what the market would accept, but recently joined a very small startup focused on building an agent. I've gone from skeptical to willing to humor to "yeah this is probably right" in about 5 months, basically I believe: if you scope the subject matter very very well, and then focus on the tooling that the model will require to do it's task, you get a high completion rate. There is a reluctance to lean into the non deterministic nature of the models, but actually if you provide really excellent tooling and scope super narrowly, it's generally acceptably good. This blog post really makes the tooling part seem hard, and, well... it is, but not that hard - we'll see where this all goes, but I remain optimistic.
- johndhi 1y agoFrom what I understand customer support chatbots have had some pretty good outcomes from ai agents. Or does that not count?
- nsypteras 1y agoI think that would be one of the success cases described in the article because HITL is an integral part of good customer support chatbots. Support chats can be escalated to a human whenever the agent is unable to provide a satisfactory answer to the user.
- jvanderbot 1y agoMy AI tool use has been a net positive experience at work. It can take over small tasks when I need a break, clean up or start momentum, and generally provide a good helping hand. But even if it could do my job, the costs pile up really quickly. Claude Code can burn $25/ 1-2 hrs, easily on a large codebase, and that's creeping along at a net positive rate assuming I can keep it on task and provide corrections. If you automate the corrections we are up to $50/hr or some tradeoff of speed, accuracy, and cost. Same as it's always been. For agents, that triangle is not very well quanitfied at the moment which makes all these investigations interesting but still risky.
- swader999 1y agoSubscription?
- jvanderbot 1y agoI have one, and upgrades don't have unlimited access as far as I can tell. Correct me if I'm wrong. This cost scaling will be an issue for this whole AI employee thing, especially because I imagine these providers are heavily discounting.
- 13zebras 1y agoRe: discounting… Given that OpenAI is burning billions and making trivial revenue in comparison, the cost per token is probably going to skyrocket when Sam runs out of BS to con the next investor. I’m guessing the only way that token cost doesn’t explode is if Claude ends up in Amazon’s hands and OpenAI is Microsoft’s. Then Amazon, Google, and MS can subsidize if they want. But as standalone businesses, they can’t make it at current token prices. IMHO
- joshvm 1y agoThere are usage limits, but the argument is that unless you're writing and modifying large swathes of code in YOLO mode, you don't hit them. At least for what I would call a small and tedious task. I'm thinking "write a docstring", "add type annotations", "write a single unit test for this case", "fill in this function". For a good prompt these are often solved in <10 interactions. Especially when combined with scoped rules that are pulled in on demand to guide output.
- atomon 1y agoIs the main point “let me mathematically prove that it’s impossible to do what I’ve already done 12 times this year?” Yes, very long workflows with no checks in between will have high error rates. This is true of human workflows too (which also have <100% accuracy at each step). Workflows rarely have this many steps in practice and you can add review points to combat the problem (as evidenced by the author building 12 of these things and not running into this problem)
- tomhow 1y ago[stub for offtopicness]
- roschdal 1y agoAI is for people without natural intelligence.
- deleted 1y ago[deleted]
- bboygravity 1y agoSo it's for 90+ percent of society? Sounds like good business to me.
- block_dagger 1y agoDownvotes are for comments like yours
- satyrun 1y agoYea just average IQ like Terence Tao. All you are really saying with this comment is you have an incredibly narrow set of interests and absolutely no intellectual curiosity.
- 7r8fze937fjd 1y ago[flagged]
- paradite 1y agoThis is obviously AI generated, if that matters. And I have an AI workflow that generates much better posts than this.
- Retr0id 1y agoI think it's just written by someone who reads a lot of LLM output - lots of lists with bolded prefixes. Maybe there was some AI-assistance (or a lot), but I didn't get the impression that it was AI-generated as a whole.
- actinium226 1y agoVery nice article. The point about mathematical reliability is interesting. I generally agree with it, but humans aren't 100% reliable, or even 99% reliable, so how do we manage to create things like the Linux kernel or the Mars landers without AI? Clearly we have some sort of goal-based self-correction mechanism. I wonder if there's research into AI on that thread?
- an0malous 1y ago> Clearly we have some sort of goal-based self-correction mechanism. Humans can try things, learn, and iterate. LLMs still can't really do the second thing, you can feed back an error message into the prompt but the learning isn't being added to its weights so its knowledge doesn't compound with experience like it does for us. I think there are still a few theoretical breakthroughs needed for LLMs to achieve AGI and one of them is "active learning" like this.
- airstrike 1y ago100% and it seems like we need a whole new architecture to get there, because right now training a model takes so much time. At the risk of making a terrible analogy, right now we're able to "give birth" to these machines after months of training, but once they're born, they can't really learn. Whereas animals learn something new every day, got to sleep, clean up their memories a bit, deleting some, solidifying others, and waking up with an improved understanding of the world.
- bot403 1y agoMaybe you're on to something. We need AI lions which will eat the models which don't learn or adapt enough.
- airstrike 1y agoI love the idea of AI lions, but you still need to find a way to allow models to continue "learning" after they're born—which is PhD worthy. Right now we train AI babies, dump them in the wild... and expect them to have all the answers.
- hannofcart 1y ago> Let's do the math. If each step in an agent workflow has 95% reliability, which is optimistic for current LLMs,then: 5 steps = 77% success rate 10 steps = 59% success rate 20 steps = 36% success rate Production systems need 99.9%+ reliability. (End quote) Isn't this just wrong? Isn't the author conflating accuracy of LLM output in each step to accuracy of final artifact which is a reproducible deterministic piece of code? And they're completely missing that a person in the middle is going to intervene at some point to test it and at that point the output artifact's accuracy either goes to 100% or the person running the agent would backtrack. Either am missing something or this does not seem well thought through.
- vrighter 1y agoHow is it that the final result is a reproducible deterministic piece of code, when the prompts become the "source code" itself, and the underlying model used is constantly changing (being updated), which is equivalent to your programming language changing its semantics every other day and refusing to tell you exactly what has changed (because they can't). Not to mention the nondeterminism that a lot of times is present due to nondeterministic order of evaluation when parallelizing?
- hungryhobbit 1y agoDid you even finish the article? The end is all about the trade-off of when "a person in the middle is going to intervene". In fact, the point of the whole article isn't that AI doesn't work; to the contrary, it's that long chains of (20+) actions with no human intervention (which many agentic companies promise) don't work.
- coliveira 1y agoHe's not wrong. The numbers are too pessimistic, however when building software the numbers don't need to be as high for a complete disaster to happen. Even if just 1% of the code is bad, it is still very difficult to make this work. And you mention testing, which certainly can be done. But when you have a large product and the code generator is unreliable (which LLMs always are), then you have to spend most of your time testing.
- alpha_squared 1y agoOne thing I'll add that isn't touched on here is about context windows. While not "infinite", humans have a very large context window for problems they're specialized in solving. Models can often overcome their context window limitations by having larger and more diverse training sets, but that still isn't really a solution to context windows. Yes, I get the context window increases over time and that for many purposes it's already sufficient enough, but the current paradigm forces you to compress your personal context into a prompt to produce a meaningful result. In a language as malleable as English, this doesn't feel like engineering so much as it feels like incantations and guessing. We're losing so, so much by skipping determinism.
- lxgr 1y agoHumans don't have this fixed split into "context" and "weights", at least not over non-trivial time spans. For better or worse, everything we see and do ends up modifying our "weights", which is something current LLMs just architecturally can't do since the weights are read-only.
- alpha_squared 1y agoI agree, I'm mostly trying to illustrate how difficult it is to fit our working model of the world into the LLM paradigm. A lot of comments here keep comparing the accuracy of LLMs with humans and I feel that glosses over so much of how different the two are.
- globular-toast 1y agoThis is why I actually argue that LLMs don't use natural language. Natural language isn't just what's spoken by speakers right now. It's a living thing. Every day in conversation with fellow humans your very own natural language model changes. You'll hear some things for the first time, you'll hear others less, you'll say things that get your point across effectively first time, and you'll say some things that require a second or even third try. All of this is feedback to your model. All I hear from LLM people is "you're just not using it right" or "it's all in the prompt" etc. That's not natural language. That's no different from programming any computer system. I've found LLMs to be quite useful for language stuff like "rename this service across my whole Kubernetes cluster". But when it comes to specific things like "sort this API endpoint alphabetically" I find the amount of time to learn to construct an appropriate prompt is the same if I'd have just learnt to program, which I already have done. And then there's the energy used by the LLM to do it's thing which is enormously wasteful.
- afro88 1y ago> Error rates compound exponentially in multi-step workflows. 95% reliability per step = 36% success over 20 steps. Production needs 99.9%+. This misses a key feature of agents though. They get feedback from linters, build logs, test runs and even screenshots. And they collect this feedback themselves. This means they can error correct some mistakes along the way. The math works out differently, depending on how well it can collect automated feedback it is doing what you want.
- whazor 1y agoCorrect, I think it is better to see it as multiple stages. Investigation stage might spin off tasks to read files, perform searches online, summarise the request. Then ‘main stage’ where it performs changes. Afterwards indeed the testing+fixing stage where it verifies the results and potentially performs a couple fixes. These plans are predictable and the models learn which steps are relevant first particular projects. For context, relevant information from steps can be cherrypicked to next stage. The math works differently because AI (mostly) ignores irrelevant results. So steps actually increase reliability overall.
- hemantv 1y agoLlm are great reflections. Issues I have come across too large of context confuse the llm. Second since llm are non deterministic in nature how do you know if the quality went from 90% to 30% there is no test you can write. What if model provider degrades quality you have no test for it
- arisAlexis 1y ago"forever"?
- oceanparkway 1y agoI think the "math" on reliability-over-steps will end up differently than described here in the long term because getting new factual input from the real world should improve the reliability of the end state, and we have all observed agentic systems at this point producing that behavior at least sometimes (e.g., a test failure prompts claude code to refactor correctly). Whether or not one term in this equation currently compounds faster is a good question, or under what circumstances, etc., but presenting agentic abilities as always flawed thinking resulting in impossible long term task execution isn't right. Humans are flawed and require long, drawn out multi task thinking to get correct answers, and interacting with and getting feedback from the world outside the mind during a task execution process typically raises the chance of the correct answer being spit out in the end. I'd agree that the agentic math isn't great at the moment, but if it's possible to reduce hallucinations or raise the strength and frequency effect of real world feedback on the model, you could see this playing out differently perhaps quite soon. There's at least a couple of examples of "we're already there".
- yunyu 1y ago"Your fancy AI scaffolds will be washed away by scale." - Noam Brown
- lmeyerov 1y agoI used to believe the error rate fallacy, but: 1. Multi-turn agents can correct themselves with more steps, so the reductive error cascade thinking here is more wrong than right in my experience 2. The 99.9% production requirement is so contextual and misleading, when the real comparison is often something like "outage", "dead air", "active incident", "nobody on it", "prework before/around human work", "proactive task no one had time for before", etc. Similar to infra as code, CI, and many other automation processes, there's mountains of work that isn't being done and LLMs can do entirely or large swathes of
- vrighter 1y agoHow about those large swaths are done with LLMs, but instead of spending all that time reviewing that code (really reviewing it, not just a brief LGTM), which would make the time savings moot, you just decide to personally assume responsibility for that code being dead wrong sometimes and the consequences it causes (you cannot blame the AI. As far as anyone is concerned, you wrote the code and signed off on it). As in, legal liability. Would you take the deal?
- ankit219 1y agoThese are all solvable problems. The issue is given the race to get to a certain ARR quickly, many startups end up not focusing on these. There is some truth to AI agents being not as useful as their promise, but the problems mentioned are engineering problems, and once we start seeing them with a different lens, they would start working. (This is not to say I believe orchestration or multi step agents are a way to go, I personally lean heavily towards RL. Just that the criticisms here assumes the state would remain the same even without AI advancement). Eg: you need good verifiers (to understand whether a task is done successfully or not). Many tasks have easier verifications than doing the task. YOu have five parallel generations with 80% accuracy, the probablity of getting one right (and a verifier which can pick that) goes to 99.96%. With multi step too, the math changes in a similar manner. It just needs a different approach than how we have built software till date. He even hints at a paradigm with 3-5 discrete step workflow which works superbly well. We need to build more in that way.
- throwaway423342 1y agoIs it reasonable to assume the five generations are independent?
- ankit219 1y agoThey are not completely independent. It's a good assumption though. If a model encounters something out of distribution then all five of the generations will fail. If the model knows and went in a wrong direction (due to lack of reliability), within five generations, it can be corrected. You need evals, runtime verifiers as basic harness for AI systems.
- jackblemming 1y agoThis is correct. Multiple different agents trying, multiple retries, and many other different solutions can help with this. I have seen agents try one method, get negative feedback, and then try another working method.
- majormajor 1y ago> Many tasks have easier verifications than doing the task. In the software world (like the article is talking about) this is the logic that has ruthlessly cut software QA teams over the years. I think quality has declined as a result. Verifiers are hard because the possible states of the internal system + of the external world multiply rapidly as you start going up the component chain towards external-facing interfaces. That coordination is the sort of thing that really looks appealing for LLMs - do all the tedious stuff to mock a dependency, or pre-fill a database, etc - but they have an unfortunate tendency to need to be 100% correct in order for the verification test that depends on them to be worth anything. So you can go further down the rabbit hole, and build verifiers for each of those pre-conditions. This might recurse a few times. Now you end up with the math working against you - if you need 20 things to all be 100%, then even high chances of each individual one starts to degrade cumulatively. A human generally wouldn't bother with perfect verification of every case, it's too expensive. A human would make some judgement calls of which specific things to test in which ways based on their intimate knowledge of the code. White box testing is far more common than black box testing. Test a bunch of specific internals instead of 100% permutations of every external interface + every possible state of the world. But if you let enough of the code to solve the task be LLM-generated, you stop being in a position to do white-box testing unless you take the time to internalize all the code the machine wrote for you. Now your time savings have shrunk dramatically. And in the current state of the world, I find myself having to correct it more often then not, further reducing my confidence and taking up more time. In some places you can try to work around this by adjusting your interfaces to match what the LLM predicts, but this isn't universal. --- In the non-software world the situation is even more dire. Often verification is impossible without doing the task. Consider "generate a report on the five most promising gaming startups" - there's no canonical source to reference. Yet these are things people are starting to blindly hand off to machines. If you're an investor doing that to pick companies, you won't even find out if you're wrong until it's too late.
- Dachande663 1y agoOP here. I posted this, this morning and then promptly forgot about it. How come the title has been changed from the blog posts own?
- wrp 1y agoThat was annoying. Saw the post, then later had a hard time finding it again.
- nextworddev 1y agoActually, author should be bullish on autonomous agents considering 90% of what he's even able to do now wasn't even possible in early 2024, so you shouldn't bet against the slope of progress
- wrp 1y agoTFA is a bit rambling and readers are getting distracted by specific claims, like the bit about 99.9%+ reliability. TFAs main point is that productive use of AI agents requires tightly specified context and frequent human intervention, which is what folks have been saying for a while.
- arendtio 1y agoThe compounding error rate in long-running processes is just one side of the coin. You can also use models to catch errors, and those success rates compound as well. So, it's not like you have no options to fight against a giant failure rate monster...
- Arn_Thor 1y agoI spoke with an Amazon AI production engineer who’s talking with prospective clients about implementing AI in our business. When a colleague asked about using generative AI in customer facing chats the engineer said he knows of zero companies who don’t have a human in the loop. All the automatic replies are non-generative “old” tech. Gen AI is just not reliable enough for anyone to stake their reputation on it.
- PaulHoule 1y agoYears ago I was interested in agents that used "old AI" symbolic techniques backed up with classical machine learning. I kept getting hired though by people who were working on pre-transformer neural nets for texts. Something I knew all along was that you build the system that lets you do it with the human in the loop, collect evaluation and training data [1] and then build a system which can do some of the work and possibly improve the quality of the rest of it. [1] in that order because for any 'subjective' task you will need to evaluate the symbolic system even if you don't need to train it -- if you need to train the system, on the other hand, you'll still need to eval
- throwehshdhdy 1y agoPlenty of tech companies have started using gen AI for live chat support. Off the top of head I know off sonder.com and wealthsimple.com. If the LLM can’t answer a query it usually forwards the chat to a human support agent.
- actinium226 1y agoAir Canada did this a bit ago, and their AI gave the customer a fake process for submitting claims for some sort of discount on airfare due to bereavement (the flight was for a funeral). The customer sued and Air Canada's defense was that he shouldn't have trusted the Air Canada AI chatbot. Air Canada lost.
- medbrane 1y agoThat was in 2022, before LLMs, and they "lost" as in they had to pay back $482 USD.
- Apocryphon 1y agoBuilding shovels during a fool’s gold rush, nice
- esac 1y agoThis is exactly right! I'm happy people are starting to care about compounding errors, we use the sigma terminology: https://www.silverstream.ai/blog-news/2sigma https://www.silverstream.ai/blog-news/2sigma Agents are digital manufacturing machines and benefit from the same processes we identified for reliability in the real world
- swyx 1y ago> The Mathematical Reality No One Talks About literally everybody talks about this lmao what are you on about https://www.youtube.com/watch?v=d5EltXhbcfA https://www.youtube.com/watch?v=d5EltXhbcfA
- deleted 1y ago[deleted]
- thedudeabides5 1y agoSee "Campbell's Completeness Conjecture" https://www.campbellramble.ai/p/dont-trust-machines https://www.campbellramble.ai/p/dont-trust-machines
- a_bonobo 1y ago>Enterprise systems aren't clean APIs waiting for AI agents to orchestrate them. They're legacy systems with quirks, partial failure modes, authentication flows that change without notice, rate limits that vary by time of day, and compliance requirements that don't fit neatly into prompt templates. Perhaps that's why MCP as a protocol is so interesting to people - MCP servers are a chance at a 'blank slate' in front of the enterprise system. You pull out only the parts you're interested in, you get to define clear boundaries when you build the MCP server, the LLM sees only what you want it to see and you hide the messiness of the enterprise system.
- eska 1y agoIt’s on the wrong layer for that. MCP is akin to putting GraphQL over an old and crufty SOAP interface. There’s some “intelligence” in the GraphQL layer, but it doesn’t fix flaws in the lower layer such as side effects that shouldn’t be there.
- physicsguy 1y agoThat's not really much different to writing a REST-ful interface over the top, and sticking an OpenAPI compliant schema on the front.
- rudderdev 1y agoSame experience. I started in mid 2023 building agents. The agents that are still working on production - they do one specific thing and the only coordination/integration I allow is via automation framework based on deterministic api only.
- lukaslalinsky 1y agoI was one of the early adopters of GitHub Copilot and generally a proponent of AI assisted coding. I've recently tried "vibe coding" and oh my god, the experience couldn't be more different to what I was used to. It feels like a super expensive machine, making all kinds of mistakes and charging me for all of them. So many trial and error attempts. I ask it to do X, it conveniently does a lot of work around it, but in the end X does not work, so it just comments it out as a minor issue. It requires so much micro management, that I don't really see the purpose. Much easier and faster to just write the code myself and let it help me in that process. With agents, I feel like I'm the one helping it get a job done. I honestly can't imagine trusting this with any kind of production process.
- hoverbot 1y agoGreat practical takes - agentic chatbots work only with real data and feedback loops. HoverBot is built this way: it automates RAG pipelines, includes confidence thresholds, and supports human-in-loop override. You can run vertical assistants (support, sales, HR) from a single dashboard with real-world reliability.