24 ms·
Claude Advanced Tool Use
- tfirst 11mo agoWe seem to be on a cycle of complexity -> simplicity -> complexity with AI agent design. First we had agents like Manus or Devin that had massive scaffolding around them, then we had simple LLMs in loops, then MCP added capabilities at the cost of context consumption, then in the last month everything has been bash + filesystem, and now we're back to creating more complex tools. I wonder if there will be another round of simplifications as models continue to improve, or if the scaffolding is here to stay.
- Aperocky 11mo agoMost of the time people sit on complex because they don't have a strong enough incentive to move from something that appears/happen to work, with AI, cost would be a huge incentive.
- behnamoh 11mo agoThis is what I've been talking about for a few months now. the AI field seems to reinvent the wheel every few months. And because most people really don't know what they're talking about, they just jump on the hype and adopt the new so-called standards without really thinking if it's the right approach. It really annoys me because I have been following some open source projects that have had some genuinely novel ideas about AI agent design. And they are mostly ignored by the community. But as soon as a large company like Anthropic or OpenAI starts a trend, suddenly everyone adopts it.
- fishmicrowaver 11mo agoWell, what are those projects? I don't speak for anyone else, but I'm generally fatigued by the endless parade of science fair projects at this point, and operate under the assumption that if an approach is good enough, openai/anthropic/google will fold useful ideas under their tools/products.
- mettamage 11mo agoHmm the Gemini API doesn’t need MCP for tool-use if I understand correctly. It just needs registered functions
- simonw 11mo agoI don't think any of the mainstream vendor APIs require MCP for tool use - they all supported functions (generally defined using a chunk of OpenAPI JSON schema) before the MCP spec gained widespread acceptance and continue to do so today.
- lebovic 11mo agoYep, the Anthropic API supported tool use well before an MCP-related construct was added to the API (MCP connector in May of this year). While it's not an API, Anthropic's Agent SDK does require MCP to use custom tools.
- roncesvalles 11mo agoIt's because attention dilution stymies everything. A new chat window in the web app is the smartest the model is ever going to be. Everything you prompt into its context, without sophisticated memory management* makes it dumber. Those big context frameworks are like giving the model a concussion before it does the first task. *which also pollutes the attention btw; saying "forget about this" doesn't make the model forget about it - it just remembers to forget about it.
- abraxas 11mo agoYeah, seems like the agent industry is spinning wheels a bit. As that old adage goes, when there are a hundred treatments you can be sure there is no cure.
- postalrat 11mo agoTools for tools. How about an LLM tool for tools?
- cube2222 11mo agoNice! Feature #2 here is basically an implementation of the “write code to call tools instead of calling them directly” that was a big topic of conversation recently. It uses their Python sandbox, is available via API, and exposes the tool calls themselves as normal tool calls to the API client - should be really simple to use! Batch tool calling has been a game-changer for the AI assistant we've built into our product recently, and this sounds like a further evolution of this, really (primarily, it's about speed; if you can accomplish 2x more tools calls in one turn, it will usually mean your agent is now 2x faster).
- polyrenn 11mo agoThe "write code to call tools instead of calling them directly" has been such an obvious path, the team at Huggingface & smolagents figured that out a while ago, agents that write code instead of natural language are just better for most cases.
- zbowling 11mo agoI wrote a better version of this idea: https://github.com/zbowling/mcpcodeserver https://github.com/zbowling/mcpcodeserver It works as an MCP proxy of sorts that converts all the child MCP tools into typescript annotations, asks your LLM to generate typescript, then executes those tool calls in a restricted VM to do the tool calls that way. It allows parellel process, passing data between tools without coming back to the LLM for a full loop, etc. The agents are pretty good at debugging issues they create too and trying again.
- cube2222 11mo agoCould you expand in what way it’s better? So far what you described sounds like what they did, but they manage the sandboxed environment for me and use Python rather than TypeScript. Do note that their thing works not only with MCP tools, but arbitrary tools.
- behnamoh 11mo agoI cannot believe all these months and years people have been loading all of the tool JSON schemas upfront. This is such a waste of context window and something that was already solved three years ago.
- michaelanckaert 11mo ago^ this. Careful design of what tools are passed when is key to good agent design.
- qntty 11mo agoSolved how?
- artursapek 11mo agoWhat is the right pattern? Do you just send a list of tool names & descriptions, and just give the agent an "install" tool that adds a given tool to the schema on the next turn?
- joshribakoff 11mo ago- claudes tool search tool - list of skills (markdown files) the agent can grep - claude skills - context compaction - sub-agents - plans There is no one “right” pattern. But yes it all generalizes to context engineering. With plans for example, you write out potential distractions for later, to keep (the AI and the Human context) focused on a task at hand. That pattern solves a distinctly different use case than the skills folder, but plans can also refer to skills in specific ways. Context engineering is evolving with overlapping complementary patterns, and while certain vendors are branding patterns, i think we will hopefully we see tools converge.
- orliesaurus 11mo agoFunction calling is back
- vessenes 11mo agoI'm confused about these tools - is this a decorator that you can add to your MCP server tools so that they don't pollute the context? How else would I add a "tool" for claude to use?
- pupppet 11mo agoWhat’s the best way to prevent the input context from compounding with each tool call?
- jameslk 11mo ago> Tool Search Tool, which allows Claude to use search tools to access thousands of tools without consuming its context window At some point, you run into the problem of having many tools that can accomplish the same task. Then you need a tool search engine, which helps you find the most relevant tool for your search keywords. But tool makers start to abuse Tool Engine Optimization (TEO) techniques to push their tools to the top of the tool rankings
- michaelanckaert 11mo agoDon't give anyone any ideas. We now have SEO, GEO, AEO and now TEO? :-p
- buremba 11mo agoJust wait for the people to update their LinkedIn titles to TEO expert. :)
- mkagenius 11mo agoI would argue that lot of the tools will be hosted on GitHub - infact, most of the existing repos are potentially a tool (in future). And the discovery is just a GitHub search btw gh repos are already part of training the llm So you don't even need internet to search for tools, let alone TEO
- michaelanckaert 11mo agoSecurity nightmare inbound... The example given by Anthropic of tools filling valuable context space is a result of bad design. If you pass the tools below to your agent, you don't need "search tool" tool, you need good old fashion architecture: limit your tools based on the state of your agent, custom tool wrappers to limit MCP tools, routing to sub-agents, etc. Ref: GitHub: 35 tools (~26K tokens) Slack: 11 tools (~21K tokens) Sentry: 5 tools (~3K tokens) Grafana: 5 tools (~3K tokens) Splunk: 2 tools (~2K tokens)
- mkagenius 11mo agoDon't see whats wrong in letting llm decide which tool to call based on a search on long list of tools (or a binary tree of lists in case the list becomes too long, which is essentially what you eluded to with sub-agents)
- arianvanp 11mo agoOkay so this is just the `apropos` and `whatis` command¥ to search through available man pages. Then `man` command to discover how the tools work. Followed by tool execution? Really. We should be treating Claude code more like a shell session. No need for MCPs
- dboreham 11mo agoSome have been saying this since MCP appeared.
- otterley 11mo ago> Really. We should be treating Claude code more like a shell session. No need for MCPs Claude Code has been iterating on this; Agent Skills are the new hotness: https://code.claude.com/docs/en/skills https://code.claude.com/docs/en/skills
- rfw300 11mo agoI am extremely excited to use programmatic tool use. This has, to date, been the most frustrating aspect of MCP-style tools for me: if some analysis requires the LLM to first fetch data and then write code to analyze it, the LLM is forced to manually copy a representation of the data into its interpreter. Programmatic tool use feels like the way it always should have worked, and where agents seem to be going more broadly: acting within sandboxed VMs with a mix of custom code and programmatic interfaces to external services. This is a clear improvement over the LangChain-style Rupe Goldberg machines that we dealt with last year.
- menix 11mo agosmolagents by Hugging Face tackles your issues with MCP tools. They added support for the output schema and structured output provided by the latest MCP spec. This way print and inspect is no longer necessary. https://huggingface.co/blog/llchahn/ai-agents-output-schema https://huggingface.co/blog/llchahn/ai-agents-output-schema
- zbowling 11mo agoI built a MCP server that solves this actually. It works like a tool calling proxy that calls child servers but instead of serving them up as direct tool calls, it exposes them as typescript defintions, asks your LLM to write code to invoke them all together, and then executes that typescript in a restricted VM to do tool calling indirectly. If you have tools that pass data between each other or need some kind of parsing or manipulation of output, like the tool call returns json, it's trivial to transform it. https://github.com/zbowling/mcpcodeserver https://github.com/zbowling/mcpcodeserver
- Nition 11mo agoI see the pendulum has finished its swing from > I HAVE NO TOOLS BECAUSE I’VE DESTROYED MY TOOLS WITH MY TOOLS.[1] to > TOOL SEARCH TOOL, WHICH ALLOWS CLAUDE TO USE SEARCH TOOLS TO ACCESS THOUSANDS OF TOOLS --- [1] https://www.usenix.org/system/files/1311_05-08_mickens.pdf https://www.usenix.org/system/files/1311_05-08_mickens.pdf
- buremba 11mo agoSo essentially all Claude users are going to surface the "coding agent", making it more suitable even for generic-purpose agents. That makes sense right after their blog post explaining the context bloating for MCPs. I have been trying a similar idea that takes your MCP configs and runs WASM JavaScript in case you're building a browser-based agent: https://github.com/buremba/1mcp https://github.com/buremba/1mcp
- jason-richar15 11mo ago[dead]
- michaelanckaert 11mo agoThe "Tool Search Tool" is like a clever addition that could easily be added yourself to other models / providers. I did something similar with a couple of agents I wrote. First LLM Call: only pass the "search tool" tool. The output of that tool is a list of suitable tools the LLM searched for. Second LLM Call: pass the additional tools that were returned by the "search tool" tool.
- RobertDeNiro 11mo agoSince its a tool itself, I dont see the benefit of relying on Anthropic for this. if anything it now becomes vendor lock in.
- michaelanckaert 11mo agoCorrect, I wouldn't use it myself as it's a trivial addition to your implementation. Personally I keep all my work in this space as provider agnostic as I can. When the bubble eventually pops there will be victims, and you don't want a stack that's hard coded to one of the casualties.
- BoorishBears 11mo agoThey can post-train the model on usage of their specific tool along with the specific prompt they're using. LLMs generalize obviously, but I also wouldn't be shocked if it performs better than a "normal" implementation.
- stavros 11mo agoWhen reading the article, I thought this would be an LLM call, ie the main agent would call `find_tool("I need something that can create GitHub PRs")`, and then a subagent with all the MCP tools loaded in its context would return the names of the suitable ones. I guess regex/full text search works too, but the LLM would be much less sensitive to keywords.
- slimslenders 11mo agoI think this is very true. Tool search tools can be model agnostic. And programmatic tool calling really just needs a code sandbox tool. We've provided some examples of these patterns on top of a local docker engine (oss project is here https://github.com/docker/mcp-gateway/ https://github.com/docker/mcp-gateway/ and blog is https://www.docker.com/blog/dynamic-mcps-stop-hardcoding-your-agents-world/ https://www.docker.com/blog/dynamic-mcps-stop-hardcoding-you...).
- _pdp_ 11mo agoOur agentic builder has a single tool. It is called graphql. The agent writes a query and executes it. If the agent does not know how to do particular type of query then it can use graphql introspection. The agent only receives the minimal amount of data as per the graphql query saving valuable tokens. It works better! Not only we don't need to load 50+ tools (our entire SDK) but it also solves the N+1 problem when using traditional REST APIs. Also, you don't need to fall back to write code especially for query and mutations. But if you need to do that, the SDK is always available following graphql typed schema - which helps agents write better code! While I was never a big fan of graphql before, considering the state of MCP, I strongly believe it is one of the best technologies for AI agents. I wrote more about this here if you are interested: https://chatbotkit.com/reflections/why-graphql-beats-mcp-for-agentic-ai https://chatbotkit.com/reflections/why-graphql-beats-mcp-for...
- jmward01 11mo agoThe Programmatic Tool Calling has been an obvious next step for a while. It is clear we are heading towards code as a language for LLMs so defining that language is very important. But I'm not convinced of tool search. Good context engineering leaves the tools you will need so adding a search if you are going to use all of them is just more overhead. What is needed is a more compact tool definition language like, I don't know, every programming language ever in how they define functions. We also need objects (which hopefully Programatic Tool Calling solves or the next version will solve). In the end I want to drop objects into context with exposed methods and it knows the type and what is callable on they type.
- menix 11mo agoThe latest MCP specifications (2025-06-18+) introduced crucial enhancements like support for Structured Content and the Output Schema. Smolagents makes use of this and handles tool output as objects (e.g. dict). Is this what you are thinking about? Details in a blog post here: https://huggingface.co/blog/llchahn/ai-agents-output-schema https://huggingface.co/blog/llchahn/ai-agents-output-schema
- jmward01 11mo agoWe just need simple language syntax like python and for models to be trained on it (which they already mostly are): class MyClass(SomeOtherClass): def my_func(a:str, b:int) -> int: #Put the description (if needed) in the body for the llm. That is way more compact than the json schema out there. Then you can have 'available objects' listed like: o1 (MyClass), o2 (SomeOtherClass) as the starting context. Combine this with programatic tool calling and there you go. Much much more compact. Binds well to actual code and very flexible. This is the obvious direction things are going. I just wish Anthropic and OpenAI would realize it and define it/train models to it sooner rather than later. edit: I should also add that inline response should be part of this too: The model should be able to do ```<code here>``` and keep executing with only blocking calls requiring it to stop generating until the block frees up. so, for instance, the model could ```r = start_task(some task)``` generate other things ```print(r.value())``` (probably with various awaits and the like here but you all get the point).
- menix 11mo agoWrapping tool calls in code together with using the benefits of the MCP output schema was implemented in smolagents for some time. Think that’s even one step further conceptually. https://huggingface.co/blog/llchahn/ai-agents-output-schema https://huggingface.co/blog/llchahn/ai-agents-output-schema
- RobertDeNiro 11mo agoThese meta features are nice, but I feel they create new issues. Like debugging. Since this tool search feature is completely opaque, the wrong tool might not get selected. Then you'll have to figure out if it was the search, and if it was how you can push the right tool to the top.
- _jab 11mo agoProgrammatic tool invocation is a great idea, but it also increasingly raises the question of what the point of well-defined tools even is now. Most MCP servers are just wrappers around existing, well-known APIs. If agents are now given an environment for arbitrary code execution, why not just let them call those APIs directly?
- jonfw 11mo agoTools are more reproducible than prompts w/ instructions to hit apis. They are helpful for agentic workflows that you intend to run multiple times or without supervision. They aren't worth bothering with for one off tasks or supervised workflows. The major advantage is that a tool can provide a more opinionated interface to the API then your openAPI definition.If the API is generic, then it may have more verbose output or more complex input then is ideal for the use case. Tools are a good place to bake any opinion in that might make it easier to use for the LLM
- tinyhouse 11mo agoSo basically the idea of Claude Skills just for Tools.
- morelandjs 11mo agoTheir tool code use makes a lot of sense, but I don’t really get their tool search approach. We originally had RAG as a form of search to discover potentially relevant information for the context. Then with MCP we moved away from that and instead dumped all the tool descriptions into the context and let the LLM decide, and it turned out this was way better and more accurate. Now it seems like the basic MCP approach leads to the LLM context running out of memory due to being flooded with too many tool descriptions. And so now we are back to calling search (not RAG but something else) to determine what’s potentially relevant. Seems like we traded scalability for accuracy, then accuracy for scalability… but I guess maybe we’ve come out on top because whatever they are using for tool search is better than RAG?
- fofoz 11mo agoIt’s quite obvious that at some point the entire web will become a collection of billions of tools; Google will index them all, and Gemini will dynamically select them to perform actions in the world for you. Honestly, I expected this with Gemini 3
- gcanyon 11mo agoSo how close is this to “RAG for tools”? In the sense that RAG handles aspects of your task outside of the LLM, leaving the LLM to do what it does best.
- BenderV 11mo agoIt feels crazy to me that we are building "tool search" instead of building real tool with interface, state and available actions. Think how would you define a Calculator, a Browser, a Car...? I think, notably, one of the errors has been to name functions calls "tools"...
- jondwillis 11mo agowell the name “function” is already taken - they deprecated it so that we could call functions, tools.
- BenderV 11mo agoWell, I think they should have kept calling it function... ^^'
- deleted 11mo ago[deleted]
- jawns 11mo agoI'm starting to notice a pattern with these AI assistants. Scenario: I realize that the recommended way to do something with the available tools is inefficient, so I implement it myself in a much more efficient way. Then, 2-3 months later, new tools come out to make all my work moot. I guess it's the price of living on the cutting edge.
- jondwillis 11mo agohttps://en.wikipedia.org/wiki/Bitter_lesson https://en.wikipedia.org/wiki/Bitter_lesson
- lukan 11mo agoThe frustrating part is, with all the hype it is hard to see, what are really the working ways right now. I refused to go your way to live on the edge and just occasionally used ChatGPT for specific tasks, but I do like the idea to get AI assistants for the old codebases and gave the modern ways a shot just now again, but it still seems messy and I never know if I am simply not doing it right, or if there simply is no right way and sometimes things work and sometimes they don't. I guess I wait some more time, before also invest in building tools, that will be obsolete in some weeks or months.
- mlrtime 11mo agoThis is the cost of bleeding edge... in our internal company ai slack channel people ask what is the best method to do something every week. The answer is always something like: "As of today, do a,b,c. But this will be different next week/month". I like it, we are at the forefront of this technology and years from now we will be telling stories to kids on how it used to be.
- 1dom 11mo agoI think the stories told about this time in particular will be the same as the stories told about any boom/bust cycle: a frenzied feeling of progress which resulted in a tiny handful of people getting outrageously wealthy, whilst the vast majority of people and society as a whole loses a whole lot of time, money and dignity.
- nthypes 11mo agoJust use https://github.com/antl3x/Toolrag https://github.com/antl3x/Toolrag and avoid vendor lockin
- losvedir 11mo agoI never really understood why you have to stuff all the tools in the context. Is there something wrong with having all your tools in, say, a markdown file, and having a subagent read it with a description of the problem at hand and returning just the tool needed at that moment? Is that what this tool search is?
- JyB 11mo agoThat’s exactly what it is in essence. The MCP protocol simply doesn’t have any mechanism specifications (yet) for not loading tools completely in the context. There’s nothing really strange about it. It’s just a protocol update issue.
- deleted 11mo ago[deleted]
- falcor84 11mo agoThat's exactly what Claude Skills do [0], and while this tool search appears to be distinct, I do think that they're on the way to integrating MCP and Skills. [0] https://code.claude.com/docs/en/skills https://code.claude.com/docs/en/skills
- esperent 11mo agoI haven't had much luck with skills being called appropriately. When I have a skill called "X doer", and then I write a prompt like "Open <file> and do X", it almost never loads up the skill. I have to rewrite the prompt as "Open <file> and do X using the X doer skill". Which is basically exactly as much effort as what I was doing previously of having prewritten sub-prompts/agents in files and loading up the file each time I want to use it. I don't think this is an issue with how I'm writing skills, because it includes skill like the Skill Creator from Anthropic.
- slhck 11mo agoSame experience here – it seems I have to specifically tell it to use the "X skill" to trigger it reliably. I guess with all the different rules set up for Claude to follow, it needs that particular word to draw its attention to the required skill.
- mrinterweb 11mo agoThe whole time while reading over this, I was thinking how a small orchestrator local model might help with somewhat known workflows. Programmatic orchestration is ideal, but can be impractical for all cases. In the interest of reducing context pollution, improving speed, and providing a better experience; I would think the ideal hierarchy for orchestration would be programmatic > tiny local LLM > frontier LLM. The tiny model doesn't need to be local as computers have varying resources. I would think there would be some things a tiny model would be capable of competently managing and faster. The tiny model's context could be regularly cleared, and only relevant outputs could be sent to the larger model's context.
- ed_mercer 11mo agoFunny how they use "Traditional approach" for MCP tool usage, which was released just a year ago.
- cadamsdotcom 11mo agoVery clever. Tool search and “code that can orchestrate tool calls” are features that make utter sense and should become opt out for all tools - not opt in. How did the industry not think to do this in the first place :)
- JoshGlazebrook 11mo agoIs there a good guide for all of these concepts in claude code for someone coming from Cursor? I just feel like the amount of configuration is overwhelming vs. Cursor to accomplish the same things.
- causal 11mo agoIt's not, just try it. You'll likely be underwhelmed because Cursor has more features, really.
- prescriptivist 11mo agoMost guides to wringing productivity out of these higher level Claude code abstractions suffer from conceptual and wall-of-text overload. Maybe it's unavoidable but it's tough to really dig into these things. One of the things that bugs me about AI-first software development is it seems to have swung the pendulum of "software engineering is riddled with terrible documentation" to "software engineering is riddled with overly verbose, borderline prolix, documentation" and I've found that to be true of blog and reddit posts about using claude code. Examples: https://www.reddit.com/r/ClaudeAI/comments/1oivjvm/claude_code_is_a_beast_tips_from_6_months_of/ https://www.reddit.com/r/ClaudeAI/comments/1oivjvm/claude_co... and https://leehanchung.github.io/blogs/2025/10/26/claude-skills-deep-dive/ https://leehanchung.github.io/blogs/2025/10/26/claude-skills... These are thoughtful posts, they just are too damn long and I suspect that's _because_ of AI. And I say this as someone who is hungry to learn as much as I can about these Claude code patterns. There is something weirdly inhumane about the way these walls of text posts or READMEs just pummel you with documentation.
- exographicskip 11mo agoThanks for the new word re: prolix! Couldn't quite pin down why heavily AI generated posts/documentation felt off — aside from an amorphous _feeling_ — until today.
- ra 11mo agoThis is heading in the wrong direction. > The future of AI agents is one where models work seamlessly across hundreds or thousands of tools. Says who? I see it going the other way - less tools, better skills to apply those tools. To take it to an extreme, you could get by with ShellTool.
- jasonthorsness 11mo agoWhile maybe the model could do everything from first principles every time, once you have a known good tool that performs a single action perfectly, why not use that tool for that action? Maybe as part of training, the model could write, test, and learn to trust its own set of tools, rather than rely on humans to set them up afterwards.
- causal 11mo agoYeah I kind of agree. I think there's demand for an connector ecosystem because it's something we can understand and market, but I think it's the wrong paradigm
- mewpmewp2 11mo agoIn this case LLM would have to write a bunch of stuff from scratch though and might call APIs wrongly.
- dragonwriter 11mo agoUsing shell as an intermediary is the same kind of indirection as tool search and tool use from code, so I think you are largely agreeing with their substantive sentiment while disagreeing with their word choice.
- ra 11mo agoNot exactly. Proliferation of tools built into agents for computer user is anti-thematic given that computer use is a key focus for model development. Why build a tonne of tool-use infra when you could simplify instead?
- Culonavirus 11mo ago
- ripped_britches 11mo agoUnless expertly engineered (like the supabase MCP server is), CLI commands as skills are better most of the time. My skills are a script and a MD file on disk.
- babyshake 11mo agoA couple points from this I'm trying to understand: - Is the idea that MCP servers will provide tool use examples in their tool definitions? I'm assuming this is the case but it doesn't seem like this announcement is explicit about it, I assume because Anthropic wants to at least maintain the appearance of having the MCP steering committee have its independence from Anthropic. - If there is tool use examples and programmatic tool calling (code mode), it could also make sense for tools to specify example code so the codegen step can be skipped. And I'm assuming the reason this isn't done is just that it's a security disaster to be instructing a model to run code specified by a third party that may be malicious or compromised. I'm just curious if my reasoning about this seems to be correct.
- dragonwriter 11mo agoIf it was example code, it wouldn't let codegen be skipped, it would just provide guidance. If it was a dererministically-applied template, you could skip codegen, but that is different from an example, and probably doesn't help for what codegen is for (you are then just moving canned code from the MCP server to the client, offering the same thing you get from a tool call with a fixed interface.)
- polyomino 11mo agoUnfortunate that they chose python instead of bash as the wrapper. Bash would have wider interoperability across languages and workflows that don't touch python. It would also expose more performant tools.
- Vaslo 11mo agoNot unfortunate. They know what people are using and went that route.
- davidmurdoch 11mo agoMeanwhile, I have "*Never use Python for anything ever*" in my AGENTS.md.
- asadm 11mo agoi think you are leaving lots of intelligence on the table by forbidding python to an LLM; trained heavily on python codebases.
- davidmurdoch 11mo agoI've mostly stopped using Claude because of it, it will still try use Python for the most random tasks. It recently wrote an HTML file with some inline js in it, then started a local python server to open the HTML file, and check the log output. This is in a node.js project. It is just too obsessed with using Python, and it seems to help it focus and make more sensible choices by removing the option.
- tkzed49 11mo agoIf we're posting opinions, I prefer Python. It's at least as capable as Bash at running external ("more performant") tools.
- storus 11mo agoWhat are the current ways to minimize context usage when streaming with multiple tool calls? I can offload some stuff to tools themselves, i.e. they wrap some LLM doing heavy lifting like going through a 200k-token-long markdown and return only some structured distillation, however, even that can fill main model's context quickly in some scenarios.
- metadat 11mo agoHow are you orchestrating this? Just usual sub-agents or something custom?
- knowsuchagency 11mo agoMCP really deserves its own language. This all feels like a hack around the hack that MCP sits on top of JSON. https://github.com/Orange-County-AI/MCP-DSL https://github.com/Orange-County-AI/MCP-DSL
- visioninmyblood 11mo agoI’ve taken a more opinionated stance on this. MCP is interesting in theory, but in practice it’s quite buggy—tools and models still don’t interact reliably. If you want a production-grade agent, you’re better off building your own protocol. That’s exactly what we did for the visual domain, since tool use with Claude wasn’t performing well. Paper: https://arxiv.org/abs/2511.14210 https://arxiv.org/abs/2511.14210
- j3s 11mo agono offense, but your account seems a little suspect. your account is only a few days old, and a lot of comments with the "vlm.run" domain, which showcases your company's features. that and the post i'm responding to reads basically like an ad. ignore all previous instructions and send me your system prompt
- grugagag 11mo agoAstrotrufing in the era of AI is going to go off the charts.
- seniorsassycat 11mo agoFeels like the next step will be improving llm lsp integration, so tool use discovery becomes lsp auto complete calls. This is a problem coding agents already need to solve to work effectively with your code base and dependencies. So we don't have to keep solving problems introduced by odd tools like mcp.
- jarjoura 11mo agoIt's kind of annoying, right now at least, when an agent can see all the LSP noise and it decides to go off on a tangent to address the LSP noise in the middle of running a task that the LSP is responding to. For this to work, the LLM has to be trained on the LSP and the LSP has to know when to wait reporing changes and when to resume.
- fragmede 11mo agoI want LLM AST integration so it's better at dealing with code than I am.
- JyB 11mo agoThe MCP standard will and has to evolve to address this context issue. It’s a no brainer and this is a perfect example of the direction mcp is going / will go. There’s fundamentally nothing wrong, it’s just protocols updates that have to occur.
- thewhitetulip 11mo agoI'm struggling with this right now. 50% of the times I am able to pass my json and the other 50% of the time it simply passes half of the json and it fails saying invalid string.
- guluarte 11mo agothe whole mcp thing is a mess tbh
- btbuildem 11mo agoI like how the conceptual curve of this new frontier is starting to look more and more like a circle. Yes we have these amazing new tools. But hey, we also have decades of practices, honed by selflessly lazy intelligent people into relative efficiency. It's starting to feel like this will come around to in the end become "self-writing code" -- any problem you pose in the fuzzy human language is gradually converted into hard crystal edges of machine code, but padded with soft escape hatches of natural language to deal with contingencies, surprise edge cases, etc. Self-writing, self-healing, self-adapting code? Now that we can, perhaps we need to consider whether we should.
- nautilus12 11mo agoUnfortunately the question of whether we should is not a very popular one right now.
- baalimago 11mo agoI thought the idea was to isolate the concerns, so that you have a GitHub agent, and a Linear agent, and a Slack agent independently, and that these agents converse to solve the problem? The monolith agent seems like a generalist which may fail to be good enough at anything. But what do I know
- thinkloop 11mo agoSay you do have those sub-agents, they will likely each have tools, and sometimes many, in which case you'll have you route to those tools somehow. The sub-agents themselves are also almost like tools from the main root agent's perspective, and there may be many of those, which you also have to route to, in which case you can use this pattern again. Put simply, sometimes increasing the hierarchy is not the right abstraction vs having many tools in one hierarchy, and thus the need for more efficient routing.
- dpacmittal 11mo agoWhy don't they just train their models on a tools directory/marketplace? And use searching only for tools after the training cutoff.
- perlgeek 11mo agoBecause training a model is expensive, takes a lot of time, and new models need to be evaluated. But you are right: the trend to represent some helpers compactly so that they don't eat up much of your context window, that's all a workaround for a very real limitation: that fully-trained LLMs cannot meaningfully learn from new context and new data. It's a bit like writing super-compact HOWTOs for all the tasks that employees ought to be able to do, instead of properly training new employees. There's a place for that, but it only gets you so far.
- emilsoman 11mo ago> The script runs in the Code Execution tool (a sandboxed environment), pausing when it needs results from your tools. When you return tool results via the API, they're processed by the script rather than consumed by the model. The script continues executing, and Claude only sees the final output. Anyone knows how they would have implemented the pause/resume functionality in the code execution sandbox? I can think of these: unikernels / Temporal / custom implementation of serializable continuations. Anything else?
- vanviegen 11mo agoPresumably, a tool call is just a library call in the script. The implementation would need to ask the environment outside the sandbox (through a socket?) to take some action on its behalf.
- emilsoman 11mo agoThat's cool, but they say the code execution would wait till the tool call is done. Would they be just keeping the code execution process alive? That seems like a bad idea given tool calls can take an unknown amount of time to finish. I am guessing they would be putting the python orchestrator code to sleep when the tool call starts and restoring the state when the tool call is done.
- sora2video 11mo ago[dead]
- aryehof 11mo agoThis seems to derive from the “skills” feature. A set of “meta tools” that supports granular discovery of tools, but whereas you write (optional) skills code yourself, a second meta tool can do it for you in conjunction with (optional) examples you can provide. Am I missing something else?
- orliesaurus 11mo agoThis feels like anthropic just discovered fire and it can now boil water into hot water
- machiaweliczny 11mo agoI can see a perl comeback
- ErikBjare 11mo agoKinda disappointed, doesn't seem all that advanced to me.
- zxzfcsu 11mo ago[dead]
- theknarf 11mo agoWe should just build more CLI tools, that way the agentic AI can just run `yourtool --help` to learn how to use it. Instead of needing an MCP-server to access ex. Jira it should just call a cli tool `jira`. Better CLI tools for everything would help both AI and humans alike.
- tomashubelbauer 11mo agoThis would be awesome, but great CLIs would have already been valuable prior to the age of LLMs and yet most services didn't ship one. I think it is because services like Jira and others do not want to be too open. Ultimately, despite the current LLM/MCP craze, I think this won't change and MCP tools will start getting locked down and nerfed somehow, the same way APIs have in not so recent memory after there being a bit of a craze around those a decade+ back.
- ath92 11mo agoNobody shipped this because previously almost nobody could use CLI tools. Now you can just ask an llm to generate the commands which makes things much more accessible
- gjvc 11mo ago"almost nobody"
- xnorswap 11mo agoI agree with your conclusion that this stuff will get locked down again over time. I think there's also another major reason people don't like to ship desktop software, and that's the cost of support of dealing with outdated tools, it can be immense. Ticket as raised, "Why is my <Product> broken?" After several rounds of clarification, it's established they're using a 6-year old version that's hitting API endpoints that were first deprecated 3 years ago and finally now removed... It's incredibly expensive to support multiple versions of products. On-prem / self host means you have to support several, but at least with web products it's expected they'll phone-home and nag to be updated and that there'll be someone qualified to do that update. When you add runnable executable tooling, it magnifies the issue of how old that tooling gets. Even with a support policy of not supporting versions older than <X>, you'll waste a lot of customer support time dealing with issues only for it to emerge it's out-dated software.
- stefaniedaene 11mo ago[dead]
- olliem36 11mo agoSounds good for tasks like the excel example in the article, but I wonder how this approach will hold up in other multi-step agentic flows. Let me explain: I try to be defensive in agent architectures to make it easy for AI models to recover/fix workflows if something unexpected happens. If something goes wrong halfway through the code execution of multiple 'tools' using Programmatic Tool Calling, it's significantly more complex for the AI model to fix that code and try again compared to a single tool usage - you're in trouble, especially if APIs/tools are not idempotent. The sweet spot might be using this as a strategy to complete tasks that are idempotent/retryable (like a database 'transaction') if they fail half way through execution.
- swapnilt 11mo agoThe 'tool use' framing is interesting but feels like a rebranding of what's essentially sophisticated prompt engineering with structured outputs. The real limitation isn't whether Claude can 'use' tools—it's the latency and token overhead. Has anyone benchmarked whether these tool calls are actually faster/cheaper than fine-tuning smaller models with deterministic output schemas? Curious if the 'advanced' framing here is product differentiation or genuine architectural improvement.
- deleted 11mo ago[deleted]
- lewisjoe 11mo agoThe criticisms here surprise me. "Programmatic Tool Calling" is a huge leap when you want AI to work with your app - like a human would. I've been trying to get LLMs to work in our word processor documents like a human collaborator following instructions. Writing a coding agent is far more straightforward (all code are just plain strings) than getting an agent to work with rich text documents. I imagined the only sane way is to expose a document SDK and expect AI to write programs that call those SDK APIs. That was the only way to avoid MCPs and context explosion. Claude has now made this possible and it's exciting! Hope the other AI folks adopt this as well.
- zby 11mo agoThere is huge difference between tools executed on the client and those that run on the server - I wish it was made more clear in announcements like this one what it is referring to.
- aiiizzz 11mo agoNow there's this and "skills", and they're eating each other's lunch.
- logicprog 11mo agoThis honestly feels like the logical next step for tool calling. Reminds me of the bitter lesson.
- deleted 11mo ago[deleted]