8 ms·
Apideck CLI – An AI-agent interface with much lower context consumption than MCP
- gertjandewilde 7mo agoWe built a unified API with a large surface area and ran into a problem when building our MCP server: tool definitions alone burned 50,000+ tokens before the agent touched a single user message. The fix that worked for us was giving agents a CLI instead. ~80 tokens in the system prompt, progressive discovery through --help, and permission enforcement baked into the binary rather than prompts. The post covers the benchmarks (Scalekit's 75-run comparison showed 4-32x token overhead for MCP vs CLI), the architecture, and an honest section on where CLIs fall short (streaming, delegated auth, distribution).
- OsrsNeedsf2P 7mo agoHow is progressive discovery not more expensive due to the increased number of steps?
- zamalek 7mo agoIn short: JSON. Plan prose or markdown is way more token efficient than JSON. I think that responding in JSON was always a mistake in the spec; it should have been free-form text (which could then be JSON if required).
- iamjackg 7mo agoIt depends on what your "currency" is: inference cost vs. models getting dumber/slower with a fuller context.
- BeefySwain 7mo agoI assume because the discovery is branching. If the an agent using the CLI for for GitHub needs to make an issue, it can check the help message for the issue sub-command and go from there, doesn't need to know anything about pull requests, or pipelines, or account configuration, etc, so it doesn't query those subcommands. Compare this to an MCP, where my understanding is that the entire API usage is injected into the context.
- vesselapi 7mo ago[dead]
- lelanthran 7mo ago> How is progressive discovery not more expensive due to the increased number of steps? Why not run the discovery (whether MCP or CLI) in a subagent that returns only the relevant tools. I mean, discovery can be done on a local model, right?
- hparadiz 7mo ago10 years from now: "Can you believe they did anything with such a small context window?"
- mbreese 7mo ago10 years from now: “what’s a context window?”
- sghiassy 7mo ago10 years from now: “come with me if you want to live” Terminator 2 Clip: https://youtu.be/XTzTkRU6mRY?t=72&si=dmfLNDqpDZosSP4M https://youtu.be/XTzTkRU6mRY?t=72&si=dmfLNDqpDZosSP4M
- berziunas 7mo ago“640K ought to be enough for anybody”
- this_user 7mo agoMore likely: "Can you believe they were actually trying to use LLMs for this?"
- rib3ye 7mo agoOSes and software engs did not end up using less RAM.
- 7mo ago
- sim04ful 7mo agoGraphql introspection queries would be a really neat application for LLM calls
- zamalek 7mo agoWhat I've done with my MCPs is turning them into a CLI, except there's still an MCP server that only has the instructions to tell the the agent about the CLI.[1] Claude and GLM-5 seem to have no problems with it. As a bonus, the entire thing now works as a plain old CLI too - which it honestly should have from the beginning. [1]: https://github.com/jcdickinson/ferrisfetch/blob/main/cmd/mcp_prelude.md https://github.com/jcdickinson/ferrisfetch/blob/main/cmd/mcp...
- patates 7mo agoYou don't need a whole server to tell agents that, I think you can just write a skill file or two and be done with it.
- zamalek 7mo agoThat would work, and is still an option. However I think this makes deployment/installation simpler (which is also why I write my MCPs in go).
- antihero 7mo agoCan the MCP tell them how to use the CLI? Surely that would mean less time wasted on discovering it each time. Going to try this with fastmail-cli and see what happens.
- zamalek 7mo agoYes, MCP instructions are a blob that is injected at the start of the context. That file is prefixed to the agent specific help[1] for the instructions (that is returned if --help is invoked with CLAUDE=1 or AGENT=1). [1]: https://github.com/jcdickinson/ferrisfetch/blob/main/cmd/agent_help.md https://github.com/jcdickinson/ferrisfetch/blob/main/cmd/age...
- antihero 7mo agoAh cool, that said: If your tool is primarily for Claude Desktop (or mobile if hosted), that surely needs the MCP to actually do anything right?
- austinhutch 7mo ago> Not a protocol error, not a bad tool call. The connection never completed. Very interesting topic, but this LLM structure is instant anthema I just have to stop reading once I smell it.
- leontloveless 7mo ago[flagged]
- nicoritschel 7mo agoWhile I generally prefer CLI over MCP locally, this is bad outdated information. The major harnesses like Claude Code + Codex have had tool search for months now.
- injidup 7mo agoCan you explain how to take advantage. Is there any specific info from anthropic with regards to context window size and not having to care about MCP?
- amzil 7mo agoFair point on tool search. Claude Code and Codex do have it. But tool search is solving the symptom, not the cause. You still pay the per-tool token cost for every tool the search returns. And you've added a search step (with its own latency and token cost) before every tool call. With a CLI, the agent runs `--help` and gets 50-200 tokens of exactly what it needs. No search index, no ranking, no middleware. The binary is the registry. Tool search makes MCP workable. CLIs make the search unnecessary.
- cruffle_duffle 7mo agoLet me guess the command: [error] Wait, better check help. is it -h? [error] Nope? Lemme try —-help. [error] Nope. How about just “help” [error] Let me search the web [tons of context and tool calls]
- binarymax 7mo agoWhy not use skills? They follow a three-tier loading approach, and you can stick an MCP as part of the toolset for the skills, so it will only load it when the skill is selected. See the progressive disclosure section in the skills docs: https://agentskills.io/what-are-skills https://agentskills.io/what-are-skills
- forrestthewoods 7mo agoThree tiers? I thought it was two?
- binarymax 7mo agoDiscovery, Activation, Execution as per the linked doc
- caust1c 7mo agoI'm getting tired of everyone saying "MCP is dead, use CLIs!". Yes, MCP eats up context windows, but agents can also be smarter about how they load the MCP context in the first place, using similar strategy to skills. The problem with tossing it out entirely is that it leaves a lot more questions for handling security. When using skills, there's no implicit way to be able to apply policies in the sane way across many different servers. MCP gives us a registry such that we can enforce MCP chain policies, i.e. no doing web search after viewing financials. Doing the same with skills is not possible in a programatic and deterministic way. There needs to be a middle ground instead of throwing out MCP entirely.
- mvrckhckr 7mo agoI agree, and it's context-dependent when to use what (the author mentions use cases for other solutions). I'm glad there are multiple solutions to choose from.
- j45 7mo agoMCPs are handy in their place. Agents calling CLI locally is much more efficient.
- yoyohello13 7mo agoIt is a weird trend. I see the appeal of Skills over MCP when you are just a solo dev doing your work. MCP is incredibly useful in an organization context when you need to add controls and process. Both are useful. I feel like the anti-MCP push is coming from people who don't need to work in a large org.
- krzyk 7mo agoNot sure. Our big org, banned MCPs because they are unsafe, and they have no way to enforce only certain MCPs (in github copilot).
- thenewnewguy 7mo agoBut skills where you tell the LLM to shell out to some random command are safe? I'm not sure I understand the logic.
- rirze 7mo agoAt this point, I feel like MCP servers are just not feasible at the current level of context windows and LLMs. Good idea, but we're way too early.
- bkummel 7mo agoThere's already an open source tool that does exactly the same thing: https://github.com/knowsuchagency/mcp2cli https://github.com/knowsuchagency/mcp2cli
- amzil 7mo agoGreat tool, however we went to a dedicated CLI client (think gh, aws, stripe) in Go.
- deleted 7mo ago[deleted]
- robutsume 7mo ago[dead]
- kristjansson 7mo agoCLIs are great for some applications! But 'progressive disclosure' means more mistakes to be corrected and more round trips to the model - every time[1] you use the tool in a new thread. You're trading latency for lower cost/more free context. That might be great! But it might not be, and the opposite trade (more money/less context for lower latency) makes a lot of sense for some applications. esp. if the 'more money' part can be amortized over lots of users by keeping the tool definitions block cached. [1]: one might say 'of course you can just add details about the CLI to the prompt' ... which reinvents MCP in an ad hoc underspecified non-portable mode in your prompt.
- amzil 7mo agoThis is a fair trade-off and the post should probably be more explicit about it. You're right that progressive disclosure trades latency for cost and context space. For some workloads that's the wrong trade. The amortization point is interesting too. If you're running a support agent that calls the same 5 tools thousands of times a day, paying the schema cost once and caching it makes total sense. The post covers this in the "tightly scoped, high-frequency tools" section but your framing of it as a caching problem is cleaner. On the footnote: guilty as charged, partially. The ~80 token prompt is a minimal bootstrap, not a full schema. It tells the agent how to discover, not what to call. But yeah, the moment you start expanding that prompt with specific flags and patterns, you're drifting toward a hand-rolled tool definition. The difference is where you stop. 80 tokens of "here's how to explore" is different from 10,000 tokens of "here's everything you might ever need." But the line between the two is blurrier than the post implies. Fair point.
- machinecontrol 7mo agoThe trend is obviously towards larger and larger context windows. We moved from 200K to 1M tokens being standard just this year. This might be a complete non issue in 6 months.
- amzil 7mo agoContext windows getting bigger doesn't make the economics go away. Tokens still cost money. 50K tokens of schemas at 1M context is the same dollar cost as 50K tokens at 200K context, you just have more room left over. The pattern with every resource expansion is the same: usage scales to fill it. Bigger windows mean more integrations connected, not leaner ones. Progressive disclosure is cheaper at any window size.
- magospietato 7mo agoContext caching deals with a lot of the cost argument here.
- amzil 7mo agoIt helps with cost, agreed. But caching doesn't fix the other two problems. 1) Models get worse at reasoning as context fills up, cached or not. right? 2) Usage expansion problem still holds. Cheaper context means teams connect more services, not fewer. You cache 50K tokens of schemas today, then it's 200K tomorrow because you can "afford" it now. The bloat scales with the budget... Caching makes MCP more viable. It doesn't make loading 43 tool definitions for a task that uses two of them a good architecture.
- hrmtst93837 7mo ago[flagged]
- dend 7mo agoOne of the MCP Core Maintainers here, so take this with a boulder of salt if you're skeptical of my biases. The debate around "MCP vs. CLI" is somewhat pointless to me personally. Use whatever gets the job done. MCP is much more than just tool calling - it also happens to provide a set of consistent rails for an agent to follow. Besides, we as developers often forget that the things we build are also consumed by non-technical folks - I have no desire to teach my parents to install random CLIs to get things done instead of plugging a URI to a hosted MCP server with a well-defined impact radius. The entire security posture of "Install this CLI with access to everything on your box" terrifies me. The context window argument is also an agent harness challenge more than anything else - modern MCP clients do smart tool search that obviates the entire "I am sending the full list of tools back and forth" mode of operation. At this point it's just a trope that is repeated from blog post to blog post. This blog post too alludes to this and talks about the need for infrastructure to make it work, but it just isn't the case. It's a pattern that's being adopted broadly as we speak.
- o_____________o 7mo ago> modern MCP clients do smart tool search that obviates the entire "I am sending the full list of tools back and forth" mode of operation How, "Dynamic Tool Discovery"? Has this been codified anywhere? I've only see somewhat hacky implementations of this idea https://github.com/modelcontextprotocol/modelcontextprotocol/issues/1821#issuecomment-3709687415 https://github.com/modelcontextprotocol/modelcontextprotocol... Or are you talking about the pressure being on the client/harnesses as in, https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-search-tool#mcp-integration https://platform.claude.com/docs/en/agents-and-tools/tool-us...
- dend 7mo agoMore of the latter than the former. The protocol itself is constrained to a set of well-defined primitives, but clients can do a bunch of pre-processing before invoking any of them.
- amzil 7mo agoThe post isn't MCP vs CLI. It covers where MCP wins. > The entire security posture of "Install this CLI with access to everything on your box" terrifies me This is fair for hosted MCPs, However I'm not claiming the CLI is universally more secure. users needs to know what they're doing. Honestly though, after 20 years of this, the whole thread is debating the wrong layer. A well-designed API works through CLI, MCP, whatever. A bad one won't be saved by typed schemas. > At this point it's just a trope that is repeated from blog post to blog post Well, "Use whatever gets the job done" and "it's just a trope" can't both be true. If the CLI gets the job done for some use cases, it's not a trope. It's an option. And I'd argue what's happening is the opposite of a trope. Nobody's hyping CLIs because they're exciting. There's no protocol foundation, no spec committee, no ecosystem to sell into. CLIs are 40-year-old boring technology. When multiple teams independently reach for the boring tool, that's a signal, not a meme. > This blog post too alludes to this and talks about the need for infrastructure to make it work When tool search is baked into Claude Code, that's Anthropic building and maintaining the infrastructure for you. The search index, ranking, retrieval pipeline, caching. It didn't disappear. It moved. And it only works in clients that support it. Try using tool search from a custom Python agent, a bash script, or a CI/CD pipeline. You're back to loading everything. A CLI doesn't need the client to do anything special. `--help` works everywhere. That's the difference between infrastructure that's been abstracted away for some users and infrastructure that's genuinely not needed.
- ekropotin 7mo agoLet me guess - another article about how CLI s are superior to MCP?
- kayig 7mo agoI k know
- nzoschke 7mo agoThe industry is talking in circles here. All you need is "composability". UNIX solved this with files and pipes for data, and processes for compute. AI agents are solving this this with sub-agents for data, and "code execution" for compute. The UNIX approach is both technically correct and elegant, and what I strongly favor too. The agent + MCP approach is getting there. But not every harness has sub-agents, or their invocation is non-deterministic, which is where "MCP context bloat" happens. Source: building an small business agent at https://housecat.com/ https://housecat.com/. We do have APIs wrapped in MCP. But we only give the agent BASH, an CLI wrapper for the MCPs, and the ability to write code, and works great. "It's a UNIX system! I know this!"
- kayig 7mo ago[flagged]
- dirk94018 7mo agoUnix approach can be surprisingly powerful. https://linuxtoaster.com/blog/gradientdescentforcode.html https://linuxtoaster.com/blog/gradientdescentforcode.html
- ycombiredd 7mo agoWhat's interesting to me is that while it was obvious to all of us who came to think in the Unix Way, that insofar as composability, usage discoverability, and gobs of documentation in posts and man pages that are hugely represented in training corpora for LLMs, that the CLI is a great fit for LLM tool use, it seems only a recent trend to acknowledge this (and also the next hype wave, perhaps.) Also interesting that while the big vendors are following this trend and are now trying to take a lead in it, they still suggest things like "but use a JSON schema" (the linked article does a bit of the same - acknowledging that incremental learning via `--help` is useful AND can be token-conserving (exception being that if they already "know" the correct pattern, they wouldn't need to use tokens to learn it, so there is a potential trade-off), they are also suggesting that LLMs would prefer to receive argument knowledge in json rather than in plain language, even though the entire point of an LLM is for understand and create plain language. Seemed dubious to me, and a part of me wondered if that advice may be nonsense motivated by desire to sell more token use. I'm only partially kidding and I'm still dubious of the efficacy. * Here's a TL;DR for anyone who wants to skip the rest of this long message: I ran an LLM CLI eval in the form of a constructed CTF. Results and methodology are in the two links in the section linked: https://github.com/scottvr/jelp?tab=readme-ov-file#what-else https://github.com/scottvr/jelp?tab=readme-ov-file#what-else Anyhow... I had been experimenting with the idea of having --help output json when used by a machine, and came up with a simple module that exposes `--help` content as json, simply by adding a `--jelp` argument to any tool that already uses argparse. In the process, I started testing, to see if all this extra machine-readable content actually improved performance, what it did to token use, etc. While I was building out test, trying to settle on legitimate and fair ways to come to valid conclusions, I learned of the OpenCLI schema draft, so I altered my `jelp` output to fit that schema, and set about documenting the things I found lacking from the schema draft, meanwhile settling to include these arg-related items as metadata in the output. I'll get to the point. I just finished cleaning the output up enough to put it in a public repo, because my intent is to share my findings with the OpemCLI folks, in hopes that they'll consider the gaps in their schema compared to what's commonly in use, but at the same time, what came as a secondary thought in service of this little tool I called "jelp", is a benchmarking harness (and the first publishable results from it), the to me, are quite interesting and I would be happy if others found it to be and added to the existing test results with additional runs, models, or ideas for the harness, or criticism about the validity of the method, etc. The evaluation harness uses constructed CLI fixtures arranged as little CLI CTF's, where the LLMs demonstrate their ability to use an unknown CLI be capturing a "flag" that they'll need to discover by using the usage help, and a trail of learned arguments. My findings at first confirmed my intuitions, which was disappointing but unsurprising. When testing with GPT-4.1-mini, no manner of forcing them to receive info about the CLI via json was more effective than just letting them use the human-friendly plain English output of --help, and in all cases the JSON versions burned more tokens. I was able to elicit better performance by some measurements from 5.1-mini, but again the tradeoff was higher token burn. I'll link straight to the part of the README that shows one table of results, and contains links to the LLM CLI CTF part of the repo, as well as the generated report after the phase-1 runs; all the code to reproduce or run your own variation is there (as well as the code for the jelp module, if there is any interest, but it's the CLI CTF eval that I expect is more interesting to most.) https://github.com/scottvr/jelp?tab=readme-ov-file#what-else https://github.com/scottvr/jelp?tab=readme-ov-file#what-else
- m3kw9 7mo agoThe thing with CLIs is that you also need to return results efficiently. It if both MCP and CLI return results efficiently, CLI wins
- enraged_camel 7mo agoWith context windows starting to get much larger (see the recent 1M context size for Claude models), I think this will be a non-issue very soon.
- mihir_kanzariya 7mo ago[flagged]
- JohnMakin 7mo agoThis matches my experience building in-house MCP servers. The mechanism I prefer on load is something like a quick FTS5 with BM25 ranking lookup to find what it needs, and then serve those. I think a lot of these things are implemented pretty naively - for instance, we ran into the huge context problem with Jira, so we just built our own Jira MCP interface that doesn't have all the bloat. If the agent finds it needs something it doesnt have, it can ask again.
- esafak 7mo agoThis is becoming a solved problem with tool search; MCP is back.
- rob 7mo agoThe real issue isn't MCP, it's these fucking bots posting here every day.
- mritchie712 7mo agoclaude code solved this about a month ago
- Havoc 7mo agoGetting LLMs to reliably trigger CLI functions is quite hard in my experience though especially if it’s a custom tool
- leontloveless 7mo ago[flagged]
- drewbitt 7mo agohttps://github.com/RhysSullivan/executor https://github.com/RhysSullivan/executor
- robot-wrangler 7mo ago> Limit integrations → agent can only talk to a few services The idea that people see this as one horn of a trilemma instead of just good practice is a bit strange. Who would complain that every import isn't a star-import? Bring in what you need at first, then load new things dynamically with good semantics for cascade / drill-down. Let's maybe abandon simple classics like namespacing and the unix philsophy for the kitchen-sink approach after the kitchen-sink thing is shown to work.
- mt42or 7mo agoTired of this shit. Be less stupid.
- bazhand 7mo agoI ran into this exact problem building a MCP server. 85 tools in experimental mode, ~17k tokens just for the tool manifest before any work starts. The fix I (well Codex actually) landed on was toolset tiers (minimal/authoring/experimental) controlled by env var, plus phase-gating, now tools are registered but ~80% are "not connected" until you call _connect. The effective listed surface stays pretty small. Lazy loading basically, not a new concept for people here.
- TheTaytay 7mo agoI’m a huge fan of CLIs over MCP for many things, and I love asking Claude Code to take an API, wrap it in a CLI, and make a skill for me. The ergonomics for the agent and the human are fantastic, and you get all of the nice composability of command line Unix build in. However, MCPs have some really nice properties that CLIs generally don’t, or that are harder to solve for. Most notably, making API secrets available to the CLI, but not to the agent, is quite tricky. Even in this example, the options are env variables (which are a prompt injection away from dumping), or a credentials file (better, but still very much accessible to the agent if it were asked). MCPs give you a “standard” way of loading and configuring a set of tools/capabilities into a running MCP server (locally or remotely), outside of the agent’s process tree. This allows you to embed your secrets in the MCP server, via any method you choose, in a way that is difficult or impossible for the agent to dump even if it goes rogue. My efforts to replicate that secure setup for a CLI have either made things more complicated (using a different user for running CLIs so that you can rely upon Linux file permissions to hide secrets), or start to rhyme with MCP (a memory-resident socket server started before the CLI that the CLI can talk to, much like docker.sock or ssh-agent)
- agenticbtcio 7mo ago[dead]
- anesxvito 7mo ago[dead]
- kayig 7mo agoHey
- maxothex 7mo ago[dead]
- featwanz 7mo ago[dead]
- gpubridge 7mo ago[dead]
- diven_rastdus 7mo ago[dead]
- anvevoice 7mo ago[dead]
- ryan14975 7mo ago[dead]
- tacone 7mo agoActually what it seems to tackle at its core is discoverability. Which should be built in in each MCP server as it's not that difficult, instead, we see MCP servers with 50+ methods. Much easier: { action: 'help' } { action: 'projects.help' } { action: 'projects.get', payload: { id: xxxx-xx-x } } And you get the very same discoverability. There are other interesting capabilities though, like built in permissions based on HTTP verb, that might be useful to someone.
- kimi 7mo ago[dead]
- seongsukang 7mo ago[dead]
- BTAQA 7mo agoUsing MCP daily as a solo founder with Claude Code. The "consistent rails" point resonates. The value isn't just tool calling, it's that the agent knows how to behave within a defined boundary. The security posture argument is underrated too. Giving a CLI unrestricted box access vs a hosted MCP server with scoped permissions is a completely different risk profile.
- techcam 7mo agoWe ran into something similar with API costs — small changes in behavior can have surprisingly large downstream effects.
- dirk94018 7mo agoMCP is a bit of a rube goldberg machine. Unix solved that problem. Pipe text in, get text out, discover capabilities incrementally. The fact that we need benchmarks to prove CLIs use fewer tokens than dumping 55k of JSON schema upfront is embarrassing. toast is just pipe stuff to an LLM and let stdin/stdout be the protocol. No schema tax, no connection lifecycle, no tool registry middleware to manage your middleware. The thing is that AIs are just not good at outputting structure, like people, json isn't natural.
- dtraub 7mo ago"You'll end up with something like MCP once you introduce enterprise users" - yeah. The token efficiency debate is a single-developer optimization. The moment you introduce teams or compliance requirements, the question shifts to who manages the credentials. With CLI, it's your machine, your keys. With direct API calls, keys live wherever the agent runs. Both work until a contractor leaves and their laptop still has active keys for your repos, your internal docs, and your CRM. Remote MCP over streamable HTTP gives you a centralized auth layer. One SSO integration, one revocation point, one audit trail. I wrote about this angle here: https://dev.to/dennistraub/missing-from-the-mcp-debate-who-holds-the-keys-when-50-agents-access-50-apis-mb3 https://dev.to/dennistraub/missing-from-the-mcp-debate-who-h...
- amzil 7mo ago[dead]
- agentictrustkit 6mo agoI like that everyone keeps separating "capability" from "authority" because they get conflated in a lot of agent-centered tooling. CLI vs MCP choice mostly changes the HOW as a side effect. It doesn't answer the bigger question and probably harder one: who delegated the rigtht to cause that effect, for how long, and with what scope? Just like with people, you need a policy decision that's independent. It should be revocable and auditable. One way that I look at it is with these long-running agents should look less like a script and more like an employee. You wouldn't give them the master key hoping they behave well. You'd give specific access and in stages probably. That's what I think we're missing with our agents is giving them appropriate authority, delegated by an owner with a audit trail