7 ms·
Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
- screm 28d agoHey! Disclaimer: I am a Co-Founder of Armature (YC P26) which sells growth services to dev tools. This study is part of our broader work on how to influence coding agents choices and get products picked. To understand how agents pick tools we measured close to 17k sessions on an environment where agents run exactly like in the real world, on various repositories, talking to different personas (vibe-coder, junior or senior engineers) in different sizes of companies. All the results are now public and we'd love to know what findings surprise you the most, here are a few we found interesting: - Claude Code rarely searches the web while Codex almost always does it and Cursor sits in the middle. - Coding agents disagree more frequently than they agree. - Some players (LangChain, Supabase, Netlify, Paypal, Adyen) are almost always mentioned in their categories but never chosen. - Modifying repository context can change the pick entirely. If you feel like digging, all the traces are there and we probably missed interesting learnings so let us know what you find!
- vivifkjo 28d agoIs there a way to force the usage of a tool for certain tasks? Example: alawys use cli "foobar" to retrieve weather starus. By tool I mean mcp server, cli, etc.
- screm 28d agoNot sure I got your question right but if you are wondering for your own coding agent then I guess the answer would be a skill? Here what I meant by "how to influence coding agents choices and get products picked" is from a vendor PoV, making sure any developer x codebase in the world asking for a tool in your category gets your tool recommended and implemented by the coding agent.
- watusername 28d agoObviously you can ask the agent to use a specific tool, but the point of this article is about what they choose when the human on the other end has no opinion/taste/clue.
- screm 28d agoExactly!
- josephg 28d agoThe data was cool. Then I tried to tap on one of the other tabs. “This content is easier to read while full screen!” - Ok I’m game. “Hey this is what makes armature special!” - I don’t care, I’m here to look at data, not onboard onto some random platform. It took ages to find the tiny “skip tour” button, hiding in black on black text. Then it gave me another popup, which I dismissed without reading. Then the full screen modal was visible but it was horizontally misaligned - the left edge was cut off and the right of my phone screen was all white. I closed the tab with great prejudice. (Safari on iOS if you wanna try reproducing it) I’m sure you - or Claude - built something you’re proud of. But I left your website frustrated.
- screm 28d agoHey, thanks for the feedback, the leaderboards aren't displaying well on mobile indeed, we are currently shipping a fix that should help with that. Thanks anyway!
- kouteiheika 28d agoFWIW I had the same reaction to the popups. Immediately closed the tab.
- hbarka 28d ago> Claude Code rarely searches the web while Codex almost always does I’m trying to understand why they are opposite. I think it is true, I find myself giving a secondary prompt to Claude to “research this” and only then will it fetch. Codex is bang on fetching already.
- edoceo 28d agoGemini CLI (at least mine) does web research all the time. I've noticed it hitting my own pages (I have to ask very specific things). I don't have any global or project rules to encourage that behavior.
- antonvs 28d agoThat’s a deliberate effort on Google’s part. Integrating AI and search is obviously pretty critical to their business.
- arcanemachiner 28d agoMust be system prompts and tool instructions guiding the agents differently.
- LunaSea 28d agoOpenAI is close to Microsoft so I assume that they have preferential and cheap access to the Bing search index. Claude is independent. Gemini should have Google search.
- 42piratas 28d agoMore likely the harness than the model or a search deal. In Claude Code, WebFetch prompts for permission per domain and WebSearch is a separately gated tool, so the cheapest path for the agent is almost always the files already in the repo. Codex has no equivalent friction in front of a fetch. Worth checking against your own permission settings before reading it as a model preference: allowlist a domain and the same agent will reach for it a lot more.
- ai_critic 28d agoHi. I appreciate that you need to make rent, but if your business is basically "we do growth hacking and SEO tricks on models and get them to use products that aren't actually best for the job", you are scum. You are perpetuating shitty practices that have hurt developers for years now. Part of the reason people use AI is because of how useless search is due to the previous generation doing the same kind of thing you propose. Please do something else with your life.
- N_Lens 28d agoFirst time?
- screm 28d ago"we do growth hacking and SEO tricks on models and get them to use products that aren't actually best for the job" -> Well this could be seen the other way around. Today, without proper promotion of services, only incumbents / leaders that are in the models priors (from their training data) are getting chosen. This is ultimately favoring the big generalist players and not the newer or more tailored solutions that benefit from less exposure. I truly think there is something to be done to improve developers' experience too!
- jdw64 28d agoLooking at this, maybe in the future, the tools that AI prefers will become the mainstream. Even now, the tools that AI gives the highest priority to are the ones people already choose. There might be a concentration effect toward the tools that AI selects
- screm 28d agoDefinitely! But about concentration I'm not so sure, there are ways to counter this effect so in the end it will be a fight like SEO is today. What is certain though is that getting recommended by coding agents will be a top prio for all dev tools.
- ex-aws-dude 28d agoIn the future: "I went ahead and built the database you requested using today's tool sponsor: Firebase"
- screm 28d agoYeah sounds kind of like the equivalent of SEA for AI agents (AEA?) except that it’s sneakier since agents can act without you noticing.. anyway this is in the hands of the labs
- pupppet 28d agoUgh...this is totally going to happen.
- Onavo 28d agoNot going to name names but it's already happening with the frontier labs as a revenue source.
- folkrav 28d agoI feel like you legitimately cannot say it for legal reasons, but I wish we would just name these freaking things.
- paulhebert 28d agoDidn’t OpenAI claim they’re going to make a billion dollars off ads in a year or something?
- drivingmenuts 28d agoI smell a money-making opportunity.
- screm 28d agoHaha there is a lot at stake for sure
- bartools_app 28d ago[flagged]
- hbarka 28d agoRedshift for databases ain’t even here. This is suspect.
- screm 28d agoWhy? We haven't benchmarked Data Warehouses yet, only prod databases where you wouldn't expect Redshift to be considered.
- ai_critic 28d agoCan we not encourage the same strip-mining and ad and SEO bullshit that previously ruined the last decade+ of the Internet? A large portion of the utility of AI is the barren ad-driven growth-hacked hellscape search has become. Don't encourage the next generation of these businesses, I beg of everyone.
- edoceo 28d agoLarge amounts of capital want this to happen? Outside of boycott, what else can be done. What can man do against such reckless ~hate~ money?
- folkrav 28d agoIt is probably already happening, and will get worse. OpenAI was boasting about their advertising revenue mere days ago. If we felt like we couldn't trust AI because of slop, soon we won't be able to trust it because it'll push whatever pays them to do it.
- DrewADesign 28d agoAnd while these sponsorship shenanigans are the tech business’s bread and butter, sponsored answer manipulation seems fundamentally more insidious. Even in a larger-scale measurement like this one, there’s no way to tell if any of that is sponsored, legitimately good recommendations, or the technical flavor of the goblins problem.
- screm 28d ago[dead]
- ttul 28d agoI built this for my own company. Armature is on to something. You start by analyzing the choices agents would make for various use cases and then glean what, if anything, you might do to start tilting the agents in the direction of your own product and away from the competitor. Selling to agents is similar to selling to humans. You dump money into marketing to make sure agents find your solution around every corner for every use case you’re well suited to.
- Freedom2 28d ago[dead]
- tiffanyh 28d agoDoesn’t this ignore that the future of ads will probably just be some type of affiliate revenue going back to the agent for any product they help recommend.
- appplication 28d agoMaybe but that future isn’t now and there’s real money to be made today with the above strategy.
- screm 28d ago[dead]
- hexapus 28d agoWell that's a horrifying thought. Thanks, I hate it.
- paidx 28d ago[flagged]
- akurilin 28d agoReally liked that "Go Full Screen" as a modal flow, surprisingly intuitive.
- screm 28d agoThanks, was considering killing it after getting the opposite feedback earlier, now I may keep both options!
- Ozzie-D 28d ago[flagged]
- luciana1u 28d ago[flagged]
- 12390asdjkas 28d ago[dead]
- thedreammachine 28d agoI've been tracking the same for a few months. All open source and available here: https://preseason.ai/ https://preseason.ai/
- negativefactori 28d agonice!
- IgorPartola 28d agoFor some reason Claude Code keeps using awk, sed, and even Python to do basic file editing. Anyone know why that changed with the 5 series?
- rcfox 28d agoI've noticed that too. Maybe the normal Write tool has to output the entire file and this is an attempt to reduce token usage?
- pjm331 28d agoThe Edit tool has been notoriously tricky to get right - it seems they have maybe branched out but I think morphllm started specifically with the pitch that they trained a small model to be good at editing files - most of their testimonials are about that But I think it’s mostly a solved problem in frontier models and the bash tool usage is more likely an attempt to be more token efficient - I’ve noticed it used for making mechanical bulk edits that would be numerous “edit” tool uses otherwise
- troupo 28d ago> this is an attempt to reduce token usage? Wouldn't generating a Python script to edit files waste more tokens than using the built-in tool?
- sznio 28d agodepends on the breadth of the edit. anything that involves multiple files might be better done with python (e.g. renaming a function, along with changing all call sites.)
- ryeguy 28d agoWhat does this question have to do with the linked article?
- Neywiny 27d agoBecause they're thinking like I did going into the article. Harnesses like Claude expose "tools" to the agent. I usually use Cline but I'm giving up on it for this exact reason. Cline tells the model "you tell me to write a file, I'll get it done" and then it messes everything up, causes tones of errors, and the model goes "wow that's a broken tool. I'm going to write a python script to write the file instead"
- harisingh1612 28d ago[flagged]
- taikhoom 28d ago[flagged]
- elzbardico 28d agoI remember when the SEO nightmare started, it looked like an innocent intelectual investigation exercise like this.
- screm 28d agoWe don't need to reproduce the same errors! Anyway I do think it is going to be different this time because generating content is becoming so easy today that the entire web would just become 99.99% slop very quickly if things don't change. That being said the solution isn't that trivial, curious if you have thoughts on it?
- elzbardico 28d agoYoutube at least, is already 99% for new content, outside the channels you already subscribed. The amount of channels that are basically and AI animation with a voice-over of a curious wikipedia article is staggering, but those are the good ones. The worse part is a deluge of stupid boomer fanfic (The ones about ungrateful kids, ungrateful employers that fired someone who secretly was a load-bearing (ha!) element for a contract, or HOA drama), self-help stuff, and red-pill incel fanfics.
- dgroshev 25d agoYou're talking to the CEO of the want-to-be-SEO company [1] that did that. It is very much content marketing for their services. [1]: https://news.ycombinator.com/item?id=49557233 https://news.ycombinator.com/item?id=49557233
- screm 23d agoIndeed, we are helping growth teams get picked by coding agents, this is not a secret. But I do think this has the potential of getting 1000x worse than SEO so I don't imagine people letting that happen. For example we just added code review here: https://armature.tech/leaderboards#app/code-review https://armature.tech/leaderboards#app/code-review -> See my comment here about how much coding agents choose themselves. I don't expect people to let this last forever for instance, otherwise it's worrying for literally any software out there.
- nijave 28d agoAzure database??? In house bot protection??? In house search??? Some of these are absolutely wild. Surprised Strands didn't even get mentioned for agent frameworks.
- screm 28d agoAzure database was mostly for enterprise use-cases. Rebuilding a lot of things in-house is a real trend, especially for Claude Code when you don't ask it explicitly to consider all solutions and avoid overhead of managing things yourself. Codex and Cursor seem to have this in mind more naturally (at least using GPT-5.6 Sol / Grok 4.6)
- natnatenathan 28d agoI keep telling people that we are living in the golden age of AI - like the first year or two of google. It is all down hill as these companies push for profit and lock-in.
- adamisom 28d agolocal models will be powerful enough, can't say the same about local search engines 20y ago
- nullbio 28d agoExactly why everyone needs to be hyper-focused on ensuring that the open-source ecosystem is healthy and that we don't let them shut that down.
- Roland303 28d agohow does one do that from their comfy chair?
- pixelatedindex 28d agoI for one am running what I can on my aging 1080Ti(!!), namely a Qwen2.5 14B (4-bit quantized). It’s not great, and the only other card I have is a 3070Ti but I need that for gaming. :( What a terrible card that 3070Ti is. Mad regrets buying it because I wanted to save $400 compared to a 3080Ti.
- pixelatedindex 27d agoI really don’t understand why this is being downvoted? Did I say something wrong / controversial? Or is it because I’m not using my resources the “optimal” way? Is it because my models aren’t open weight? I was only trying to contribute my experience :(
- oefrha 28d agoBy sending some cash (and maybe some non-sensitive training data too) the way of labs creating open models? At work we have expensive Anthropic and OpenAI subscriptions but also have some in house workflows plugged into DeepSeek and GLM APIs.
- lhk931122 28d agoI only get web search from Claude Code when I ask, only one of my last 8 sessions accessed the web at all. But they were the case of Opus here. Curious what Fable 5.1 model does instead of Opus.
- rf15 28d agoSo we build an LLM... that grabs new tokens based on stochastics... train it on all the programming teaching material and projects available on the internet... and then analyse the output... ...for the distribution of content of the source material? what? You learn nothing.
- small_scombrus 28d agoBecause of weightings around other things you aren't going to get a 1:1 50% of input code used node so it uses node 50% of the time. There's enough randomness and other stolen content to (in theory) bias it towards weird outputs/choices
- martypitt 28d agoI like the look of this - and it's a problem that we're thinking about right now at work. But, the pricing of this is .. really high .. - starting at $5k / month? I'd find that difficult to justify.
- screm 28d agoHey, thanks! I'm wondering if it's clear from our website that this is the price of a fully managed service, not just access to a platform or reports. Think of an SEO agency model.
- killix 28d ago[flagged]
- trimethylpurine 28d agoTell it what tools to use. Choosing an architecture is pretty important if you are going to lead a project. That's not the best part to skip, I don't think.
- xnorswap 28d agoThey didn't measure which programming language because we all know it's python, which it tries to use, every session, despite repeated memory files to not use python. Recently it's even taken to installing python to get jobs done.
- screm 28d ago[dead]
- olmo23 28d agoPerhaps it depends on your previously saved memories, because over here if given the choice it's always reaching for either TypeScript or C.
- tedmartynov 25d ago[flagged]
- alankritxghoshx 28d ago[dead]
- tuberreact 28d agoI appreciate the eval design here. And man Claude not doing web searches is killing me bc it just won't offer the most up to date information. At the same time, there's research saying Claude relies heavily on Brave search so I'm not sure how to reconcile
- screm 28d agoIf you want your Claude Code to search you can always tweak your own with a good skill, this should work perfectly! It's more a problem for vendors who can't tell all people in the world to download a specific skill first.
- runtime_lens 28d ago[dead]
- theootzen 28d agoWhich sandboxes do you yourself use to run those agents? And how did you choose this provider?
- screm 28d agoWe decided to use three different providers so we could verify that this choice doesn't impact the result of our experiments (E2B, Blaxel and Daytona)
- aaron_m04 28d agoI've learned so much awk from Claude!
- claud_ia 28d ago[flagged]
- OpenQuota 28d ago[flagged]
- agentislandpro 28d ago[flagged]
- ahmedelsama 28d ago[dead]
- JUEJINs 28d ago[flagged]
- anduril22 28d agoLM Studio or Unsloth are both great
- malinono 28d ago[flagged]
- dearlinko 28d ago[flagged]
- felixlu2026 28d ago[dead]
- saadyousfi 28d ago[flagged]
- deleted 24d ago[deleted]
- nixoda8041 24d ago[flagged]
- xfor 24d agoI suggested a new sector - "Code Review" - there's a growing landscape of apps trying to solve this in a new agentic paradigm
- coldbootHq 15d ago[flagged]