11 ms·
Composer: Building a fast frontier model with RL
- ibash 1y agoVery cool, congrats!
- romanovcode 1y agoWhere is the comparison with Sonnet 4.5? That would be the only thing that matters, really.
- matheist 1y ago> "Best Frontier" includes GPT-5 and Sonnet 4.5, which both outperform Composer.
- yodon 1y ago>> "Best Frontier" includes GPT-5 and Sonnet 4.5, which both outperform Composer. Looking at the graph, it would appear there's an implicit "today" in that statement, as they do appear poised to equal or surpass Sonnet 4.5 on that same benchmark in the near future.
- ciphix 1y agoWhat Cursor is really emphasizing here is speed — they’re claiming it runs about four times faster than GPT-5/Sonnet, while still offering roughly the same level of performance.
- timcobb 1y agoDoes anyone code with GPT-5? I've never had it work in Cursor. I mean, like, at all.
- srush 1y agoA lot of people use it! It scores very well on our benchmarks, significantly better than Composer-1.
- alyxya 1y agoI wonder if this custom model is trained on cursor users. There’s a lot of potential on how much better a custom model could be the closer it is integrated with the product. Having the model learn to adapt to different user preferences would make it stand out compared to memoryless frontier models.
- Sammi 1y agoThe fact that you are wondering this is bad. You definitely should know this. _ALL_ the online ai providers are training on your data. They have more expensive enterprise plans if want to opt out.
- alyxya 1y agoI’ve generally seen providers allow you to opt in or out. What may vary is what the default is and what they may offer in exchange for using your data (perhaps they could offer higher rate limits).
- nu11ptr 1y agoI love Cursor. I've tried Copilot/Claude/etc. but keep coming back to Cursor. I just want to work, and Cursor tab complete is dang accurate, esp. for refactoring tasks.
- Sammi 1y agoI tried going back to VS Code + Copilot a month ago. I only lasted 4 days because it was to bad. It was super slow and gave poor suggestions, but mostly it just flat out did not suggest anything. Cursor feels snappy in comparison and the suggestions are more often than not useful. The most annoying thing about Cursor tab complete, is that it is so fast that when I am doing something unusual then it will keep on jumping in with useless suggestions. They have a snooze function for this though.
- WanderPanda 1y agoDamn TIL, I always used > Cursor: disable completions and forgot to turn it on again I need to try snooze then!
- stared 1y agoWhile I am excited to see a new model, I am skeptical when there is so much vagueness - charts with "frontier models" without actually spelling out which ones, charts with no numbers (time axis, or in one chart - entirely).
- srush 1y agoThere is a footnote that should help with the models. Training is a harder thing to report on, but roughly our finding here is that RL scales.
- solarkraft 1y agoPeople on here love to be contrarian about Cursor, but I’ve tried all the popular alternatives (Copilot, Claude Code, Codex, Gemini CLI, Cline) and found Cursor’s overall experience to just be unmatched. A big part of that is its speed, another its reliability. It’s the only coding agent I’m actually really motivated to use out of the box because it really does make me feel more productive while the others keep messing up the project, from way too large changes I didn’t ask for all the way to constant syntax and request errors. It’s the only coding agent I’ve used that feels serious about being a product rather than a prototype. Their effort in improving their stack is totally paying off.
- pqdbr 1y agoI dropped cursor for the precise reason you mention: reliability. Countless times my requests in the AI chat just hang there for 30+ seconds more until I can retry them. When I decided to give Claude Code a try (I thought I didn't need it because I used Claude in Cursor) I couldn't believe how faster it was, and literally 100% reliable. EDIT: given today's release, decided to give it a go. The Composer1 model _is_ fast, but right at the second new agent I started I got this: > Connection failed. If the problem persists, please check your internet connection or VPN
- davidgomes 1y agoA lot of progress is being made here on the Cursor side I encourage you to try it again. (Cursor dev)
- cleak 1y agoThis is the exact reason I left Cursor for Claude Code. Night and day difference in reliability. The Windows experience might be especially bad, but it would get constantly hung or otherwise fail when trying to run commands. I also had to babysit Cursor and tell it to continue for mid sized tasks.
- jonasnelle 1y agoThey've improved performance dramatically in the last few weeks, might have fixed your issues.
- jasonjmcghee 1y agoMaybe I'm an outlier but Sonnet 4.5 quality is about as low as I'm willing to go. It's generation speed is not the problem or the time sink. It's wrestling with it to get the right output. --- And just to clarify as maybe I misunderstood again but people are comparing cursor to Claude Code and codex etc here- isn't this whole article all cursor just using different models?
- swyx 1y ago> Sonnet 4.5 quality is about as low as I'm willing to go. literally a 30 day old model and you've moved the "low" goalpost all the way there haha. funny how humans work
- vidarh 1y agoYes? Because why should we settle for less now that it is available?
- swyx 1y agobecause engineering is the art of "good enough" and composer is clearly "good enough but a lot faster" which makes up for intelligence gaps in interesting ways
- vidarh 1y agoIt's not good enough for a lot of us, though, clearly.
- tomashubelbauer 1y agoFor me the bar for barely good enough is and always has been Codex. Before I found frontier models more trouble than they're worth. And there is still a massive amount of room to grow before I can genuinely say working with these tools is more enjoyable than frustrating for me and now I use them (and how I think they should work).
- jasonjmcghee 1y ago
- jonasnelle 1y agoCursor has the best Tab model, and I feel like their lead there has kept growing - they're doing some really cool things there. https://cursor.com/blog/tab-rl https://cursor.com/blog/tab-rl I wonder how much the methods/systems/data transfer, if they can pull off the same with their agentic coding model that would be exciting.
- srush 1y agoWe also are big Tab users here at Cursor. In the blog we talk about the motivation for this project came from thinking about a Tab-like agent.
- dagss 1y agoIt's great. BUT: Wish they had selected another shortcut like shift+tab. Every time I write code myself I find myself racing the AI to get an indentation in before the AI is done... gets annoying
- RosalieCodes 1y agoYou can change the key bind, I personally set it to ctrl+tab
- vidarh 1y agoI feel like that's like having a lead in producing better buggy whips. I run Claude Code in the background near constantly for a variety of projects, with --dangerously-skip-permissions, and review progress periodically. Tabbing is only relevant when it's totally failing to make progress and I have to manually intervene, and that to me is a failure scenario that is happening less and less often.
- lubujackson 1y agoThis is just a completely different use of LLMs and has little to do with working at a real business with a live site and users. Cursor is great when you want to gain understanding of an issue quickly, or resolve something clear and specific quickly. I'm not against YOLO vibe coding, but being against tab completion is just insane to me. At the end of the day, LLMs help you achieve goals quicker. You still need to know what goal you want to achieve, and tab completion basically let's me complete a focused goal nearly as soon as I determine what my goal is.
- kilroy123 1y agoWhat I can't stand about cursor is the constantly changing and confusing billing and usage. I think competition in the space is a good thing, but I'm very skeptical their model will outperform Claude.
- srush 1y agoHi everyone, I am an ML researcher at Cursor, and worked on this project. Would love to hear any feedback you may have on the model, and can answer question about the blog post.
- alyxya 1y agoIs the new model trained from scratch? What training data went into it?
- dfltr 1y agoIs it true that Cheetah is Grok Code Fast 2? Does this mean that the new Cursor model is also based on Grok?
- srush 1y agoCheetah was an earlier (and dumber) version of this model that we used to test production speed. They are both developed in-house. If you liked Cheetah, give this model a try.
- dfltr 1y agoAwesome, thanks for the clarification. So are the rumors around Cheetah being based on a Grok model just straight up untrue? I want to try Composer but have a pretty strict no X/Grok policy.
- srush 1y agoStraight up untrue.
- carlosbaraza 1y agoThis is nice. I liked Cheetah for grunt work that I want to get out quickly and is not too hard. The speed is really awesome. A model that would run at even higher speeds like the OSS models at groq/cerebras would really be workflow changing, because the slowness of SOTA models really breaks the flow. I find myself taking a ton of breaks and getting distracted while I wait for a model to complete a task (e.g. just now).
- swyx 1y agosee also https://cursor.com/changelog/2-0 https://cursor.com/changelog/2-0 and https://cursor.com/blog/2-0 https://cursor.com/blog/2-0 other links across the web: https://x.com/amanrsanger/status/1983581288755032320?s=46 https://x.com/amanrsanger/status/1983581288755032320?s=46 https://x.com/cursor_ai/status/1983567619946147967?s=46 https://x.com/cursor_ai/status/1983567619946147967?s=46
- swyx 1y agomy very small nit is... why is the model called Composer?? of all things?? when there was already a Cursor Composer from 2024. Cursor Cheetah wouldve been amazing. reusing the Composer name feels like the reverse OpenAI Codex move haha
- srush 1y agoWe like the name Composer and were sad to see it go. Excited to bring it back. (Agree Cheetah is a cool name too.)
- OsrsNeedsf2P 1y agoOne thing no competitor is serious on is average response completion time. Cursor lapped everyone there
- srush 1y agoThere are lots of good models we like here. But we agree that getting the right point on the smart+fast graph can make agentic coding feel really good. (Cursor researcher)
- 80hd 1y agoInsane velocity from the Cursor team. I wonder how they move so fast?
- numbers 1y agoPlease keep the naming of your models sane, I'd like to know that composer 1 is the first model and composer 2 is second but composer 1o is not yet another 1 variant that's actually newer and better than 2, that's just dumb. Not that you're doing that, some other companies do that.
- srush 1y agoWe will do our best. Luckily I don't think there are major telecom companies called Composer-2.
- Sander_Marechal 11mo agoThere is also a very polular package manager called Composer. Do companies not search for name collisions? Or do they squat on community projects on purpose?
- carlosbaraza 1y agoCursor 2.0 keeps crashing on me while having an agent running and opening the IDE part of the application. I might have to rollback.
- amilich 1y agoHey - really sorry to hear this - could you email me andrew@cursor.com? Here are 3 suggestions to try- 1. Reset your settings.json - if shared with vscode, sometimes settings can cause perf regressions 2. Could you try cmd-shift-p -> "capture and send debugging data"? Will send us some profiling data to debug 3. Clear your user data (will delete chats) as a last resort - cmd-shift-p, "reveal user data," close the app, then delete this folder and restart the app
- carlosbaraza 1y agoCould anyone explain how to use multiple agents and subagents in Cursor, Claude Code, or others? It is already challenging to me taming one model doing work, let alone synchronizing multiple parallel workers. Do you have to split the plan in parallelizable tasks that could be worked in parallel in one codebase without breaking and confusing the other agents?
- asdev 1y agoyou can use git worktrees and just have multiple Claude Code terminal instances working on each worktree. That way they don't clash, just delete the worktree when the task is done.
- carlosbaraza 1y agoI have never leveraged git worktrees... That is such a crazy useful tool that I am almost ashamed of not having researched it before. Git is such a beautiful piece of software.
- asdev 1y agoI built an open source project to make the whole workflow easier: https://github.com/built-by-as/FleetCode https://github.com/built-by-as/FleetCode
- asdev 1y agois Cursor Bench open? Would like to see an open benchmark for agentic coding
- srush 1y agoUnfortunately not, as we used our own internal code for the benchmark. We would also like to see more benchmarks that reflect the day-to-day agentic coding use.
- gabriel666smith 1y agoIs there any information at all available, anywhere, on what Cursor Bench is testing and how? It's the most prominent part of the release post - but it's really hard to understand what exactly it's saying.
- srush 1y agoRoughly, we had Cursor software engineers record real questions they were asking models, and then had them record the PR that they made that contained the result. We then cleaned these up. That is the benchmark.
- ukblewis 1y agoWhich programming languages/tools/libraries did the teams questions/code involve?
- gabriel666smith 1y agoAre you able to give a sense of how many questions, which domains they were split over, and how that split looked in % terms? As a user, I want to know - when an improvement is claimed - whether it’s relevant to the work I do or not. And whether that claim was tested in a reasonable way. These products aren’t just expensive - it requires switching your whole workflow. Which is becoming an increasingly big ask in this space. It’s pretty important for me to be able to understand, and subsequently, believe a benchmark - I find it really hard not to read it as ad copy where this information isn’t present.
- neuronexmachina 1y agoFor anyone else who was wondering, it looks like the within-Cursor model pricing for Cursor Composer is identical to gemini-2.5-pro, gpt-5, and gpt-5-codex: https://cursor.com/docs/models#model-pricing https://cursor.com/docs/models#model-pricing ($1.25 input, $1.25 cache write, $0.13 cache read, and $10 output per million tokens)
- lubujackson 1y agoI'm curious if their near-term expectation is that this is be better than these models or is this a model they tend to use in Auto mode, or if the focus is really if you want speed...? I guess my question is why would I actively chose this over Auto?
- cwyers 1y agoThe lack of transparency here is wild. They aggregate the scores of the models they test against, which obscures the performance. They only release results on their own internal benchmark that they won't release. They talk about RL training but they don't discuss anything else about how the model was trained, including if they did their own pre-training or fine-tuned an existing model. I'm skeptical of basically everything claimed here until either they share more details or someone is able to interpedently benchmark this.
- criemen 1y agoI understand where you're coming from, and I'd love to have learned about pre-training vs. off-the-shelf base model too. But > their own internal benchmark that they won't release If they'd release their internal benchmark suite, it'd make it into the training set of about every LLM, which from a strictly scientific standpoint, invalidates all conclusions drawn from that benchmark from then on. On the other hand, not releasing the benchmark means they could've hand-picked the datapoints to favor them. It's a problem that can't be resolved unfortunately.
- nickpsecurity 1y agoIn high-security systems, we solved this problem with trusted, independent evaluators who got all the data. They replicate the results themselves. They analyze every artifact for flaws. They also pen test the system offensively. If they say it's good, then maybe it is good or maybe less, obviously bad. We could have third-party groups with evaluation criteria who don't make models or sell A.I.. Strictly evaluators. Alternatively, they have a different type of steady income with the only A.I. work they're doing being evaluation.
- cwyers 1y agoI'm not saying SWE-Bench is perfect, and there are reports that suggest there is some contamination of training sets for LLMs with common benchmarks like SWE-Bench. But they publish SWE-bench so anyone can run it and have an open leaderboard where they attribute the results to specific models, not just vague groupings: https://www.swebench.com/ https://www.swebench.com/ ARC-AGI-2 keeps a private set of questions to prevent LLM contamination, but they have a public set of training and eval questions so that people can both evaluate their modesl before submitting to ARC-AGI and so that people can evalute what the benchmark is measuring: https://github.com/arcprize/ARC-AGI-2 https://github.com/arcprize/ARC-AGI-2 Cursor is not alone in the field in having to deal with issues of benchmark contamination. Cursor is an outlier in sharing so little when proposing a new benchmark while also not showing performance in the industry standard benchmarks. Without a bigger effort to show what the benchmark is and how other models perform, I think the utility of this benchmark is limited at best.
- timcobb 1y agoI wish it was easy to find out how much it costs relative to Claude :)
- skeptrune 1y agoFacts. They really need to make pricing more clear across the entire product.
- sebdufbeau 1y agoAs a stealth model, it was priced as $1.25M in / $10M out Right now, it seems free when you are a Cursor Pro user, but I'd love more clarity on how much it will cost (I can't believe it'll be unlimited usage for subscribers)
- Jcampuzano2 1y agoA bit late but it's actually not free. You can see it on their models page. It's similarly priced to GPT-5 and Gemini 2.5 Pro. https://cursor.com/docs/models#model-pricing https://cursor.com/docs/models#model-pricing
- Jayakumark 1y agoThis looks like a model RLed on top of Qwen3-Coder or GLM 4.6 as per their graph and foot note.
- netcraft 1y agoI love cursor, the tab completion and agent mode. But I really dislike vscode after using intellij for so many years. I really wish the underlying editor was better, or I could get cursor features in intellij instead. The editing of the files is mostly fine, but its everything else around it that a full IDE provides thats just so much better. Right now its intellij + claude code for me, and its fine, but I wish I could get the AI power of cursor in a better package.
- Jcampuzano2 1y agoBuilding off of VSCode was probably Cursors silver bullet and the best decision they could have ever made. It made migrating for everyone using VSCode (probably the single most popular editor) or another vscode forked editor (but at the time it was basically all VSCode) as simple as install and import settings. I do not think Cursor would have done nearly as well as it has if it didn't. So even though it can be subpar in some areas due to VSCodes baggage, its probably staying that way for a while.
- netcraft 1y agoI dont disagree with anything you said. If I was in their shoes, I would have done exactly the same thing. Maybe my complaint is that I wish vscode had more features like intellij, or that intellij was the open source baseline a lot of other things could be built on. Intellij is not without its cruft and problems, dont get me wrong. But its git integration, search, navigation, database tools - I could go on - all of these features are just so much nicer than what vscode offers.
- pbowyer 1y agoIntellij's tab-complete is coming along; it's hit and miss if it will work but for similar edits I'm finding it picks up the pattern quickly and I can tab - tab - tab to make them happen. Still not up to Cursor standards though :)
- DefineOutside 1y agoI find Cursor's tab completion to be distracting enough with multi-line changes that I just disabled it, while I use IntelliJ's tab completion regularly. Cursor's tab completion is better, but it doesn't seem to have a concept of not trying to tab complete. IntelliJ is correct half the time for completing the rest of the line and only suggests when it is somewhat confident in its answer.
- simonw 1y agoHere's the Composer 1 pelican riding a bicycle: https://static.simonwillison.net/static/2025/cursor-1-pelican.png https://static.simonwillison.net/static/2025/cursor-1-pelica...
- arresin 1y agoSame price as GPT-5
- SafeDusk 1y agoI think both Cursor and Cognition and going in the same direction of SWE-grep[0]. SWE-grep was able to hit ~700tokens/s and Cursor ~300token/s, hard to compare the precision/recall and cost effectiveness though, considering SWE-grep also adopted a "hack" of running it on Cerebras. I'm trying to kickstart a RL-based code search project called "op-grep" here[1], still pretty early, but looking for collaborators! [0]: https://cognition.ai/blog/swe-grep https://cognition.ai/blog/swe-grep [1]: https://github.com/aperoc/op-grep https://github.com/aperoc/op-grep
- koakuma-chan 1y agoI just gave it a try and it's reaally fast. Didn't expect this from you Cursor, good job.
- toobulkeh 1y agoI used the new system tonight and it felt like a definite downgrade. Generated a few non-working basic apps, couldn’t handle CSS in a NextJS environment. Terminal context didn’t work. And it went back to not reasoning through the problem until resolution. And kept slowing down. I’m assuming major release vs stable, but this is pretty lackluster so far. Switched back to Sonnet reasoning. Here’s to improving!
- ianberdin 1y agoFeels like the comments are fighting of prepaid influencers.
- ciphix 1y agoThe metrics in the post seem quite abstract. Does anyone know the detailed metrics of this mysterious model? Was it fine-tuned from open models or trained from scratch?