10 ms·
Cursor Introduces Composer 2.5
https://twitter.com/cursor_ai/status/2056415413077233983 https://twitter.com/cursor_ai/status/2056415413077233983
- XCSme 5mo agoCan we use Composer 2.5 via API/OpenRouter?
- asar 5mo agoThe model is (like Composer 2) based on Kimi K2.5 and they claim SOTA performance for 1/10th of the cost. The tweet also mentions that they've started a new model from scratch on Colossus 2 (xAI/SpaceX Cluster). Really impressive how they've made this jump from being called the vscode fork with no moat just a couple of months ago.
- Lionga 5mo agoThey are still a vscode fork with no moat? Like they lost about 70% of users in half a year which goes to show how there is not even the tiniest of moat.
- GenerWork 5mo agoI feel like they've been targeting enterprise pretty hard. I know my company uses them, and the companies that hire us also use Cursor.
- kvetching 5mo agoCursor will definitely win the enterprise for coding. Enterprises aren't going to trust a TUI
- esafak 5mo agoWhy not? That makes no sense to me.
- Squarex 5mo agoAll enterprises I know use GitHub copilot as they already have Office, Teams, … wonder how will it change with the recent pricing changes
- pjmlp 5mo agoI can tell my company wants nothing with them.
- kilroy123 5mo agoI think it's going to be brutal for them to compete with OpenAI and Anthropic. I switched to claude code because of usage. For $200 a month, I would run out of usage halfway through the month. Then be forced to use their composer model or whatever slow, dumb model they served up in their "auto" mode. For that same $200 a month, I could use claude code and basically never hit usage limits. I don't understand what people are doing who run into the limits on that max x20 plan. I NEVER have.
- whywhywhywhy 5mo agoIt's still a VsCode fork just now with a Kimi fine tune and still no moat... I won't debate that it turns out none of this mattered when it came to being as successful company though and kinda makes anyone who tried to roll their own instead of fork look a little silly.
- hkleppe 5mo ago"No moat", well... How I see this is that its so important to bundle the model with the right tooling. Like a racecar, having the best engine doesn't help if the rest of the car lacks other winning properties (reliability, aerodynics etc). So for Cursor, which IMO, they put themself in a strong position by having both a solid IDE __and__ a solid+cost efficient model. Those two working great in combination for the task they are designed to solve (coding) is more important than benchmarks
- aurareturn 5mo agoI doubt it's a brand new model. It's likely just Kimi K2.5 further trained on coding.
- enraged_camel 5mo agoThey didn't say it's a new model... in fact they said exactly what you just said.
- liuliu 5mo agoSince the frontier is only 8-month ahead of DeepSeek, it is hard to see how model training can be a moat as all the tricks are available from open labs in China. You really just need <100m to bootstrap at this point.
- onlyrealcuzzo 5mo ago> Really impressive how they've made this jump from being called the vscode fork with no moat just a couple of months ago. Impressive, yes. But they still don't have a moat...
- kkukshtel 5mo agoAnd its still just a vscode fork
- icemelt8 5mo agoCursor 3 is a complete rewrite, its no longer a fork.
- gkbrk 5mo agoIt's still a VSCode fork. Even Cursor's own About window tells you it's VSCode. Cursor Version: 3.4.20 VSCode Version: 1.105.1
- muhfournik 5mo agoI believe the agent view is a complete rewrite, and maybe the other parts but not the editor itself
- alach11 5mo agoIsn't a large user base and the data collected from those users a moat of sorts?
- AussieWog93 5mo agoHonestly the data itself is probably worth heaps even in the company itself collapses. Early attention engineering when humans were still in the loop!!!
- NitpickLawyer 5mo ago> Early attention engineering when humans were still in the loop Exactly. Cursor was the first product used by tons of devs on real codebases. Just the signal "acceptance rate" is huge and can't be easily captured w/ synthetic data.
- deleted 5mo ago[deleted]
- wg0 5mo agoThis was the only way forward.
- antirez 5mo agoHow much the RL they are doing really improves Kimi K2.5 is to be seen. So, right now, the ground truth is that they combined what they had with a strong open weights model. The RL improvement may be both marginal (since may folks report strong results with vanilla K2.6) and may mostly bias the model towards coding tasks: when a model like this is trained to be generalist, there is a tension between being good at one thing and the other, in terms of SFT and RL. You can see this in the DeepSeek v4 Flash training report for instance but it is a known fact. So if you have the GPUs and a decent RL pipeline that does not run the model you can indeed specialize it a bit more for a given task at the expenses of tasks people will not do inside Cursor. But, so far, the measurable reality is that Cursor uses an open weight model like most could do, and the RL story could be partilly a marketing move to call to Composer 2.5 more than a real strong gain, given that there is no way to verify and K2.5 was already strong. And we also know that they had to partner to do the training, which is also not a good news.
- the_duke 5mo agoIn my opinion cursor actually has one of the best harnesses again at the moment.
- make3 5mo agowhy is that part impressive specifically? they got purchased by SpaceX, they have access to infinite compute and cash now. & now they're still losing all of their users to Claude Code and Codex.
- DeathArrow 5mo ago>& now they're still losing all of their users to Claude Code and Codex. Why pay for Cursor when I can use GLM 5.1, Kimi K2.6, MiniMax M2.7, Xiaomi MiMo V2.5 Pro and Deepseek v4 for cheap and use whatever harness I want, including Claude Code. It's not like Cursor harness is the best out there. And even if I want to edit the code, I don't need to run the agent harness in an IDE.
- make3 5mo agothese are in the trillion parameters range, not sure it's actually that cheap to have at a reasonable speed without quality degradation & without like.. your own DGX B200
- DeathArrow 5mo agoI didn't say to run them at home. There are some cheap coding plans that gets you plenty of usage for the Chinese models.
- wmichelin 5mo agoNot a cursor shill by any means, I do use it at work but that's because it's what they pay for. But Cursor has a CLI harness.
- DeathArrow 5mo ago>Really impressive how they've made this jump from being called the vscode fork with no moat just a couple of months ago. With so much money and computing from SpaceX, is not so impressive.
- farco12 5mo agoOne would hope the vscode fork with a $50B valuation and no moat, would wisely spend the money they raised to build a moat.
- everfrustrated 5mo agoFull details https://cursor.com/blog/composer-2-5 https://cursor.com/blog/composer-2-5
- dang 5mo agoThanks! Link belatedly changed above.
- scuderiaseb 5mo ago[dead]
- jdlyga 5mo agoIt's a bit odd that they're not comparing it against Sonnet
- jjice 5mo agoI don't think so. They're comparing it to the highest tier available models from Anthropic and OpenAI. Generally speaking, Opus is better than Sonnet in almost every way, so why have the redundancy?
- 3836293648 5mo agoPrice to performance?
- jjice 5mo agoI think their comparison to how their benchmarks compare to Opus are a great way to show "look at similar benchmarks for a fraction of the cost". If it has Opus benchmarks (I don't actually take benchmarks seriously, but for their comparison purposes) and Sonnet is still more than half the price of Opus, I figure it's close enough where it doesn't matter.
- CodingJeebus 5mo agoThe tweet specifies that the new model is geared towards long-running tasks, which is what you'd use a model like Opus for anyway.
- svclaws 5mo agoTheir previous Composer was already marketed as a cheap model capable of competing with SOTA on most tasks. The evals they shared back then backed this up but in my day-to-day usage it fell short across the board. Canceled my cursor subscription and switched to Claude Code a few weeks ago. It has its own shortcomings but in terms of model capability and UX quality Cursor will have a hard time competing in the long term. Elon Musk will be a very good way out for them.
- PUSH_AX 5mo agoThey set themselves up for flack when they use whatever these evals are… they did the same for composer 2 which was evaled in close competition with frontier models, spoiler alert, it wasn’t even close in practice. So now 2.5 is supposed to compete with opus 4.7? Sure…
- criemen 5mo agoWell is that a statement about the quality of Opus 4.7 or about compose 2.5? :P
- tuo-lei 5mo agothey say it themselves in the post - behavior dimensions "not well captured by existing benchmarks". that was the exact problem with composer 2. not dumber on individual tasks, just bad at session-level decisions like when to stop editing, how much context to carry forward, when to re-read a file vs assume. you don't catch any of that in an isolated eval.
- infecto 5mo agoAs I have said before in prior composer threads. The proof is in the usage. I am inclined to somewhat believe the results as I use composer and also take the results for the given context. It’s not a general purpose sota model. It’s a model that runs inexpensively in their coding workflow that is creating results similar to opus or gpt.
- jmcqk6 5mo agoThat does not match my experience. Composer 2 was fantastic for my uses, and I hit Composer 2.5 with some very difficult things last night, which it handled fast and effectively. I don't really care about benchmarks. I care about practice, and in practice, it's been very very good for me.
- sergiotapia 5mo agoCongratulations on the launch! I'm interested in trying Cursor but it's very confusing what I should buy. What does the Pro $20 plan get me in usage if I only use Composer 2.5? How fast is the model?
- darkwi11ow 5mo agoI use $20 plan on daily basis for more than a year now, and have yet to exhaust that limit. The plan includes $20 in api costs for non-Cursor premium models and $20 for Composer and Auto models provided by Cursor themselves. That said, I am pretty old-fashioned coder and use LLM mostly to overcome the blank page problem, which means I review and often rewrite LLM output by hand and avoid prompt loops for a single task. People who are aiming to not read code any more might find this $20 plan lacking for their needs, however for my needs it fits perfectly.
- kaizoku156 5mo agoThe limits are probably even higher than that, i seem to get about 100$+ of usage on composer and about 45-50 usd on non composer models
- ChrisArchitect 5mo agoNon-x link: https://cursor.com/blog/composer-2-5 https://cursor.com/blog/composer-2-5 (https://news.ycombinator.com/item?id=48182126 https://news.ycombinator.com/item?id=48182126)
- vanuatu 5mo agoIt's always great that more companies are throwing their hat in the ring, especially focusing on value (latency + intelligence + cost)
- re-thc 5mo agoDid they just upgrade Kimi 2.5 to 2.6?
- lukebrichey 5mo agostill uses 2.5
- jtwaleson 5mo agoOk this might be weird but I've moved everyone in my 4 person team to our team plan and costs seem to have sky rocketed compared to the individual plans. Where before most people spent 20-100 USD, now the total bill is more like 1k USD. I haven't gone into the details but it feels like I'm being scammed.
- PUSH_AX 5mo agoMy cursor costs sky rocketed recently too
- danbrooks 5mo agoCheck which model you're using. The fast version of composer is the default now (which costs ~x3 as much).
- infecto 5mo agoKeep in mind I believe there is a larger buffer given to personal plans. If they have 50% extra with the personal plan you now only get 25%.
- DedlySnek 5mo agoMy company is shifting us from Cursor to Claude due to increased costs.
- mohsen1 5mo agoWe moved off Cursor and onto Codex + Claude Code. Cost went from multiple thousand per engineer per month to about $500
- zackify 5mo agoBest deal currently: Cursor team Codex team Claude team Swap between the models when limited. I am saving our company a lot of money vs Claude enterprise usage cost
- skeptic_ai 5mo agoI did some monitoring. 15 accounts, 300 millions tokens input, 200k output went to 0 the 5h quota in 7 hours. 4 parallel tasks. I think 300 million is too low. For reference before I could do more than 1 billion on same conditions.
- lukebrichey 5mo agothis feels super bullish on cursor/spacexai's ability to train a frontier level model. could be truly SOTA on coding given that their RL data is this powerful
- granzymes 5mo agoSurprised this got pushed off the front page so quickly! It’s exciting to see what the Cursor team has been able to do with significantly fewer resources than the frontier labs. I do wish they weren’t joining xAI. Something tells me there will be a contingent of researchers that departs Cursor if that merger is consummated.
- dang 5mo agoIt set off the flamewar detector, a,k.a. the overheated discussion detector. We'll turn that off.
- granzymes 5mo agoThanks, dang! The blog post[1] might be a better source than the twitter thread. Also I regret my typo above (lab -> labs) but too late now! [1] https://cursor.com/blog/composer-2-5 https://cursor.com/blog/composer-2-5
- dang 5mo agoThanks! I had been just about to add that maybe the link wasn't the most informative. We've switched it now from https://twitter.com/cursor_ai/status/2056415413077233983 https://twitter.com/cursor_ai/status/2056415413077233983. As for the typo, s's are cheap and I've added one :)
- memoryleakgame 5mo agoIf these benches from their site hold up (they likely wont) Wouldn't this compress ai revenue like 15x quickly If they really have a 4.7 opus high equivalent at 1/16 the cost wouldn't this significantly effect all the current capex and planing Maybe they are getting elon to cover cost
- zackify 5mo agothis thing is so awesome on fast mode, so far i am impressed, some of its observations feel similar to opus. i use gpt 5.5 and opus 4.7 a lot every day, if i can get good results at this speed, hopefully the usage level holds up on my team plan haha
- infecto 5mo agoThe way I have read their benchmark results is that they trained a model to work insanely well in their coding workflow. It’s not a general purpose model. One of the surprisingly hardest problems to solve is to get a model to use the tools you give it access to.
- 2001zhaozhao 5mo ago> compress ai revenue like 15x that roughly just puts it on par with OpenAI and Anthropic subscriptions in terms of pricing per token
- smallnamespace 5mo agoAI revenue has been going up while the cost per token has been rapidly falling. The Jevons paradox applies here. The cheaper software is, the more software is written. There is not a finite demand for software.
- rafaelmn 5mo ago> AI revenue has been going up while the cost per token has been rapidly falling Every model release now has been straight price increases since what GPT 4 ? When was the last time a new flagship model decreased prices compared to the previous one ?
- polski-g 5mo agoI don't know why their model isn't on Openrouter yet. They must not have enough capacity to offer it.
- uf00lme 5mo agoI wonder why they didn’t train off Kimi 2.6, I hope is it because they already had a good base and not that they messed up that relationship.
- re-thc 5mo agoThat's 3.0
- NitpickLawyer 5mo ago> and not that they messed up that relationship. There's nothing to mess up. The license is MIT w/ attribution, and the attribution clause can be easily sidestepped w/o any legal repercussions. The "drama" was simply content creators going nuts over some misunderstandings and poor comms from some kimi related devs.
- big-chungus4 5mo agoCan you please train Qwen 3.5 like 0.8B to 9B using the same training techniques
- deleted 5mo ago[deleted]
- m_mueller 5mo agoIt's a bit confusing to me why they'd make this 'fast' version the default, as it appears to be much more expensive than Composer 2. Wasn't it supposed to be a very cheap alternative to SOTA models?
- mrklol 5mo agoIsn’t it a really cheap alternative to sota models (according to benchmarks)?
- goyozi 5mo agoI kind of want to try it, to see if and how far they can take an open model and improve it but I really don’t miss the Cursor user experience. Constant UI changes, half-baked features, smaller and smaller limits, useless AI change attribution; I think I’ll wait for others to report if it’s any good.
- epolanski 5mo agoGood point. One of the things I've came to appreciate about the cli tools like Codex or Claude is that the interface is so limited that every feature they release is still limited and constrained to the same UX limitations, whereas those "funkier" IDEs change from month to month giving me further fatigue.
- jstummbillig 5mo agoIsn't there a cli version of cursor by now?
- vorticalbox 5mo agoYes https://cursor.com/cli https://cursor.com/cli
- yourboirusty 5mo agoIt's a bit better than the VSCode fork, but still much worse than competition: - lags constantly, - if you type while it's generating you'll get missed inputs, - 'plan mode' doesn't clear context before starting work, - you can't directly edit the plan, you can only ask the bot to do it, - you can't immediately whitelist commands, only accept once or allow all.
- rubyn00bie 5mo agoDamn do I feel the UI changes being a pain point. It’s a near constant regression in my workflows. “Multiple agents” got destroyed recently, and the new interface for it some sort of command isn’t as good or reliable. Then you’ve got modals everywhere[1] and truncated bits (like long branch names) that make it insanely frustrating to use. They’re constantly changing the UI without actually improving it at all. I’ll likely cancel it and use opencode for personal stuff with Deepseek and only use it at work because I have to. There was a time when I appreciated the harness but it’s becoming less useful, or at least noticeable, over time… all the while the actual UI becomes substantially more painful and awkward to use (like @ in the “agents” window being completely unable to find a file because it’s some sort of “global” scope). One thing that surprises me about this whole segment is that JetBrains haven’t eaten these folks lunch. Their IDEs are leagues better than VSCode but their AI integration is awful by comparison (and the bar is low). I can’t even see how much of the context window I have left. [1] it’s insane I have to answer questions in a tiny input box I cannot resize or adjust the size of. Let alone the fact the text area I input prompts into cannot be resized. Truly feels like the UI/UX is done by people without any experience.
- contextcost 5mo ago[dead]
- SadErn 5mo ago[dead]
- throwaw12 5mo ago> Composer 2.5 is built on the same open-source checkpoint as Composer 2, Moonshot's Kimi K2.5. Really nice to see they're giving credit to the company and I am optimistic Kimi K open models soon will outperform Opus models
- howdareme9 5mo agoOnly because last time they tried to hide it lol
- trymas 5mo agoYes and if I remember the drama correctly - Kimi's license or terms of use says that for commercial use cases (or was it user count?) - you must declare credit to Moonshot and Kimi.
- Lennie 5mo agoIt's important to mention: they were compliant, because they trained the model at an AI hosting provider that had a partnership with Moonshot AI, but Moonshot didn't know Cursor was a customer.
- maxdo 5mo agoHow can distilled opus become better than original? There are numbers of reports including anthropic that kimi team was participating in fraudulent activities
- throwa356262 5mo agoDo we know the "fraudulent " requests really came from moonshot engineers and was not QA team running a ton of benchmarks against other models? I feel distilling something as big as Opus would require many many more samples, but I dont really know much about this subject
- maxdo 5mo ago
- bingud 5mo agoSeems like a promising and useful model but its probably scary how much customer data they fed into it to reach this performance
- try-working 5mo agoA lot of people saying Cursor have no moat. Sure. Neither do OpenAI or Anthropic.
- DeathArrow 5mo agoI think anybody will be much better by acquiring a coding plan from Kimi.com and using Kimi K2.6, with whatever harness they like, including Claude Code, instead of paying more for Cursor's version of Kimi K2.5.
- Dongyu_Jia 5mo agoWill this be the cursor's last dance? LoL
- I_am_tiberius 5mo agoI hope people soon wake up to the fact that they use user data for model fine tuning.
- zurfer 5mo agoKudos to the team. Please consider making the model available via API!
- bg24 5mo agoThey shipped an SDK recently. https://cursor.com/blog/typescript-sdk https://cursor.com/blog/typescript-sdk
- deleted 5mo ago[deleted]
- luodaint 5mo agoBenchmarks measure turn-level capabilities: you feed a task into the system and then grade the result. Capability for production-level usage concerns session-level decision making: does the agent know when to stop editing, retain the right amount of context, or go back and reread the file if the state has changed? This is not a property of the model, but a property of the discipline; it can be operationalized by what you have documented before the session begins. Without "stop editing where you can no longer follow your changes to the spec" and "go back and read the migration file before changing the schema," there is nothing to halt the process until it fails integration. Those teams who get consistent results independent of the model being used typically do so because they have operationalized their discipline first. Those switching out models monthly tend to expect the model to supply them.
- 0fes911 5mo agoI found composer 2 pretty good as a subagent delegating tasks like auditing for bugs after finishing implementation, but hopefully composer 2.5 will be more reliable so it can be used to implement and execute long running tasks.
- brunooliv 5mo agoAny reason why they indexed on Kimi K2.5 model? I have tried many open-source ones in Opencode, and, in my experience (standard backend development, Java, Python, Spring, etc) Qwen3.6 is SO MUCH BETTER that's shocking. Kimi can't even get most tool calling arguments right.
- Bombthecat 5mo agoCheaper to run?
- CuriouslyC 5mo agoThere's a lead time on models, and there's some tuning gotchas they probably already figured out with Kimi, so they weren't ready to just drop everything and switch. I'm sure they will switch models eventually.
- roflcopter69 5mo agoI recommend reading the entire article Together with SpaceXAI, we're training a significantly larger model from scratch, using 10x more total compute. With Colossus 2's million H100-equivalents and our combined data and training techniques, we expect this to be a major leap in model capability.
- grim_io 5mo agoI guess this will largely decide if xai is going to pay 60 or 10 billion, depending on the success of the new coding model.
- KaoruAoiShiho 5mo agoKimi 2.5 has the best long context. For raw coding benchmark scores you can just post train on top of it with more specialized data. 2.5 is kinda old, 2.6 is the current release which is exactly just that and catches up to the frontier in most aspects.
- Glohrischi 5mo agoHahah wtf? They are training on colossus 2? Their own model? Dude what the hell happened to Musks Grok? How incapable are they that they give away training compute to Cursor like this? Weird that the genius Musk doesn't need his own compute, after all shouldn't Macrohard (no joke) already building the worlds software from scratch?
- mgambati 5mo agoWords on the street is that xAI will buy cursor.
- Glohrischi 5mo agoYeah for 10-60 BILLION. which again makes this even stupider. For this amount of money you can rebuild cursor and everything else on the market, and with the rest of 9-59 Billion, you just hire experts in coding and let them code real high quality code examples. And then you just use your existing grok pipeline and just add this functionality. This xAI stuff has to be run by idiots
- radu_floricica 5mo agoBuy "Cursor", not "Cursor's IP". This means brand, users, and a shitton of data. And if you combine a shitton of data with a lot of compute, large userbase and good engineers, you have a pretty good chance of doing something interesting.
- Glohrischi 5mo agoYeah you know how much 10-60 Billion are? You could literaly just give your compute away for free for a year to pull people in. Make an API Endpoint for free with the caviat that they are allowed to use the data for traing, what everyone else does too.
- mgambati 5mo ago
- jorl17 5mo agoI want to like composer, but I just can't. - Its communication style is completely opposite to Anthropic models. It's not as bad as OpenAI's models, which are obsessed with "shapes", "wrinkles", hyphenated-words, and other cryptic formulations that make you feel like you're not on planet earth after a while talking to them. But it is nonetheless markedly "rude", "dry", "cold", gives off this "entitled I'm right, you're wrong" attitude. I once had composer2-fast accidentally run `rm -rf $HOME` (no harm done) as part of a bug in an install script it wrote and all it could say once it realized it was: "Running script with proper hardening". Qwen's models have clearly been distilled from Anthropic models because they have a much closer communication style and that's why I hope cursor will one day release a new family of composer models derived from that. A damn joy to use. - It's just dumb. I don't know what they're doing with benchmarks, but for my work (python, bash, docker, whatever), cursor is just incredibly dumb. Always does in 10 lines what could be done in one. Doesn't know loads of internals of things that other models know. Never places things in the right files, constantly makes terrible edits (inline imports, edits without testing). Everything is so complicated when done by composer2, it's just a joke to me at this point. It clearly needs more handholding than Opus 4.x or GPT-5.x. I tried 2.5-fast and it seemed more of the same. And this would sort of be acceptable if it owned up to its incompetence, but it is so confidently incompetent that it's revolting. I know that for many people the "tone" of the models is not relevant, or maybe they even prefer models like these. I simply cannot work like that. Ever since Gemini started blowing benchmarks out of the water while being a clearly inferior model incapable of producing anything (and pretty much just doing tool calls without any feedback to the user), I gave up on benchmarks. Composer has been more of the same in that regard. As a GPT model would say: "Small wrinkle: the production-ready benchmark results were tainted by real-world data points. I've assimilated the inconsistencies and added guardrails so that v2 has the right shape for future evaluations."
- vinzdg 5mo ago[flagged]
- machiaweliczny 5mo agoTested and it's good. Fast version is bad though. I like planning model in Cursor that it works more like human written design doc instead of too detailed AI plan. Seems like this is more responsible for results that model but still on fast it failed but on normal got good results.
- WhitneyLand 5mo agoSay what you want about Cursor but they don’t lack for ambition. Forking VS Code, going big on bleeding edge features like cloud agents, and now they’ve thrown down the gauntlet directly challenging frontier labs by training their own model (“much larger” than Kimi 2.5’s 1T parameters) from scratch. They’ve been highly successful so far. Raised $50B, $2B in revenue, forecast to end 2026 above $6B. But even at these heights, they’re just not in the same league as OpenAI/Anthropic/Google. And if building a state of the art multitrillion parameter model is not challenging enough, it’s a mountain you don’t climb just once. Every few months you need to push it farther with a new release. Fall off for a couple cycles and like Facebook you may never catch up again. Not for the faint of heart.
- causal 5mo agoYeah I want them to do well. I find Cursor to be a much better tool for actually working with the code the agent writes than whatever the big vendors provide.
- worldsavior 5mo agoThem raising this much money doesn't mean they're successful, it only means they know how to fool the investors well. A project that is basically an extension to VSCode only adding a chat interface, isn't really worth this much money. Obviously, it's the users, but people think it's something genius and revolutionary, but no.
- infecto 5mo agoThis is rsync all over again. Go create it yourself if you think it’s just a simple extension.
- worldsavior 5mo agoYou're right, I regret I didn't have the sense to do the same as them at the time.
- 5mo ago
- enraged_camel 5mo agoI tested it yesterday. It is pretty bad. Just like with Composer 2, it's fast, but quality is nowhere near what Cursor claims with their benchmarks. It is not even at Opus 4.5 level. I gave it a mix of refactoring tasks and new feature tasks. For each one, I had it write a plan, then I had Codex review it. Codex found major issues with every plan: patterns that don't match the rest of the code base, hallucinated variable/function names, and even outright bugs in the way the plan was written. I fed the feedback to Composer 2. After it made the changes and implemented the revised plan, I had Codex and Opus 4.7 do code reviews, and once again both of them found major bugs. Overall it was a very frustrating experience. I feel like I wasted a whole day. Which is sad, as I have been looking for an excuse to come back to Cursor. But as things stand, Codex + CC combo cannot be beat, not just in terms of price but also quality.
- steviedotboston 5mo agoIt's very confusing that they use the same name as the very well known PHP package manager, composer https://getcomposer.org/ https://getcomposer.org/
- wesammikhail 5mo agoI dont know what it is with products names these days. Antigravity, Antimatter, Composer, Clay, Ramp, Bolt, etc. You'd think the founders would Google for naming conflict before choosing a name.
- NikolaosC 5mo ago[dead]
- joka88xj 5mo ago[flagged]
- rcleveng 5mo agoI have to say the new model is quite good at the basics, I've been handing over more and more tasks from Linear straight to it instead of the copy-paste into Claude dance lately. At this point, more of my complaints are on the harness side, which is odd since originally they were by far the best harness out there. Support - This is pretty much non-existant, it's community support or sales support. Interacting with GitHub - this should work and be awesome, Claude code does this well (responding to lint errors and comments). Cursor you have to poke the agent to look at the comments or lint errors, and even then it's about 10% good. Even GitHub Copilot is better here. Bugbot - I have it setup to trigger manually, but it still seems to wake up and burn 80-120k tokens just to notice it's configured to be manually invoked. When it does run, it tells me there's no issues (but claude or copilot both find real things) App - When you have both agent window and the ide windows, it's hard to open up the code in the right directory. A simple "cursor ." from the terminal used to do it, now it'll often open the agent window, you have to try a few times for it to work. I love that they are running super fast, it's just hard when many of the basics break or don't work.
- khazhoux 5mo ago> I've been handing over more and more tasks from Linear straight to it instead of the copy-paste into Claude dance lately Tangent: we've been using Linear at work and I still don't understand why it claims to be "task tracking for agents". Is there anything at all that lends itself better to agentic workflows compared to JIRA or gitlab/github issues or whatever else? Seems like Linear just hopped on the buzzword hype train at the exact right moment...
- dbalatero 5mo ago> Seems like Linear just hopped on the buzzword hype train at the exact right moment... I think you nailed it. Provided an agent can connect and ingest the information in the ticket, that's basically what's needed. I guess it's nice to be able to nudge ticket status and post back to it, but all of those seem like wiring up existing APIs to an MCP and calling it good. I don't see why JIRA couldn't execute on that, despite being Atlassian.
- Armonsrer 5mo agoIt looks a massive update from cursor and i like their platform Let hope its good
- k3ymaker 5mo ago[dead]
- wunderlotus 5mo agoI love Cursor as a tool, but I'm skeptical bc: 1/ CursorBench is so opaque [1] that it makes it hard to trust. Not to mention the v3.1 eval is a newer iteration and there's no insight into the tasks or if the model was just tuned to max it out. Composer 2 previously scored between 60-65% on the previous benchmark eval [2] but scores between 50-55% on CB v3.1[3]. 2/ I've experienced Composer 2's performance and it leaves much to be desired as a daily driver for a knowledge worker. but KWs are obviously not the target users and I can see how it's cost-efficient for executing on clearly-defined, discrete coding tasks. Obviously that's their value proposition and they're figuring out how to communicate it well to the target customer. It just doesn't feel like CursorBench is that. [1] https://cursor.com/blog/cursorbench#building-cursorbench https://cursor.com/blog/cursorbench#building-cursorbench [2] https://cursor.com/blog/composer-2-technical-report#performance https://cursor.com/blog/composer-2-technical-report#performa... [3] https://cursor.com/blog/composer-2-5 https://cursor.com/blog/composer-2-5
- chemex 5mo agoI've been using Claude Code as my daily driver on a React Native + iOS codebase for the last few months. The thing that surprised me wasn't quality differences on individual edits — those are pretty close once you control for harness wiring — but how differently I'd ended up structuring my workflow around each style of tool. Tab completion + chat-in-sidebar feels like an extension of my editing. An agentic harness feels more like delegating a 20-minute task and coming back to review. Different cognitive load, different bug profile. The "which is better" framing tends to skip over the fact that they reward different working styles. Two things I'd watch on Composer 2.5 specifically: 1. How it handles long-running multi-file refactors that touch 10+ files. My experience with smaller models in that slot is they lose track of which files they've already edited around 30% of the way through. Frontier models keep the plan coherent for longer. 2. How it deals with non-obvious file boundaries. The thing that takes me out of "let it work" mode is the model deciding it needs to edit a config file I didn't think of. Usually that's right, but occasionally it's spelunking somewhere I don't want it to be. The Kimi K2.5 base is interesting on its own. Open weights below frontier closed models is the thing worth watching from the harness side. If anyone's set up to fine-tune for a specific harness, this is the moment.
- chis 5mo agoAI slop detected, you're under arrest
- ryanshrott 5mo agoThe cost claim is the easy part to sell. The real test is whether it stays useful in ugly codebases, long files, and repos with a bunch of half-broken conventions. That’s where these assistants usually fall apart, even when the benchmark numbers look great.
- sofumel 5mo agoI'm currently using Claude Code, but should I cancel it at the next renewal and switch to Composer 2.5?