5 ms·
It seems to be extremely economical - 4x better reasoning efficiency compared to Opus while being priced at $2/$6. For comparison, GPT 5.4 is $2.5/$15, GPT 5.5/
by Tiberium 3mo ago
It seems to be extremely economical - 4x better reasoning efficiency compared to Opus while being priced at $2/$6. For comparison, GPT 5.4 is $2.5/$15, GPT 5.5/5.6 are $5/$30, Opus 4.8 is $5/$25, Fable is $10/$50.
And by benchmarks (unless they gamed them), seems to be at around Opus 4.7 level, which is what Elon mentioned in https://x.com/elonmusk/status/2074911038286295049 https://x.com/elonmusk/status/2074911038286295049.
I guess the Cursor data was very useful.
- minimaxir 3mo agoThe comparison may be better against GPT 5.6 Terra (instead of Sol), which is $2.5/$15.
- Tiberium 3mo agoWe don't yet know Terra's results for DeepSWE/TerminalBench though.
- giancarlostoro 3mo agoNow if they could have an "equivalent" to Claude's $100 plan with similar compute limits. I have the $40 a month version of Grok and I get a max of like 8 hours of "non-stop" Grok Build coding, per month.
- BoumTAC 3mo agoGrok Build sucks compare to composer 2.5. Just use compose 2.5 and you'll have basically unlimited usage on the 40$ plan.
- DoesntMatter22 3mo agoComposer 2.5 is so underrated IMO. I built a really feature rich application, insanely complicated, close to 200k LOC since it came out and for the most part it ran like a champ. Only used CLaude a couple times to get it unstuck. 8 hours a day and I'm paying about 30 a month.
- embedding-shape 3mo ago> I built a really feature rich application, insanely complicated, close to 200k LOC If you listed it, how many features/LOC or vice-versa? Really hard to know if 200K LOC is good or bad, at the surface it sounds like too much, but I don't know what the application was either.
- DoesntMatter22 3mo agoIt’s a fantastic signal processing / engineering app. There are 5 major players and this app isn’t quite as good but it’s in the ballpark. I’d day when I release early next nonth this will be the biggest fully featured vibe coded app I’m aware of.
- alchemist1e9 3mo agoI’d be curious to hear more about your dev setup and what tips you have for other aspiring vibe app coders.
- DoesntMatter22 3mo agoI currently just use a Mac Mini M4 with cursor. I had 30 years programming before this (in my mid 40s now) so I am not a great person to give vibe coding advice because I have so much actual coding m. I’m pretty sure I couldn’t have done it without knowing how to code because there were some times even the best models would get stuck and I’d have to figure it out myself. If I had to give advice I’d say to just do it. Work on some project that interests you and go for it
- esafak 3mo agoHow do you like its design mode?
- giancarlostoro 3mo agoSuppose eventually that gravy train will disappear, might as well use it then.
- bhouston 3mo agoEvery time I use Composer 2.5 I have to spend a bunch of time cleaning up its mistakes. It is unusable compared to GPT 5.4 or 5.5. My time is more valuable that I will use a model that doesn’t f** up my code base.
- subhobroto 3mo agoI think we need to be explicit about the domains we're applying Composer 2.5 to in these discussions. I mentioned here (https://news.ycombinator.com/item?id=48766275 https://news.ycombinator.com/item?id=48766275) how poorly it handles my specific use cases. My coworkers in DevOps and frontend UI swear by its cost-effectiveness, whereas I strongly prefer the reasoning capabilities of Opus 4.8 and Fable 5. Composer 2.5 seems to be SOTA for Helm charts and React/Vue, but, for my usecases it absolutely struggles spectacularly when tasked with rigid body dynamics or kinematic logic.
- sroussey 3mo agoKeep composer away from anything configuration related—it will ruin your day.
- trollbridge 3mo agoIsn't Composer 2.5 designed to be used from the Cursor harness, and is otherwise not that useful?
- bhouston 3mo agoI use it from the Cursor harness.
- AgentMasterRace 3mo agogive it a structured plan and it it does really well compared to similar priced models. I'd never use it for anything that required heavy reasoning and it's not built for that.
- DoesntMatter22 3mo agoWeird I use it 8 or so hours a day and I haven’t had that issue
- tengbretson 3mo agoIt is hard to evaluate the model performance of Composer 2.5 when Cursor's harness is so awful compared to the others on the market.
- brightball 3mo agoIn what way? I spend more of my time managing than hands on lately so I legitimately don’t know.
- seunosewa 3mo agoNot true. The only issue is cost of frontier models.
- sergiotapia 3mo agoNot my experience at all. I've been using Cursor hardcore for about two weeks now and Composer 2.5 and it's wonderful. Now with Grok 4.5 I'm quite excited about the possibilities.
- sroussey 3mo agoYeah, it’s not great—except for debugging. It shines there.
- bdlowery 3mo agoCursors whole moat is their harness. Other people have benched opus and GPT models through Claude code, codex, and cursor, and cursor came out on top everytime because of their harness.
- AgentMasterRace 3mo agoThe harness is commonly ranked one of the best. what specifically had you hating it?
- steve1977 3mo agoCare to elaborate?
- Tiberium 3mo agoThe model is available through Cursor which has $20, $60 and $200 plans. I assume the $60 version might work better for you?
- giancarlostoro 3mo agoWill have to give that a try I suppose.
- kesor 3mo agoYou get "Grok Build" (the CLI) that uses the Cursor and/or Grok 4.5 models when you buy SuperGrok, which is like $300/year. I don't know if there is a feature-to-feature comparison by anyone on this, but you can get access to these models with unmetered tokens with SuperGrok.
- srid 3mo ago> you can get access to these models with unmetered tokens with SuperGrok [$USD 300/year] Could you support this statement with an official reference?
- 2001zhaozhao 3mo agoAround Opus 4.7 level would be the same as Sonnet 5 while being cheaper overall. I wonder how good their subscription discount is on both their subscription types.
- Tiberium 3mo agoSonnet 5 is a huge token hog, though, it uses far more reasoning tokens than Opus models while being priced at $2/$10 with promo, and $3/$15 (usual Sonnet price) afterwards.
- giancarlostoro 3mo agoI'll probably get hate for it, but I was not impressed by Fable, I felt like it was just Opus with more tokens for thinking. I feel like the second I turned on Fable I drained my usage more quickly, despite them billing it as though it were Opus level of usage. The value is just not there for me. I wish they could make Haiku remain low-cost and drastically more capable to the point you could use only Haiku.
- kittoes 3mo agoDid you explicitly tell it to use Sonnet or Opus subagents and stick at or below high effort? Asking because such practices make a huge difference in the quality of output and the amount of tokens burned. I used one of my accounts to explore ultramax and it was just a token hog that might be worse than Opus.
- giancarlostoro 3mo agoI had it on whatever the recommended settings was, but maybe I should have told it to use Sonnet for most subtasks. Even so, I'm just not that impressed, I felt like I got more done by just using Opus.
- sroussey 3mo ago
- conradkay 3mo agoAnnoying they didn't show benchmarks for several effort modes, since it seems like it might close the gap with Opus 4.8 by cranking tokens up? Noam Brown (OpenAI) "Implications of Large-Scale Test-Time Compute" https://xcancel.com/i/article/2064210146558136827 https://xcancel.com/i/article/2064210146558136827
- HarHarVeryFunny 3mo agoThe $2/6 pricing seems to only apply for context under 200K. Above that (max context is 500K) pricing doubles to $4/12. https://docs.x.ai/developers/models/grok-4.5 https://docs.x.ai/developers/models/grok-4.5
- jadbox 3mo agoThat's very notable and left out of the announcement.
- gabriel-uribe 3mo agoWomp. Didn't see this anywhere else. No longer feels as inexpensive. Will likely just include this in the rolodex of <200k context tasks, like being one of my review agents.
- seer 3mo agoYeah but depends how you use it - with superpowers and it’s prevalence of splitting things into smaller focused subagents - this could seriously reduce costs … I wish my company gave me more options than just using Claude to test these things out
- braebo 3mo agoClaude Fable and Opus 4.8 1M are by far the best, smartest models. Anything else is a downgrade so you’re not missing anything.
- HarHarVeryFunny 3mo agoThe recent Databricks comparison has GLM 5.2 performing identically to Opus 4.8 on high effort, and some early Twitter reports (e.g. from the OpenCode developers) strongly favor GPT 5.6 Sol over Fable. As always it depends on what you are using them for, and how you are using them.
- GodelNumbering 3mo ago
- deleted 3mo ago[deleted]
- game_the0ry 3mo agoI have a theory that xAI has one of the largest clusters but with far less traffic + tokens to process bc its less popular than its competition, and xAI can pass the savings on to the end user.
- goos 3mo agoWhy would having more costs and less income allow them to pass savings on to the end user?
- rjh29 3mo agoThey already invested in the massive datacentres of GPUs sitting idle. They have fewer users so they can deliver more inference per user - more thinking, larger models.
- mvdtnz 3mo agoSo where are these mythical savings coming from? You're saying they have spent more per user therefore can charge each user less or something? I'm not following.
- spacebanana7 3mo agoThe (optimistic?) take is that xAI is genuinely better at building datacenters at scale than anyone else, and the freedom to use Nat Gas as the primary energy source allows them to have lower marginal costs. The (pessimistic?) take is that they have loads of idle GPUs and want to get some revenue out of them rather than none. Compare this to OpenAI/Anthropic where every token used by a consumer has to compete with enterprise spenders, and there’s not enough to go around for everyone.
- rjh29 3mo agoIt's also sensible for them to provide a cheap, intelligent model to users if they have capacity, then once they built a user base, tighten the screws. All the other AI providers have done that.
- numpad0 3mo agoHow does it compare to Chinese APIs? It doesn't seem like xAI is meaningfully more competent or any single bit more honest than Chinese labs anyway, so you might as well send tasks straight to China unless theirs is substantially cheaper.
- numpad0 3mo ago[dead]
- bashtoni 3mo agoBut very expensive compared to Deepseek v4 Pro, which performs similarly. Grok is stuck in a difficult place - not the best model at anything, and not the cheapest either. It's hard to make a case for using it on any dimension, even before you factor in the history (I'm not sure suggesting the company uses the model that refers to itself as "MechaHitler" is the way to a promotion).
- cousinbryce 3mo agoIt’s the #1 model for creating CSAM
- tick_tock_tick 3mo agoAn AI model can't create CSAM unless you're claiming it's managing to hire people to commit crimes in the real world?
- insane_dreamer 3mo agountil I exceed my Claude Max/$100 sub (have not hit a wall so far), the pricing of these models isn't relevant to me
- gertlabs 3mo agoGrok 4.5 is a huge step up from their next best model and now around the same performance as GLM 5.2, but it's not exactly at the frontier of the cost efficiency curve in our coding evaluations. That curve is defined by the 2 lighter GPT 5.6 models. However, the fact that they finally have a strong post-training and RL setup bodes well for future releases. They certainly are not compute-constrained anymore. Data at https://gertlabs.com/rankings?mode=oneshot_coding https://gertlabs.com/rankings?mode=oneshot_coding
- Chu4eeno 3mo agohuh, why is "Positive-Sum" such an outlier when comparing grok 4.5 and GPT-5.6 Sol?