11 ms·
Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index
- satvikpendem 2mo agoCursor, since Grok 4.5, has had an incredible deal for frontier level models, their subscription now goes way further than OpenAI or Anthropic. Even on their lower tier plans you can use a lot tokens on their of their first party models (Grok and Composer) and not really run out comparatively. Combine them with an orchestrator and implementor type setup and it goes even further.
- aliljet 2mo agoCan you explain what you mean? These days courtesy of an addictive reset game OpenAI is playing, I can't find anything with frontier intelligence that's more cost efficient...
- timr 2mo agoIf they didn’t constantly reset, they’d be about the same as Anthropic. Right now, I find that Grok offers better value, uses fewer tokens per turn, and makes better code. I haven’t tried Cursor because I don’t want to change editors again, but maybe I should try it…
- apitman 2mo agoAre there any projects that track how much usage of each model translates to how much percentage drop in weekly/5hr windows?
- timr 2mo agoNot that I know of. AA's token use metrics (mentioned in this article) are indicative, however. They say explicitly here that the Grok models are notably token efficient. This is my experience.
- taosx 2mo agoUsage? Not exactly. But I tried to make something that can estimate dollars per tokens in actual usage while taking into account multiple factors. https://harness.eveid.com/lazy-harness-cost-simulation https://harness.eveid.com/lazy-harness-cost-simulation
- hyldmo 2mo agoJust curious, was this coded with Claude or Codex? Copy reads very Claude to me but I’m curious if thats an actual pattern or just me
- deleted 2mo ago[deleted]
- esafak 2mo agoThat is not true; GPT is the most reasoning efficient model family on the market.
- pickleRick243 2mo agoYeah, even without the resets, chatgpt subscription currently goes quite a bit further than an equivalent anthropic plan. The main reason to have an anthropic plan is to get access to Fable 5 if you feel the quality of output makes it worth it.
- timr 2mo agoThe benchmark article we're replying to shows that Grok token usage is at least on par with the latest OpenAI models [1], and significantly cheaper per token: https://artificialanalysis.ai/models/grok-4-6#token-use https://artificialanalysis.ai/models/grok-4-6#token-use So depending on how you want to define "token efficiency", Grok is either tied with OpenAI, or in the lead. [1] Though I grant that 4.6 appears to be wordier, on the order of Terra max.
- jesse_dot_id 2mo agoGoes even further to exfiltrate your data, yeah.
- greenavocado 2mo agoThat would be Muse Spark Contributor Tier. 12-21x price reduction at the expense of your digital existence.
- CuriouslyC 2mo agoI'd be the first model I'd reach for if I was providing a free service to AI gooners though. Serves them both right.
- greenavocado 2mo agoThat's exactly what's going on LOL
- hmokiguess 2mo agoI believe they are the only western provider that has Kimi K3 on a subscription plan today as well. I would love to ditch Anthropic and be on Kimi if there were a subsidized plan like that with ZDR
- jeffyaw 2mo agoyou can use Kimi K3 on the typed++ model tier: https://typed.cloud https://typed.cloud
- homakov 2mo agokimi k3 credits end in just a few sessions. Only Grok models allow generous use in Cursor Pro/+
- pkaye 2mo agoGitHub Copilot does have Kimi K3.
- timr 2mo agoWhat’s the multiplier? GH copilot nerfed their product so badly that I unsubscribed.
- pkaye 2mo agoThey don't do request based pricing anymore. Its just token based (1 credit = $0.01) plus some bonus credit based on which plan you subscribe. So for example a $39 plan get $70 of credits. https://github.com/features/copilot/plans https://github.com/features/copilot/plans https://github.blog/changelog/2026-08-06-kimi-k3-is-now-available-in-github-copilot/ https://github.blog/changelog/2026-08-06-kimi-k3-is-now-avai...
- timr 2mo agoYeah, I know, but "credit" translates differently because the models bill at different rates, which gets turned into "multipliers" (or at least, it did). Have they converted entirely to transparent API rates + base allocation now? One of the reasons I left was that if I was going to be billed at API rates anyway, I'd just rather use the APIs. The value proposition still sucks for individuals now, when the other major providers are bundling at below-API rates.
- nomilk 2mo agoHow does Grok 4.5 compare to Opus >= 4.8 though? I'm willing to pay 2x for a 10% smarter model. Intelligence matters that much (because 10% smarter probably saves, on average, several hours of human time).
- redox99 2mo agoIt's a bit worse. I haven't tried so it's pure speculation based on benchmarks, but I'd assume Grok 4.6 is around Opus 4.8 in real world use, but clearly below Opus 5.
- douglee650 2mo agoI've found Fable 5 to be so much better than 4.8. For building a full stack custom CRM and media pipeline tool with video conversion, transcription, and indexing. Supabase, AWS, Meili, NextJS, GCS - lots of surfaces and planes. 4.8 basically couldn't do it, I abandoned the project as the fallback was, "current business processes". With F5 it's been 4 weeks and almost ready for production release.
- mandeepj 2mo agoI have the same quality results with Fable. With just a brief prompt, it created a great static website with a beautiful animation of a workflow. Gemini's output was so poor that I closed the chat. And with Codex, the results were bad, so I discarded them.
- douglee650 2mo agoAgree. Opus 4.8 could make nice little toy and demo apps. This app had a lot of surfaces and pipeline, Postgres, vector search, S3 -- couldn't handle that. Fable 5 is still a lot of work and you have to check it, but it really does perform at senior eng level. Shipping good size features daily.
- deleted 2mo ago[deleted]
- Computer0 2mo agoWhen I last used Cursor their subscription covered usage of ~$20 per month. Have they switched to a subsidized subscription model like ChatGPT and Claude?
- satvikpendem 2mo agoSubsidized for their own models now, plus 20 dollars of API credit for non first party models.
- johnnyApplePRNG 2mo ago>their subscription now goes way further than OpenAI or Anthropic. Until it doesn't... Honestly, this entire OpenAI reset credit fiasco this past week has convinced me to rip off the Codex and Claude Code bandaids and start building my own proper Pi Coding Agent running models that I select and pay for on openrouter. And I am feeling a lot better about it now that I've finally got it working.
- unglaublich 2mo agoBut still for US frontier you're paying 10-20x more per token compared to their limited subscriptions. For China frontier you'll be good though, and that might be the future anyway.
- maxdo 2mo agoGrok is cheaper vs real Chinese frontier aka kimi. Sponsored or not.
- indigodaddy 2mo agoWhat's the cheapest way to use grok models for coding?
- satvikpendem 2mo agoCursor subscription as I mentioned
- johnnyApplePRNG 2mo agoRelying on a single frontier model to just zero-shot all the work is so 2025. Deepseek V4 Flash 0731 is surprisingly capable and cheap. [0] Checkout pi coding agent. You can create as many different sub-agents as you wish, to specialize and understand and tackle or pass off any problem you like. It's refreshing, really. I feel like a coder in control again. [0] https://arcprize.org/results/deepseek-v4-flash-0731 https://arcprize.org/results/deepseek-v4-flash-0731
- 2mo ago
- everfrustrated 2mo agoCursor also allows disabling Grok Fast mode which means tokens last forever. Fast is great tho, but nice to have the option.
- indigodaddy 2mo agoThe value in their subscription is bound to the Cursor agent/software only though correct?
- btreecat 2mo agoSounds like you have your final solution
- thiago_fm 2mo agoI often wonder if there's a chance, even if minimal... that they stole the weights of the Anthropic models they run on their datacenter... or are actively destillating it.
- connicpu 2mo agoI think the more likely explanation is that the Cursor data they effectively acquired for $10B was extremely valuable for their training when combined with the insane number of GB300s xAI has for training.
- winstonp 2mo agoCursor was 60B. The 10B number was the breakup fee if the deal fell through.
- bpodgursky 2mo ago$60B in SPCX stock, which is sort of magic money.
- qudat 2mo ago> ... or are actively destillating it. I just assumed every model manufacturer is distilling from the frontier models. If they aren't they are definitely trying to do it.
- scottyah 2mo agoI wonder if the distillation was part of the compute deal.
- nylonstrung 2mo agoI have never met a single human being who uses Grok for coding
- sidcool 2mo agoHello. Nice to meet you.
- supriyo-biswas 2mo agoI'm only being forced to use it at $WORK since some people overran their Cursor bill, so everyone gets Cursor Auto enabled by default which routes to Grok 4.5.
- kvirani 2mo agoFolks working in US govt tend to, based on convos I've had with one such person.
- deleted 2mo ago[deleted]
- dogmayor 2mo agoNot the best endorsement given the current US gov
- nozzlegear 2mo agoYeah, sounds a bit like a selection bias, i.e. the people who managed to survive the firings by DOGE and the current admin are the kind of people who might prefer Grok.
- dfedbeef 2mo agoThey make good stuff
- Gigachad 2mo agoGrok probably doesn’t object when the government asks it how to bomb schools.
- pzo 2mo agoSeems the cache read pricing almost doubled from $0.30 in Grok 4.5 to $0.50 in Grok 4.6. In my experience in heavy coding sessions most pricing is just cache read and cache write like 80% of my token bill.
- sidcool 2mo agoGrok is not the best model around, but it's decent. It gets the basic job done at a low price. I don't think it can advance frontier Math, yet.
- MrBuddyCasino 2mo agoIt is also not annoying to use. It doesn’t overcomplicate things, and its quick. Much better than eg GLM 5.2. Pretty good bang for the buck.
- leerob 2mo agoProbably can't advance frontier math yet, yeah. But please let us know other places you want to see Grok improve for future models!
- solid_fuel 2mo ago[flagged]
- freejazz 2mo ago[flagged]
- trollbridge 2mo agoGrok doesn’t have those features, and people who like to make adult content have been complaining for a while how Grok has made it a lot harder to do so.
- freejazz 2mo ago[flagged]
- trollbridge 2mo agoYou’re asking for deepfake porn creation and for children? What the heck is wrong with you???
- petesergeant 2mo agoInteresting. Grok 4.5 is a capable model, although not quite at Fable/Sol levels. Will be interesting to see how this holds up. Musk appears to have made a savvy choice buying Cursor's data.
- TSiege 2mo ago[flagged]
- dmix 2mo agoMechaHitler was something that existed only on X's grok chatbot, due to a one-line system prompt change they reverted after half a day. That's different than using Grok as a model for coding.
- freejazz 2mo ago[flagged]
- losvedir 2mo agoI think "system prompt" is the key bit they're getting at. It doesn't necessarily reflect poorly on the underlying model if the system prompt was bad. It does reflect somewhat, in terms of alignment (how well the model does what the training company wants) and instruction following (how well the model does what the user wants). But it's not so clear to me what exactly the right answer is here. E.g., a model that scrupulously follows its system prompt and does what the user wants is a pretty useful, if very sharp, tool, albeit perhaps dangerous in the wrong hands.
- freejazz 2mo agoThe point isn't that it's the underlying model, it's that it happened at all in the first place...
- GlickWick 2mo agoIt's true, but I do worry about governance when it comes to these models. That shows a surprising lack of discipline in their deployment pipeline.
- dmix 2mo agoAgreed but there were similar controversies with how OpenAI was generating images. The only pass is these are the early days of chatbots and this stuff is so non-deterministic and experimental. For context, this was the change Grok's team made, that was later reverted: > - The response should not shy away from making claims which are politically incorrect, as long as they are well substantiated. https://github.com/xai-org/grok-prompts/commit/c5de4a14feb50b0e5b3e8554f9c8aae8c97b56b4 https://github.com/xai-org/grok-prompts/commit/c5de4a14feb50...
- alpha_squared 2mo ago[flagged]
- macinjosh 2mo agoThis comment is literally off topic. You opened the comment saying as much. So I have dutifully downvoted it.
- appplication 2mo agoNot disagreeing with you at all, but welcome to our new AI enhanced world. Nothing you enjoyed regarding human interaction, trust, or social norms is safe. HN will not survive 5 years, and likely less. There is too much money to be made by capturing discourse on the major (and minor) forums of the internet. The more trusted that community is, the more valuable it is to pillage with AI astroturfing.
- scottyah 2mo agoI love how this comment is both off-topic and heavily biased.
- small_model 2mo agoSpaceXAI is the only frontier model company that had its own compute/date centres and soon chip making factory, I think they will pull ahead with cheaper tokens similar intelligence and better harness/tools. Grok build is 2-5x faster than Claude Code in my opinion.
- bdangubic 2mo ago[flagged]
- hackyhacky 2mo ago[flagged]
- UltraSane 2mo agoExcept EUV lithography is the most complicated industrial process that exists and they won't have usable yields for many years if ever. I don't think Musk actually expects these fans to ever actually make sense they just let him hype and distract.
- krakrum 2mo agoLight generation is the most complicated part and Elon plans on doing Free Electron Laser (FEL) which is not as complicated as self contained tin based solution that ASML uses now
- UltraSane 2mo agoWhy do you give Elon any credibility? FEL is not proven at all and is much much riskier than using existing EUV machines. Plus if it breaks all your machines are down until you fix it.
- jamesrcole 2mo ago[flagged]
- insane_dreamer 2mo agoWhy is Grok so much cheaper than Claude or GPT?
- petesergeant 2mo agoElon owns a lot of compute
- iSloth 2mo agoHe owns the DCs and isn’t scrambling for revenue to justify an upcoming IPO
- mrguyorama 2mo agoDemand so much lower they had to resell capacity.
- stickfigure 2mo agoThe answer is simple: They're willing to burn money faster than the others. Nobody's profitable in this space, they can price it however they want as long as investors keep pouring money in. And SpaceX just got a lot of money poured in.
- mohamedkoubaa 2mo agoMy strat until the bubble pops is to just use the most subsidized model with acceptable performance
- slopinthebag 2mo agoBecause the model is probably smaller. Go look at openrouter’s costs for other open models around this performance level, they’re similar.
- deleted 2mo ago[deleted]
- HardCodedBias 2mo ago[flagged]
- mchusma 2mo agoNice to see SpaceX on the model frontier! They have been chasing it for a while.
- paimapi 2mo agocool here's Stanford HAI's graph on the carbon emitted from model training per model: https://spectrum.ieee.org/media-library/chart-showing-estimated-carbon-emissions-from-training-of-ai-models-from-2012-to-2025-with-grok-3-and-grok-4-the-chart-shows-a.jpg?id=65506058&width=980 https://spectrum.ieee.org/media-library/chart-showing-estima... note that Grok's training, thanks to its portable gas generators that are magnitudes less efficient than even other integrated, permanent gas turbines, means the training for this model is dramatically less efficient than models like DeepSeek a lot of the CO2 emission debate on AI is overblown but it's accurate for Grok
- deleted 2mo ago[deleted]
- rd 2mo agoI'm just wondering why they sold compute to Anthropic if they were planning on still competing in this race?
- JLO64 2mo agoLikely because they had the capacity to spare. Prior to Grok 4.5, I doubt there was much demand for their models.
- scottyah 2mo agoFor distillation deals lol
- Strom 2mo agoThe revenue was critical to making their IPO numbers look a bit less insane.
- HarHarVeryFunny 2mo agoCompeting doesn't mean winning The rental deal can be terminated by either side with 90 days notice, and presumably Musk would do so if he needed the compute or generally thought it advantageous to do so. For now he doesn't need the compute. The rental deal may also have been at least in part to juice the SpaceX IPO and to help Anthropic stick it to his enemy OpenAI.
- everfrustrated 2mo agoTiming. They had a massive amount of compute coming online and a serious pipeline of more arriving.
- Rover222 2mo agoa lot of that compute is used for inference, which is demand-based
- mortenjorck 2mo agoAs a point of comparison: Samsung has sold smartphone chips and later OLED displays to Apple for over fifteen years. Deals like this that look awkward from the outside but are mutually beneficial to both participants exist everywhere.
- ipaddr 2mo agoImagine 2.0 is out as well. Reviews say people look like plastic. Image and video generation quality is extremely low.
- t1234s 2mo agoReading the SWE bickering back and fourth in this thread about Claude vs Grok reminds me of IE vs Netscape bickering way back when.
- andrewinardeer 2mo agoNetscape 4 Lyf
- sillysaurusx 2mo agoIt still survives in the cookie jar format.
- not_a_bot_4sho 2mo ago/me eating popcorn from my Lynx term
- thegagne 2mo ago/me emailing my hot takes and emoticons to newsletters from pine
- dmode 2mo agoCan someone explain to me what's the point of Grok anymore? I don't understand why we need a third or fourth closed frontier model. It is clear that chatGPT has locked down the consumer play, and may be Gemini is there. Claude has enterprise locked up, followed by chatGPT and Gemini. Enterprise switching costs are notoriously high, and even if they switch, they have chatGPT or Gemini to choose from. Beyond that, you have a vast array of open source models (DeepSeek, Kimi, and now Meta's Spark and Glimmer). So, why would anyone need a third or fourth frontier model and why would SpaceX spends billions in CapEx for a very small market share
- siliconc0w 2mo agoCompetition keeps service quality high and pricing low - even if you aren't using Grok, the mere existence of Grok keeps pricing for whatever provider you use lower and service faster and more reliable.
- dmode 2mo agoThat's fair. But my point is from a business POV, why would SpaceX want to invests hundreds of billions of CapEx on a third or fourth frontier model, which cannot compete with chatGPT and claude on the high end, and getting squeezed by open weight models on the low end
- siliconc0w 2mo agoThat is more FOMO and maybe some end-game where SpaceX essentially owns a vertical slice of ISP+Datacenter+AI+application stack.
- throwaway613746 2mo ago[dead]
- Gigachad 2mo agoScamming investors long enough for the founders to cash out.
- 2mo ago
- deleted 2mo ago[deleted]
- osinix 2mo agoThat is good news for Grok team. However, most of the time cost comparing to the result is secondary, and better results and conclusions can come from mixing AI brains together.
- mwkaufma 2mo ago[flagged]
- theyliesoeasily 2mo agobrilliant, I'm going to mentally append this to all benchmark results from now on.
- LZ_Khan 2mo agoWell this makes me bullish on Gemini if its this easy to reach the frontier
- heaney-555 2mo agoWho said it's easy? xAI staff are putting in 80+ hour weeks and building datacenters faster than anyone.
- LZ_Khan 2mo agotheir entire founding team left recently for one. they are the weakest in mission and dont particularly pay much either
- heaney-555 2mo agoIf their entire founding team left and now Grok has caught up to the frontier, that says more about the founding team than xAI.
- btreecat 2mo agoWhen you cut regulatory corners, it does get easier to move quickly
- suttree 2mo ago[flagged]
- maxdo 2mo agoCursor ultra is great . For 200 you got essentially unlimited capacity vs Claude. I used auto in cursor it’s much faster va Claude code and as good.
- mpalczewski 2mo agoI've been using grok 4.5 with grok build soon after it came out and dropped claude. primarily for personal code. It communicates better. While that might not sound like a big deal it is. It doesn't give me a wall of text, tells me what I need to know and I'll make the actual decisions. It is very quick as well which means the sessions are far more interactive, I'll be steering it more. I sometimes cross check with codex and sol, but the daily driver is grok for me. I found it has improved my productivity and output over claude where it felt like claude was giving me work to do. furthermore with the recent claude watermarking thing, I'd rather use grok or openai. If anyone is curious download grok cli and throw a couple of prompts at it. you'll be surprised.
- Rover222 2mo agoAgreed, I've been using it on personal projects, and prefer it to Claude and GPT at the moment.
- nopurpose 2mo agoSounds like a caveman skill.
- zingababba 2mo agoSame, sad how far Claude has fallen. I still think Claude code patterns are amazing so I just port those.
- 3abiton 2mo ago> Same, sad how far Claude has fallen. I still think Claude code patterns are amazing so I just port those. I am curious, is that a plugin, or skill?
- soVeryTired 2mo agoOut of interest, do Musk's politics impact your decision on whether or not to use Grok? I'd be interested to know where folks lie on the (Agree / Disagree) and (Use / Don't use) axes.
- onesandofgrain 2mo agoDamn hn is full of ai shilling, jesus fucking christ
- touwer 2mo ago[flagged]
- Rover222 2mo agothought police on patrol. maybe look in the mirror and whisper "who is really the fascist?"
- theyliesoeasily 2mo agoThe main issue I have with grok is that Musk, its owner, did 2x seig heil at the presidential inauguration, and proceeded to gaslight the world about it (this is a strong form of dogwhistling, kind of a dog bullhorn). Therefore all services which have anything to do with Musk are ineligible for use - they are directly funding the worst kind of person.
- Rover222 2mo ago[flagged]
- theyliesoeasily 2mo agoWhy? The timing was precise, the motion was precise, his facial features were grimaced with intense determination. I cannot fathom calling it anything else, and I see denying it as a kind of shibboleth for "we know but we are pretending it wasn't, wink wink nudge nudge". Either it was deliberate, or the richest man on earth is so incompetent that he accidentally made a motion exactly mimicking a seig heil while welcoming a self-proclaimed king and dictator. Twice - once towards the audience, and once towards the flag of the united states. Neither option is good. What was it?
- Rover222 2mo agoYou can find pictures of AOC and Mamdani doing the exact same gesture. It’s obviously a common gesture when emoting to crowds. But you just want to believe the narrative. If you’re worried about Nazis, take a look at all the pro-Hamas people.
- theyliesoeasily 2mo agoI've seen the claimed pictures in context with their videos, they are distinctly different gestures in timing and emphasis - normal waves, without the sharp hand to chest -> locked elbow with hand outstretched motion. Again, the seig heil is a particular gesture with a particular emphasis and timing which Musk imitated precisely, twice in quick succession. Those other photos are decontextualized from gestures which were clearly different motions entirely. You're spouting easily refuted nonsense, and then immediately making ad hominem attacks because your points are extremely weak and not backed up by the evidence you vaguely cite. You're also getting ratioed because people here are not idiotic ideologues. Please do better. --- I'll also leave you with a nice quote from Sartre which was directed at the fascists of the time - we've seen this shit before, we know what you're doing: “Never believe that anti-Semites are completely unaware of the absurdity of their replies. They know that their remarks are frivolous, open to challenge. But they are amusing themselves, for it is their adversary who is obliged to use words responsibly, since he believes in words. The anti-Semites have the right to play. They even like to play with discourse for, by giving ridiculous reasons, they discredit the seriousness of their interlocutors. They delight in acting in bad faith, since they seek not to persuade by sound argument but to intimidate and disconcert. If you press them too closely, they will abruptly fall silent, loftily indicating by some phrase that the time for argument is past.”
- dizlexic 2mo agoGemini 3.5 flash lite is all I use.
- dudeinhawaii 2mo agoGrok is quite interesting. I run comparisons almost daily on tasks and Grok is its own beast, in a good way. It's good to have model diversity. When I run a task across Sol, Terra, and Luna, I get variations of the same thing with diminishing quality. It makes the lineup pointless. Ditto for Anthropic. Gemini-3.6-Flash and 3.1 Pro genuinely behave differently. Opus 5 and Fable are.. cousins. I find that when I want to test a complex creative challenge, having 4 "families" to choose from makes the experience interesting since they will excel in different areas. Grok might implement unique lighting, Opus, elegant primitives, Sol, accurate snowfall in one pass, Gemini, silky movement. Combined, you can pick and choose best. For what its worth, Grok always feels "messy" but finishes. Grok 4.6 though is no longer "smart and fast". It's about as fast as Sol though. A big improvement I noticed in 4.6 was tool use for verification. Previously, Opus/Fable were the only models to consistently screenshot things that they can't directly interact with easily. Now Grok is probably right behind them, perhaps tied with Sol on propensity to verify visually. Grok 4.5 notably did not do this often.
- 0xEnsp1re 2mo agoStarted using it with grok build cli. So far so good.