8 ms·
Grok 4.7
- jmward01 13d ago[flagged]
- andsoitis 13d agoTry it for software development.
- jmward01 13d agoI have even less trust in their not training on my data/credentials/everything on my computer.
- solid_fuel 13d agoSeriously. They already get caught uploading everyone’s private credentials once before, one would have to be a particularly gullible rube to trust grok again. Especially with musk in charge.
- sejje 13d agoMaybe comment on model releases you've got some experience, or insight about.
- ls1911 13d agoafter using cursor grok & trae.ai for several months , grok curor is highly superior results to trae.ai
- simianwords 13d agoI guess it’s only my opinion but having used grok for personal chat: it’s by far the worst one amongst Claude, ChatGPT and even Deepseek, Gemini etc. The personality is bland and it doesn’t work nearly as hard or even tries to help.
- artemonster 13d agoI used openrouter to send same prompt to qwen, derpseek, gemini and grok and found that grok does good research and produces less bullshit, especially when prompted to be critical of an idea
- xutopia 13d agoAsk it to be critical of the birthday photos and see where that gets you.
- artemonster 13d agocan you elaborate?
- Paracompact 13d agoElon's mother recently posted an AI-generated photo of her son's birthday party. The tag indicating such was scrubbed as soon as it was pointed out.
- slowin 13d agoThis has been my experience as well. Grok will end tasks almost immediately and claim "Done!". It's definitely the laziest and most "dishonest" of all the models. The others aren't perfect, but I can't use Grok for any serious coding task.
- Capricorn2481 13d ago> The personality is bland I don't use Grok, but do you want your LLM to have a personality? "Personality" is exactly what people don't like about Claude.
- Razengan 13d agoI want my sexbot to have a personality
- moojacob 13d agoApparently Grok 4.7 has 40% more weights than Grok 4.6, but the price ($6 output token, $2 input) is the same. Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise. However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby. My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.
- jasonjmcghee 13d agoFor what it's worth - over the last few years or whatever, it seems like Anthropic benchmaxxes the least. That being said, I currently prefer Sol / Astra to Opus / Fable as I find both to be a better cost payoff to me.
- vessenes 13d agoI was going to say the reverse - claude has been the less satisfying normalized by benchmark for me in the last year. Both astra and fable have their quirks, but I am 90% codex this year up from 10% last year.
- vintermann 13d agoIt's not just about benchmaxxing. Sincerely targeting those long-autonomy benchmarks is questionable in the first place, because naturally it drives the model to assume more and more about what you want.
- svachalek 13d ago
- kristofferR 13d agoWhat's with the deceptive graph on top? Not including Astra can't have been an oversight, did the model compare poorly to it?
- Iolaum 13d agoI wonder if that means that SpaceX evals show that they consider astra better than fable or that they hate Sam&co so much they don't want to show their stuff.
- babelfish 13d agothis is exactly it.
- Jcampuzano2 13d agohttps://openai.com/index/our-decision-on-cursor-following-its-acquisition-by-spacex/ https://openai.com/index/our-decision-on-cursor-following-it... Its because of this. You can't use Astra in Cursor, and cursorbench uses cursor as the harness. They can't actually benchmark it using their harness hence why its not included.
- babelfish 13d agoThey have Astra in other benchmarks lower on the page. They just don't want to show it winning
- Jcampuzano2 13d agoThe chart is cursorbench though and they asked about the "deceptive graph"
- ryeguy 13d agoThey can benchmark it because you can use an openai api key with cursor. Astra is just not included in the cursor plan.
- bluecalm 13d ago
- toader 13d ago[flagged]
- AtlanticThird 13d agoWeird, that's the main reason I purchase all of Elon's products https://time.com/5936036/secret-2020-election-campaign/ https://time.com/5936036/secret-2020-election-campaign/
- thoman23 13d agoПривет, fellow American!
- ctrlkctrls 13d agoJudging by Elon's staggering success in all of his ventures I'd say you're out of touch.
- thereitgoes456 13d agoHe has had many failures, SolarCity and xAI and X and DOGE to name a few, but he has often bailed them out with his larger ventures. Even with his successes (Tesla, SpaceX) he has built them up in large part by bending levers of government to his advantage.
- voidfunc 13d ago> Even with his successes (Tesla, SpaceX) he has built them up in large part by bending levers of government to his advantage. So what? Thats called being a maverick. He is very very good at executing on making money which is the point of business.
- andsoitis 13d ago> He is very very good at executing on making money which is the point of business. Also pushing technology forward.
- sidgtm 13d agoIn my experience Grok especially inside Grok build is pretty solid choice, it’s a no nonsense model and stays on its course. Another surface where I truly enjoy the experience of using Grok model is Grok bot
- guywithahat 13d agoI've had really good experiences with Grok 4.6 and grok build. I've been playing around with tscircuit and it can write code with an understanding of spacial reasoning, while also importing cad components from different file formats into tsx, I've been having claude come in and try to error check it and so far claude hasn't found anything to improve in my three projects. I'm excited for 4.7 although I share skepticism with other users whether 4.7 will be significantly better, since they didn't raise the price.
- vessenes 13d agoNice to see this release cadence increasing and some continued improvement in quality. I am guessing these models are basically still outcomes of the cursor team integrating with the massive amount of compute they now own: I’d imagine we will see significant step up improvements with grok 5 later this year as the team gets more experienced and confident with larger training deployments. Here’s hoping for another competitive frontier model!
- Tsarp 13d agoWaiting on simonw "Generate an SVG of a pelican riding a bicycle " benchmark to judge this model
- rvz 13d ago[flagged]
- jcims 13d agoWe're allowed to have our ceremonies.
- kridsdale3 13d agoThank you. If this whole thing isn't fun, it isn't worth doing.
- user43928 13d agoYou don't think it's useful to learn whether a model's "intelligence" generalizes beyond the tasks and modalities it is usually optimized for?
- TylerE 13d agoAbsolutely not. Makes about as much sense as judging a car based on how good an airplane it makes.
- lumirth 13d agoHave you considered that the single most impressive breakthrough of LLMs as a technology is their ability to generalize beyond what they were explicitly trained on? Great analogy, pal, but LLMs aren't cars.
- user43928 13d agoI disagree. If GPT-7 can draw the Mona Lisa in MS Paint via computer use, this would be interesting. That it isn't the most efficient way to achieve the same end result is irrelevant.
- Saline9515 13d agoI tried in Omp (Oh-my-pi), and so far it's really problematic. It will loop in thinking mode ("Let me implement those fixes: Fix 1, Fix 2, Fix 3 .... Fix 80, Fix 81"), ignore the AGENTS.md instructions, corrupt plan files, etc etc... I have 5.6 Sol as advisor/watchdog, and it blocks every turn, I never saw this. Quite a shame, 4.6 wasn't so bad.
- xmorse 13d agoOMP is a joke. don't use that garbage
- Saline9515 13d agoCan you explain your opinion? I'm curious but such vague comments won't convince me.
- samtheprogram 13d agoProbably the same reason as oh-my-zsh, you don't need 90% of it. Further compounding the problem in an agent harness is that you are polluting the context window by throwing the kitchen sink at it.
- marwatk 13d agoI've been experimenting with omp because: - it allows different models within one session via roles (I only have API, so pay per token) - it's much more likely (ime) to use the LSP over grep for determining how code fits together But I agree a 20k+ starting context is way overkill. I find it's very hard to get information on harnesses people are using. I have to stay model agnostic so I avoid claude, codex, cursor, etc. I've used and tried opencode, which worked well, but obviously lacks the above features. Does anyone have a resource for following what people are actually being productive with? With so much vibe going on it's hard to separate the wheat from the chaff.
- xmorse 13d agothis summarizes the average OMP user and dev https://x.com/greg_horvay/status/2100764473392820433?s=20 https://x.com/greg_horvay/status/2100764473392820433?s=20
- andsoitis 13d agoCongratulations to the team!
- AM1010101 13d agoDid 4.6 not have an x-high reasoning level? Why are they comparing 4.7 x-high with 4.6 high?
- ssutch3 13d agoIt did not. xhigh is new to grok.
- forgot-my-pw 13d agoNot sure on the API side, in Cursor you can always use 4.6 at xhigh.
- ssutch3 13d agoWe've only used it through API - but you're right, now API supports xhigh for 4.5-4.7.
- deleted 13d ago[deleted]
- everfrustrated 13d agoI think 4.6 got an xhigh after launch. The benchmarks seem to all have been against 4.6 high.
- maz1b 13d agoEither way, the fact that xAI or SpaceXAI or whatever the name is, I can commend the team behind it on their rapid ascent and progress by being close and or on the frontier in several respects.
- avazhi 13d agoYour comment is like 6 months to a year late. There for awhile it seemed like we’d have 3 big competitors but then Grok 4.2 or 4.4 was just diabolical while OAI and Claude continued their significant improvements. Grok was/is so bad that I was convinced musk was gonna shut it down and just fund Anthropic compute once they reached their compute agreement.
- 6thbit 13d ago( why is the x-axis on the first chart in descending order ? )
- meerita 13d agoGrok it's really expensive. I'm getting really amazing results using DeepSeek 4.1 Flash for fraction of the price.
- parineum 13d agoBrought to you by...
- meerita 13d agoBy no one. For the price of 1M token you can get more and with better results with other models.
- includenotfound 13d agoSure, if you're doing easy work. But Grok is a lot more intelligent and can handle harder tasks.
- testfrequency 13d agoWhat is the most secure way to use this model as someone who is lazy
- user43928 13d agoI understand DeepSeek 4.1 Flash is available on US providers with Zero Data Retention if that is what you are asking.
- sparkling 13d agoYes, but with subpar caching and higher cached token pricing, compared to directly using the DeepSeek platform.
- brcmthrowaway 13d agoLink for the lazy?
- zug_zug 13d agoWell I "tried it out" I asked it one question, and it gave no answer and said "Sign up to use more!" I don't think I'll be doing that, no. I can't think of a single dimension grok is winning on (capability, cost, voice), but want to stay open-minded -- anybody want to vouch for its capabilities in any domain?
- grim_io 13d agoIt's probably the most aligned (to a single person) model out there!
- puszczyk 13d agoFor me it works well for agentic coding tasks and terminal/unix/bash (in cursor and grok build); it's also token efficient and cheaper than gpt 5.6. It's def not as good as Fable for me (I haven't used Astra much, can't comment). So it's not the cheapest, not the most capable, but it has a good mix of it for my backend, go, infra work. The voice is the same AI slop as the others imho. (This is about Grok 4.6, I didn't test 4.7 yet). edit: clarified I mean agentic coding tasks
- sejje 13d agoIf you haven't used it, how do you know if it's winning? I think it's winning on UI for normies (grok bot) and they made some claims about being pareto SOTA (lowest cost per task completed) a while back with 4.6. I find it to be a perfectly capable model for implementation (there are many in this class--deepseek flash, spark1.3, luna, etc). I find the usage to be very generous w/ supergrok. I find the model to be just fine for 90% of what I want to do, but I use a smarter model to plan complicated things.
- bluepeter 13d ago[dead]
- simonw 13d ago$2/million inout and $6/million output but I couldn't see any pricing information for cached input tokens?
- MuffinFlavored 13d agoIf the CursorBench 4.0 score diagram is the headline, I read it as "Grok 4.7 xHigh is almost the same as Fable5.1 on low". Is there a metric for like... time taken when comparing these two? I see score and cost. If Fable5.1 can knock it out more quickly on low but Grok4.7 might take twice as long to stumble through a problem (and leave behind a bunch of yucky comments or un-needed extra unit tests), are they really comparable? Or like... the "quality" of the solution? "It works" versus "it's unmaintainable/very messy/hacky".
- WarmWash 13d agoGood thing they used 5.6 sol instead of Astra for benchmarks, the EEbench one is crazy[1] [1]https://eebench.org/ https://eebench.org/
- saejox 13d agoNot even close to astra. Astra is something else. It is expensive, but uses way fewer tokens do my tasks. xAI missed its chance, Ball is on Anthropic's court.
- enraged_camel 13d agoAstra fails in similar ways, and at similar frequency, as GPT 5.6 Sol does. It often goes way out of scope, or just stops prematurely, or tries to find odd and even dangerous workarounds when it gets stuck. It's phenomenal at computer use and 3D stuff. I've been using it less and less for coding.
- brink 13d agoSame, Astra is extremely RL fried, and nobody is talking about it. I used Astra for a few days on my personal project, and load times went from less than 3 seconds to almost 30 seconds because it kept using the wrong sync primitives and bad architecture overall.
- haellsigh 13d agoHuh, I've had a totally different experience. I've used it extensively, maxing out the 200€ plan on personal projects and it's the best model I've ever used, so easy and pleasant to use. It's great for frontend design and using it in Rust I've had Coming from Opus 5, it's a breath of fresh air.
- redox99 13d agoNot surprising considering Grok 4.7 is a 2T model, so Sol/Opus class, not Astra/Fable class.
- jstummbillig 13d agoHow many parameters do Astra or Fable have?
- 13d ago
- simonw 13d agohttps://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F58dc4e8ec482330856fce89dac670727 https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - default reasoning level. Here's reasoning level high: https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F8ab126bda2b384264b3ad931e3ebb8b4 https://tools.simonwillison.net/markdown-svg-renderer?url=ht... For some reason reasoning effort low and medium used similar numbers of tokens, and xhigh used less than high. I think I need to try without OpenRouter in the middle. UPDATE: I tried again with the xAI API directly: https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F5a1819a2bd24bb642f38c4bc6733090f https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - not a great deal of difference between reasoning levels, and this time xhigh and low used the same number of reasoning tokens for some reason. For comparison here's a fresh run against Grok 4.6: https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ffedc404b9aa8e6fca31d59c898dabba0 https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
- MattDamonSpace 13d agoAre there good tools for doing context audits? I feel I have no good way to visualize what a new session is getting by default in a given repo without crawling through every potentially included markdown file
- datsci_est_2015 13d agoPoor fella doesn’t have a seat. Intriguing design where both pedals are on the same side of the frame. Balancing must be a challenge.
- kiliancs 13d agoWhat is the default reasoning level?
- forgot-my-pw 13d agoI tried in Cursor and see a lot of improvements over Grok 4.6 svgs. The AA numbers indicate it's not very token efficient though: https://artificialanalysis.ai/agents/coding-agents?agents=codex-deepseek-v4-pro-0813-max%2Ccodex-gpt-6-astra-max-reasoning-effort-max%2Cclaude-code-opus-5-max%2Ccodex-gpt-5-6-sol-max-reasoning-effort-max%2Cdevin-fusion-cli-claude-fable-5-1-xhigh-swe-2-medium%2Cclaude-code-qwen3-8-max%2Cdevin-fusion-cli-gpt-6-astra-xhigh-swe-2-medium%2Cantigravity-sdk-gemini-3-8-flash-high%2Cmuse-code-muse-spark-1-3-max%2Ckimi-code-cli-kimi-k3%2Cclaude-code-fable-5-1-max-with-fallback%2Copencode-glm-5-3-reasoning-effort-max%2Cgrok-build-grok-4-7-xhigh%2Cgrok-build-grok-4-6-xhigh#coding-agents-token-usage-chart-tabs https://artificialanalysis.ai/agents/coding-agents?agents=co...
- dom96 13d agoIt’s a shame this model has such negative political baggage associated with it. It’s the only one I decided not to run in my LLM benchmarks[1]. 1 - https://bench.killswitch-lang.org https://bench.killswitch-lang.org
- sejje 13d agoYou'll have to include it in the future, or your benchmark won't be relevant. For now, I doubt anyone would notice your protest if you didn't announce it.
- lirolero 13d ago[dead]
- peder 13d agoI think you're seeing a big shift around it.... since it's been markedly cheaper and also still easily available from OpenCode, it's getting large enterprise traction.
- mempko 13d agoYes, and that's a bad thing.
- mempko 13d agoNot sure why you are being downvoted. Until Musk owns up to his Nazi salute, I won't have anything to do with Grok, no matter how good or cheap it is. And yes, we need to keep talking about this because it's absurd.
- felixgallo 13d ago[flagged]
- eleventen 13d ago[flagged]
- drop_star 13d agoI wont touch his products and neither will my organization
- ForrestN 13d agoI completely agree. But this has been true for many years. This sort of head in the sand compartmentalization seems to be a core feature of the culture here.
- moolcool 13d agoIt’s either compartmentalization, or something else
- mlindner 13d agoI have to say I'm a little tired of whenever a Musk related product comes up there's random nolifes that arrive to rant about politics. Luckily they're relatively rare on hacker news. Also it's kinda hilarious how you think any money spent on Grok will go toward furthering climate change versus literally any other AI model that does the same thing. Grok at least seems to be more efficient than most models.
- grokgrokgrok 13d agoI have to say grok, grok, grok, grok, grok. Also, anyone who doesn't modulate across models and run their own memory system is an idiot.
- TheOtherHobbes 13d agoMusk is literally burning methane for funsies, and generating CSAM and getting sued for it. Handing corporate code secrets to his AI model is... unusually trusting.
- deleted 13d ago[deleted]
- TylerJaacks 13d ago[flagged]
- 786562354238 13d agoDid you come up with that by yourself?
- TuxSH 13d agohttps://knowyourmeme.com/memes/mechahitler-grok https://knowyourmeme.com/memes/mechahitler-grok
- mrtesthah 13d agoEveryone should be clear that this is what they’re cheering on when they celebrate a Grok performance win. A technology is no longer neutral when wielded by a self-proclaimed white supremacist whose actions have killed over a million black and brown people, mostly children and babies.
- oulipo 13d agoExactly. And it's DEEPLY DISTURBING that all comments on HN that point out that X is pro-nazi no longer has an upvote button
- mavamaarten 13d agoYeah. I'm actively avoiding giving mr far right any $$
- swalsh 13d agoCodex has become my goto tooling. I used to be a Claude Max subscriber, but I was becoming disappointed with the quality of the output from Opus 5. Fable chewed through my usage too quickly to be practical. Moving to a Pro account w/ Codex was a big improvement. Sol had great output, and the usage was more than sufficient for most of my needs. However astra does tend to chew up usage, so when i've done to much of that, and it's became an issue Grok Build has beocme my second go to account. The output especially after the cursor purhcase has become quite good, and the usage has always been very generous.
- johnfahey 13d agoNo doubt xAI has seen rapid progress, but it's been several months of them being "just behind" OpenAI and Anthropic. It seems the gap between just behind the frontier and pushing it is a lot wider than most people thought it was a year ago, and that's why a clear third contender in the frontier model space has yet to materialize.
- shdtabasum 13d agoWhy Chinese models from Kimi, Deepseek are not added in comparison benchmarks?
- xquce 13d agoSame reason Coca-Cola only mention Pepsi and Pepsi only mention Coca-Cola. It's an proven way to capture the market. You would rather split the pie in two rather than in 4,12 or 50 right?
- thih9 13d agoI refuse to use Grok. Mostly because of the usual reasons - somehow this high profile AI model seems more disgusting than others and it is in a way impressive. But also Xai doesn’t seem to care about user experience and long term support.
- swozey 13d agoI can't take anyone seriously who uses grok seriously. I like to look at the cybertruck owners forum every so often because it's just... hilarious. And the amount of superfluous grok use over there is just insane. Half the posts I click in there will have a bunch of people dumping entire grok takes "why do people hate cybertruck owners?" "Because they're jealous and poor," sort of stuff that they just LOVE to post. As a technical point of reference to compare against other llm stuff, sure, I'll glance at a report or benchmark but I really couldn't care less about anything to do with the project and it could blow other options away and I wouldn't touch it.
- ElectronCharge 13d agoPossibly interestingly, I can't take you seriously for having such a superficial approach. You probably shouldn't cut off your nose to spite your face.
- mempko 13d agoI don't know man, Musk doing Nazi salutes doesn't seem that superficial. He did help get Trump in power and also killed a lot of aid to children that need it. What's superficial about refusing to use a product from someone like that? Or are you one of those 'technology isn't about politics' people? That's a superficial take if you ask me. All technology is political, and understanding that is a deep, not superficial take. It requires systems thinking which unfortunately many people building technology seem to lack, despite software being a sophisticated complex system.
- brandonagr2 13d agoYou should try it, it is less sycophantic than other models and is faster and better at most reasoning levels, don't confuse the twitter bots and services also named Grok with the frontier model itself
- gslepak 13d agoDoes anyone have any experience with Grok's subscription? How does it compare price-wise to the API?
- andreyvit 13d agoWell when I ran out of Grok SuperHeavy subscription ($300) once and tried to use extra credits to cover half a day remaining till reset, $50 in extra credits went in two hours. Based on that, subscription definitely lasts longer; Grok subscription just about covers a week of my work (sometimes a bit extra remains unused, sometimes it runs out half a day to a day early). And as a point of comparison, it lasts for doing same tasks as 2.5-3 weekly limits of Codex on 5.6 Sol did (using xhigh on both Sol and Grok); I needed 3x$200 Codex subscriptions to cover my weekly usage.
- nwienert 13d agoBy far the worst value subscription of any. I tried Superheavy and got about 5-10% the usage of CC/Codex.
- everfrustrated 13d agoI find I can just about get by with coding every day on a Cursor $60/mth sub with Grok fast mode disabled. Doing pretty heavy coding work/requirements etc, but not much sub agents and no loops. For me and what I’m doing that’s insanely good value. I find grok build chews through my SuperGrok sub very quick - but I think that is due to it having the 500k context window which uses more credits. Cursor limits it to 256K (tho I see in today’s update for Grok 4.7 there’s now a toggle for context size).
- thefourthchime 13d agoThere are two ways to subscribe, and it’s very confusing, but the best value is to get cursor ultra for $200 a month. I basically have infinite tokens with that plan, plus grok bot, which I really like
- daquisu 13d agoThere are some users reporting it improved a lot in the last few weeks. The max sub usage for Grok is around $12,000 of API pricing now, so a 40x multiplier for the $300 plan. It is the same multiplier for Sol with subscription. For Astra though the multiplier is ≈20x, so half of Sol usage. For Claude it seems to be ≈40x too for Opus, but less for Fable (similar to Astra in GPT). All on the most expensive plan. Previously, Grok usage escalated linearly from the $100 plan to $300 plan. That would be a really good $100 plan if it is still true. Some sources: 1. https://x.com/kunchenguid/status/2098256018836963382 https://x.com/kunchenguid/status/2098256018836963382 2. https://x.com/stevenzhang/status/2092110386569089311 https://x.com/stevenzhang/status/2092110386569089311 3. https://github.com/openai/codex/issues/43731 https://github.com/openai/codex/issues/43731 4. https://redd.it/1wciwc1 https://redd.it/1wciwc1 5. https://x.com/SemiAnalysis_/status/2064815044085318040 https://x.com/SemiAnalysis_/status/2064815044085318040 6. https://redd.it/1vx0k69 https://redd.it/1vx0k69
- notduckrabbit 13d agoSignificant regression in token efficiency compared to Grok 4.6 suggested by artificialanalysis.ai Intelligence Index Comparisons.
- everfrustrated 13d agoThat is comparing Grok 4.6 high to Grok 4.7 xhigh tho.
- notduckrabbit 13d agoNo, you can add Grox 4.7 high to the chart. 36k vs 66k
- sourcecodeplz 13d agoOutput tokens from Intelligence Index: - grok 4.6 (xhigh): 97M (for 44 score) - grok 4.7 (xhigh): 240M (for 46 score)
- c0rruptbytes 13d agoas someone who is limited by amazon bedrock support at work (no idea why we got stuck with the worst one) - grok is literally the only budget-ish model option, so nice to see it updated, Sol and Opus are just too rich for my blood. Luna is good but so slow at getting things done (tps wise it's fast)
- GodelNumbering 13d agoEvery Grok release obscures their cache pricing while highlighting their input/output pricing From their headline comparison: Grok: $2/$6 per million Fable: $10/$50 per million What this doesn't say: Grok costs 0.50/M cache read, Fable $0.25/M cache read Long running agentic workflows are dominated by cache reads. Just makes Grok sound deceptive, and more importantly, reliant on user's lack of understanding of costs aka predatory (which in turn is more infuriating)
- sourcecodeplz 13d agomuse spark 1.3 contribs cache read is $0.002 btw (~220x diff).
- qwerpy 13d agoI've been using 4.6 for some one-off game mods/utilities and it has done very well. "I have a very niche keyboard (Moonlander) and I play this very niche space sim, make me a SVG keyboard cheatsheet for it". Told me to grab keymap.c for the keyboard and inputmap.xml for the game's key bindings, churned for a while, then spit out a pretty good first attempt, along with the python script used to generate it. Spent another hour of back and forth to refine the script, and now it generates great diagrams that will adapt as my keyboard firmware and game bindings evolve: https://files.catbox.moe/x0u76x.svg https://files.catbox.moe/x0u76x.svg Excited to try 4.7. I hope they fixed the "it's not X, it's Y" that showed up in 4.6.
- Theodores 13d agoImpressive! I had to peek at the SVG file and it superficially looks good, however, as is the case with everything AI, the more you look, the more it doesn't make any sense. By now AI should know of the DRY concept. But no. Hence the keys have a rounded rectangle for the key shape and another rounded rectangle for a clip path, to prevent text overflow. There are 72 * 2 = 144 identical rectangles, when just one would suffice (in the defs), with this being cloned once for the clip path, and 72 times for the keys. I would not expect SVGO levels of optimisation (rounding numbers, that sort of thing), however, the human, if writing out the same thing for the 72nd time, might think 'is there a better way', to get the manual out. A graphics program such as Illustrator would not do that, but AI 'should' because AI. The above is not criticism of your work, just an observation regarding AI SVG capabilities.
- enraged_camel 13d agoThis thing is DOA. They compared 4.7 xhigh to 4.6 high to make it look like it improved. The reality is pretty bad: https://x.com/chetaslua/status/2102087511367618942 https://x.com/chetaslua/status/2102087511367618942
- trentor 13d agoLooks like they have still problems with caching. Prize is double the other providers for cache hits... which is most of what I do. :/
- deleted 13d ago[deleted]
- BoumTAC 13d agoVals AI just affirm that Grok 4.7 is worse than Grok 4.6 (It ranks #24 on the Vals Index at 54.2%, down 5.0 points from Grok 4.6 (#14, 59.2%)) https://x.com/ValsAI/status/2102086608476590432 https://x.com/ValsAI/status/2102086608476590432
- nostrebored 13d agoIs this an ad for Vals AI? Looking at their website, the rankings don't mesh with my observed utility for almost any model outside of fable and astra being good-ish.
- BoumTAC 13d agoAbsolutely not. I know Elon retweet them a lot when Grok is good. This is how I discover the company. I like to follow them and look for benchmark for each LLM release.
- nostrebored 13d agoAh gotcha, not on twitter so just hadn't seen them before!
- brcmthrowaway 13d agoDumb question. Are these products really winner-take-all? Why is there such a furious rate of development?
- hdhdjdif 13d agobecause boomers will give you free money + tip musk can fund the space stuff with this
- sourcecodeplz 13d agolooks like token efficient/verbosity took a big hit. Output tokens from Intelligence Index: - grok 4.6 (xhigh): 97M (for 44 score) - grok 4.7 (xhigh): 240M (for 46 score)
- oh_no 13d agowhich is crazy because this was grok's competitive advantage, worse than OpenAI models but better than everything else, now it's less efficient than Opus or Fable 5.1
- usumgallu 13d ago[dead]
- gaigalas 13d agoPacing the frontier, with an aggressive release cadence. Gotta love the US tech industry.
- oh_no 13d agothe AA numbers are generationally bad. double token use (the one thing Grok was good at was low reasoning usage!) to gain 5% in the benchmark score. with reportedly a larger model. maybe it shows gains IRL but wow, I've never seen a new generation model look so underwhelming compared to the last.
- johnnyApplePRNG 13d agoI thought elon had agreed to "pace the frontier" along with the rest of the gatekeepers? And then he releases a stronger model like a week later? Fuck these jokers
- alansaber 13d agoAs anthropic/openai subscription allocations get squeezed you'll see more people using "second rate" closed models like grok. The token allowance with a Cursor subscription is crazy.
- jascha_eng 13d ago32 on the omniscience index. Not terrible but far from Astra and fable: https://artificialanalysis.ai/evaluations/omniscience https://artificialanalysis.ai/evaluations/omniscience
- mempko 13d agoUntil Musk owns up to his Nazi salute, I won't be using Grok, sorry. I don't care how good or cheap it is. And no, I won't stop talking about it either.
- DaSHacka 13d agook.
- inshard 13d agoAny real world experience with Grok Ultra $300 monthly subscription vs Claude Code Max in terms of overall built work mileage, or general token limits?
- Invictus0 13d agoSpaceX AI releasing "Grok" has to be some of the worst branding I've seen in my lifetime
- outside1234 13d agoWho uses this trash?
- jjcm 13d agoIt's definitely gotten better at image->html workflows. Here's a test comparing Astra (currently SOTA at this) vs Grok 4.7: Designs: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.webp https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we... Astra's build: https://html.non.io/annui/ https://html.non.io/annui/ Grok's build: https://html.non.io/Annui-grok/ https://html.non.io/Annui-grok/ Additional prompt instructions: "Add scrolling clouds behind the statues. Dynamically light the statues based on mouse position. Use diffui to generate the normal maps/depth maps/roughness maps of the objects, and to separate out the assets on to different layers." Overall I find these models are getting good at following image as a source of instructions, but their refinement of the output varies heavily between the models. Astra's final output feels more polished, has better visual contrast, and the animations between the pages are smoother. Grok also chose to light all of the background elements, which imo overcooks it a bit. Still though, for the price it's a great starting point.