3 ms·
Apparently Grok 4.7 has 40% more weights than Grok 4.6, but the price ($6 output token, $2 input) is the same. Given that the decrease in their margin and the
by moojacob 14d ago
Apparently Grok 4.7 has 40% more weights than Grok 4.6, but the price ($6 output token, $2 input) is the same.
Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise.
However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.
My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.
- jasonjmcghee 14d agoFor what it's worth - over the last few years or whatever, it seems like Anthropic benchmaxxes the least. That being said, I currently prefer Sol / Astra to Opus / Fable as I find both to be a better cost payoff to me.
- vessenes 14d agoI was going to say the reverse - claude has been the less satisfying normalized by benchmark for me in the last year. Both astra and fable have their quirks, but I am 90% codex this year up from 10% last year.
- vintermann 14d agoIt's not just about benchmaxxing. Sincerely targeting those long-autonomy benchmarks is questionable in the first place, because naturally it drives the model to assume more and more about what you want.
- svachalek 14d agoThe target market for frontier models is CEOs who want to lay off entire departments of their company. So the long autonomy benchmarks would seem to be sending exactly the right signal.
- vintermann 14d agoThey're still going to have to communicate with the bots replacing those departments they lay off, or they're going to have a bad time.
- dumberquestions 14d agoToken price doesn't tell you much without knowing token efficiency.
- deleted 14d ago[deleted]
- user43928 14d agoTheir leading benchmark with cost per task shows a tough sell compared to Fable 5.1 Low and doesn't reach the performance of Fable 5.1 Medium. How representative that is of real world usage, I don't know. In their benchmark GPT 5.6 Sol performs suspiciously poorly compared to the former models.
- deleted 14d ago[deleted]
- Lucasoato 14d ago> I simply cannot stand Claudish I totally agree, it’s like that as models become more intelligent, they are less understandable by most of people... but aren’t we humans doing the same?
- Aperocky 14d agoThe best ideas are usually the simplest to elaborate. If someone comes up with a convoluted scheme that are hard to understand or be adequately explained, it's usually fraud. When claude speak in convoluted mess, they are often going off on tangents in real work that you asked it to do, too.
- fragmede 14d agoThat believes that the world can be simplified into dichotomies, or at least, simplified. Sometimes problems are complex, and the solutions to them necessarily so. For example, cancer. I order to begin to understand that problem, you have to understand the utter complex scheme it has devised in order to exist. A 20 minute YouTube video isn't going to be able to begin to cover the basics of the subject, although there are some good ones, with clever analogies. Just because something is difficult to understand doesn't mean it's fraud, although if someone is trying to dazzle you with clever words and names of institutions you recognize because they are selling you something, there's a good chance they're lying to you in order to get some money from you.
- Pannoniae 14d agoNo but almost all good ideas can be reduced down to a few sentences if you're good at explaining things. It's a different kind of intelligence than what's commonly called IQ but it's something like that regardless. Sure the explanation will oversimplify a lot but then you can expand it recursively if needed, you gotta start somewhere.
- includenotfound 14d ago> That believes that the world can be simplified into dichotomies, or at least, simplified. Sometimes problems are complex, and the solutions to them necessarily so. For example, cancer You just simplified most of the problems people work on down to cancer complexity. Ironic, isn't it? That's also simply not the case, most people are building CRUD apps with some frontend code and some accessory stuff like build systems etc., which while complex, can still be expressed in very plain, easy to understand language for anyone who's a bit technical. Does not excuse the Claude slop.
- atniomn 14d agoI expect the next Anthropic release to finally reduce the prevalence of Claudish
- sscaryterry 14d agoBased on?
- moojacob 14d agoIf they fix Claudish, they've earned me back as a max customer! Fable 5.1 is not there quite there yet. They need to get that Sonnet 3.5 magic back.
- rfgplk 14d agoSame. The issue with Anthropics models is that (speaking regarding code generation) they REFUSE any kind of comment override instructions. I've tried everything and no matter what, after a few turns, they resort to generating the same overtly verbose junk. Bun's codebase is littered with them See // `HANDLE` is an opaque kernel handle (kernel32 validates and returns 0/FALSE // on a non-console handle); every out-param is `&mut T` to a `#[repr(C)]` POD, // ABI-identical to the Win32 `LP*` pointer (thin non-null). The reference type // encodes the only pointer-validity precondition, so `safe fn` discharges the // link-time proof. (`bun_windows_sys::kernel32` declares these with `*mut`; // redeclared locally so the legacy-conhost cursor path below is plain calls.) or // Progress's terminal handle is the canonical `output::File` (vtable-backed // stderr/File from `OutputSinkVTable`). The duplicate `ProgressTerminalVTable` // from B-0 round 1 is removed; tty/ansi/winsize route through the new // `OutputSinkVTable` slots so `bun_core` stays T0 (no `bun_sys` dep). from src/bun_core/Progress.rs
- WarmWash 14d agoPerhaps you haven't had the chance to use it, but 3.8 flash is the best model for talking too. Even routing Claudes output through 3.8 to have it explain whats going on is a breath of fresh air
- AustinDev 14d agogemini 3.8 flash?
- jtwaleson 14d agoyes
- esafak 14d agoI would if they let me bring the subscription I have to the harness of my choice.
- svachalek 14d agoAgreed. It's very capable for something carrying the "flash" label, super fast, and very clear to read.
- moojacob 14d agoI'll have to try Gemini Flash for coding. The reason I haven't I used Gemini for coding is last time I tried it couldn't call tools very well. I am a huge fan of Gemini Pro for chat... gemini somehow just knows the most obscure stuff. I'll double check something Gemini said and find the source is deep inside a hard to access scientific paper. Google just has the best index of the internet.
- xmorse 14d agoit's definitely not bigger. smaller if anything looking at how much faster it is
- smashers1114 14d agoFYI a quick fix for claudish is to ask for the response to be in ASD-STE100 (Simple Technical English). Then it is far more readable. But I would agree that this is an annoyance and shouldn't require user workaround to get something readable.
- _boffin_ 14d agoDoes not work for Claude, at least for me and I put it as the system prompt
- LPisGood 14d agoI don’t think system prompts are particularly reliable way to do much at all. It’s better to put it as a hook after each response, or a skill at least so you can trigger it at will if you don’t want it everytime.
- kekebo 14d agoDo you think they're unreliable based on the position in the conversation or other factors?
- MisterMunchkin 14d agoAnthropic has probably RL’d the system prompt into nothing because of their fear of the user being able to control the model. If it listened to you about the slop language, it might listen to you if you asked it to help you with no-no tasks.
- ffsm8 14d agoIt does work, you however have to put it into every single prompt in which you didn't want a rubbish response Literally every one, even 1-2 prompts later it starts to go back
- bel8 14d agoFor me it works at first but Claude models forgets it after some prompts, despite only using like 100k tokens.
- Waterluvian 14d agoUsing a variety of models feels similar to the benefit of having a team of individuals from different backgrounds.
- tk90 14d ago> I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby. Wonder if we'd benefit from a much more specialized + task-specific benchmarks to paint a clearer picture like this. A benchmark solely for frontend, ruby, hardware, etc.
- dmix 14d agoAgreed, Claude has a "Claude Design" tool but doesn't publish any frontend brenchmarks. Maybe the industry will develop one.
- algoth1 14d agoI've noticed Chatgpt 5.6 Sol High, on the chat interface, inventing words that are a mixture of Portuguese and English. Like "hardcodar" a mix of "hardcode" and the most common verb ending in Portuguese "-ar". Some don't have a single google hit
- shawabawa3 14d agoDo you have any connection to Portugal? I imagine if you have Portuguese in any of your prompts that might bleed into your user profile which becomes a part of every prompt. Alternatively it might use browser language settings
- rayiner 14d ago> My favorite part of the new Groks has been how they speak in plain english. I don't know if it's the plain english or what, but I really like Grok for legal research (as opposed to code). It's got a noticeable edge in getting to the point compared to Opus 5.
- johnsimer 14d agoI've found grok 4.6 speaks heavily in Claudish. It especially likes using verbs as nouns.
- pietz 14d agoLooking at AA and Vals, your theory seems to check out.
- petesergeant 14d agoGrok and Zai have both been excellent as adjunct code-reviews, on their cheapest plans, for me. Fable plans, Opus writes, Codex as primary reviewer, but Grok and Zai usually find something worth fixing that the others have missed. Both are well worth whatever the $20 or so I'm paying for them
- Forgeties79 14d agoI do not understand how anyone can seriously use a tool that has "Be funny and irreverent when appropriate" baked into the system prompt. I don't want to waste money because my calculator is cracking jokes. They don't deserve their paltry 5% marketshare or whatever it is they have currently. I'm not even getting into Musk as a person or the horrid things we've seen Grok spit out on twitter. I just don't trust his companies with my data and I have seen very little evidence that it's ever the best tool for the job. I'm sure those cases exist but I can't imagine it's worth it.
- StilesCrisis 14d agoI am on the exact same page as you, but there is definitely a market for LLMs which speak more conversationally and less like Claude! Non-programming use cases abound and most users don't like the rigid, exact tone that engineering demands.
- attentive 14d ago$0.50 for cache reads, which is 25% of input. While other models are 10% of input. And like that grok4.7 cache reads are more expensive than sol's (at $0.40/mil).
- quater321 14d ago[dead]
- giancarlostoro 14d ago> Claudish I do wonder why a frontier model does this to be honest. It still does good coding wise, but it seems strange to me. r/Claude is full of "load bearing" jokes in every thread.
- quater321 14d ago[dead]