4 ms·
For those who would like to know the total and active parameter count of this model: even though Google doesn't disclose the model technicals, we can infer them
by easygenes 5mo ago
For those who would like to know the total and active parameter count of this model: even though Google doesn't disclose the model technicals, we can infer them within relatively tight margins based on what we do know.
We know they serve the model on TPU 8i, which we have plenty of hard specs for (so we know the key constraints: total memory and bandwidth and compute flops). We can also set a ceiling on the compute complexity and memory demand of the model based on knowing they will be at least as efficient as what is disclosed in the Deepseek V4 Technical Report.
We can also assume that the model was explicitly built to run efficiently in a RadixAttention style batched serving scenario on a single TPU 8i (so no tensor parallelism, etc. to avoid unnecessary overheads... Google explicitly designed the 8th-generation inference architecture to eliminate the need for tensor sharding on mid-sized models).
We know Google intends to serve this model at a floor speed of around 280 tok/s too.
Putting all these pieces together, we can confidently say this model is ~250-300B total, and 10-16B active parameters. Likely mostly FP4 with FP8 where it matters most.
Visual:
┌────────────────────────────────────────────────────────┐
│ TPU 8i VRAM (288 GB) │
├───────────────────────────┬────────────────────────────┤
│ Static Model Weights │ Dynamic Allocations & │
│ (250B - 300B @ Mixed │ Compressed KV Caches │
│ FP4/FP8) │ (RadixAttention / SRAM) │
│ ~110 GB - 150 GB │ ~138 GB - 178 GB │
└───────────────────────────┴────────────────────────────┘
I do model serving optimization work. This is napkin math.
Edit: There's one factor I under-rated in my initial estimate... TurboQuant. This is a compute to KV memory use tradeoff. It's plausible with TurboQuant at a quality-neutral setting they've gotten the model up to 400B with similar economics. This is a variable effecting concurrency and the the way they decided total model size was likely based on what they see for the average user's average KV cache depth in real-world usage.
- zacksiri 5mo agoDo you have similar math for the flash-lite variant of the models? I'd be curious. Based on my testing / benchmark i think it's around the 100-120B mark. With the Pro variant being around 600B - 800B My testing is comparing it's performance / output to other models in the same size range, so not as scientific as yours.
- Maven911 5mo agoTell me more about what your day looks like. What do you think of the LLMOps books from Abi, in case you have read it ? Any other resources you can recommed?
- anthonypasq96 5mo agogiven this, is it safe to assume that inference pricing is barely related to cost to serve at this point and there is considerable margin?
- daemonologist 5mo agoIf this is accurate it raises the question: why is this model so expensive? DeepSeek v4 Flash is 284B total/13B active, FP4/FP8 mixed, and only costs $0.14/$0.28 - even less from OpenRouter. Of course Gemini 3.5 Flash is most likely a better product, and therefore it can command a higher price from an economics perspective, but does this imply Google is taking roughly a 90% profit margin on inference? If so they're either very compute-limited or confident in the model and wanting to recoup training/fixed costs (or both).
- xmonkee 5mo agoWell, we use flash models extensively (both 2.5 and 3.1) and I cannot overstate this, google cannot fucking serve them without 503s 70% of the time on most days I think it’s pure economics. Flash models are OP for the price, leads to too much demand, google cannot serve it. This is likely expensive to reduce load and hey, if it still makes money just keep the margin.
- WarmWash 5mo agoRumor is that GCP was happily selling compute to competitors. After all, under the hood, Google is closer to a federation than a corporation. The state of GCP doesn't care about the state of Gemini.
- happyopossum 5mo ago> Rumor is It’s not a rumor - there are many public announcements about $B deals around compute for other Ai companies
- collabs 5mo ago>> Rumor is that GCP was happily selling compute to competitors. After all, under the hood, Google is closer to a federation than a corporation. The state of GCP doesn't care about the state of Gemini. > It’s not a rumor - there are many public announcements about $B deals around compute for other Ai companies The last time I read a public announcement, the commentary I read was this is because Anthorpic doesn't want to run out of cash or capacity before a funding round / IPO so they gave Google some equity and in return Google gave it some compute resources? You could reframe it as Google is buying into Anthorpic — which is how Claude tends to frame it as but the end result is the same. Equity for spare capacity. You could even argue that at a hyperscaler like Google's scale — all capacity is spare capacity and no capacity is spare capacity at the same time. GCP seems to have deals with Anthorpic, OpenAI, Meta, Apple, Healthcare companies, Banks, LG(?), Best Buy(?) so in my mind Google (and all AI vendors) are hyping up AI to drive up interest and building up as fast as they can to capture that interest and convert it into cold, hard cash. It honestly feels like this is out of my mental capacity though because these AI vendors had the cold hard cash that they spent on these data centers that we don't even know might become obsolete within a decade(?) but I guess meanwhile they could make beaucoup bucks. There is also the idea that they had to be seen as conspicuously spending on AI or investors might see them as falling behind, triggering a selloff. So yeah I guess it is a fact that GCP is selling compute to other AI companies but it makes sense because basically you can build capacity potentially for Gemini to use in the future while having other companies pay for some of that cost today. In my mind, for hyperscalers — Google, Amazon dot com, Microsoft — competitors are not really enemies but rather partners. The real fear or threat is market uncertainty and customers souring on AI altogether. As long as customers are interested in this AI stuff, you could compete on merit or cost benefit ratio but if competitors start failing because they ran out of capacity or cash, that could send an unwanted message to the market. To summarize though, I have to agree that the supposed rumors are better than rumors, they are facts and we could even make an educated guess that this is a part of a strategy, as much as you can strategize when it comes to an "industry" with a high fixed cost and an uncertain demand.
- gertlabs 5mo agoWe've been really impressed with the performance of ~30B parameter class models and how close they are to the frontier from ~6-12 months ago, which begs the question, are the frontier labs really serving 10T parameter models? Seems unlikely. If these Gemini 3.5 numbers are accurate, then I'd wager GPT 5.5 and Opus 4.7 are a lot smaller than people have speculated, too. It's not that frontier labs can't create a 5T+ parameter model, but they don't have the data to optimize a model of that size. Gemini 3.5 Flash is really smart in one-shot coding reasoning, btw. Near the frontier. But it doesn't do so well in long horizon agentic tasks with arbitrary tool availability. This is a common theme with Google models, and the opposite of what we see with Chinese models (start dumb, iterate consistently toward a smart solution). Data at https://gertlabs.com/rankings https://gertlabs.com/rankings
- easygenes 5mo agoWe know from NVIDIA's public Vera Rubin inference engine marketing materials that the frontier lab models are ~1-2T total. Mythos is an exception that's larger.
- MisterPea 5mo agoI exclusively use gemini models and this has been my experience. I mitigate it by creating dense planning docs for everything and executing iteratively. Lot's of time wasted on procedure unfortunately
- beacon294 5mo agoI agree with this sentiment but the reasoned anecdotes do not agree. I imagine the flagship models have modalities/usages that we hn-ers don't imagine easily.
- nl 5mo agoElon says Opus is 5T (and I would expect he'd know) > It's not that frontier labs can't create a 5T+ parameter model, but they don't have the data to optimize a model of that size. The have plenty if data. They use very large amounts of verifiable synthetic data in (lots in coding and math) cover the gap. Also the frontier labs are paying people to do tasks, tracking the trajectories and training on that. Most of the optimization is in RL based on these trajectories.
- nilstenura 5mo ago[flagged]
- smnscu 5mo agoNice post! You piqued my curiosity, so after a bit of research it turns out that, with techniques like MTP/MLA/CSA, it's quite probable that these models are much more efficient (and maybe bigger? tho 400B sounds about right) than a simple RAM breakdown would suggest. MTP - https://blog.google/innovation-and-ai/technology/developers-tools/multi-token-prediction-gemma-4/ https://blog.google/innovation-and-ai/technology/developers-... MLA - https://machinelearningmastery.com/a-gentle-introduction-to-multi-head-latent-attention-mla/ https://machinelearningmastery.com/a-gentle-introduction-to-... CSA - https://deepseek.ai/blog/deepseek-v4-compressed-attention https://deepseek.ai/blog/deepseek-v4-compressed-attention
- Doxon 5mo agoThese techniques are used by DeepSeek, and work well with the commodity (NVIDIA) GPU's they use. Google designs their entire AI stack from the custom silicon up. So they have different optimization approaches. (Though Gemma does use MTP)
- stared 5mo agoA nice estimate! Since „you can compress knowledge, but not factual knowledge” https://x.com/bojie_li/status/2049314403208896521 https://x.com/bojie_li/status/2049314403208896521, it is likely we can actualy measure its size.
- stared 5mo agoI tried to run it, but estimate is 24–33T parameters, vide https://gist.github.com/stared/a86d7380937e6d0ab7920014866ace4c https://gist.github.com/stared/a86d7380937e6d0ab7920014866ac.... It seems to be a huge overshot, vide Hy3 model, which this model claims to be 2.4T, while it is 295B.
- rawoke083600 5mo agoI like your chain of thought there !
- 4ggr0 5mo agometa - i think that's the first time i've seen a table in a hn comment, and i'm surprised/impressed! nice are these pre-generated in a different tool with plain unicode and then just copy-pasted, or is it a built-in feature of hn?
- DCKing 5mo agoIf two things hold up - 1) this is actually a 2-300B parameter model and 2) this is actually competitive with frontier OpenAI and Anthropic models (and not just benchmaxing), the implications are pretty big. It would mean you could run "frontier level" performance in one box at home. 300B models at least fit in a single maxed out Mac Studio or a small stack of DGX Sparks or AMD Strix Halo boxes. For comparison, DeepSeek V4 Flash is all the rage now for small efficient models. It's very good for its size but far from the performance of the latest GPT Pro and Opus models. The vanilla variant has 284B parameters. It fits on both 256GB and 512GB Mac Studios and hits about 20-30 tokens/second. The implication of all this here is that you could have a (somewhat sluggish) Opus in a small box at home. At least once competing models and hardware to run them will be available (high end Mac Studios have been discontinued). Something tells me that this means that Google's performance numbers here are inflated.
- stymaar 5mo ago> the implications are pretty big. It would mean you could run "frontier level" performance in one box at home. That wouldn't surprise me at all actually, models like Qwen3.6-35B are comparable to frontier level models from a year ago and I wouldn't be surprised if we had self-hostable open weight models matching Opus 4.7 in a year. Assuming that Google has one year of advance against Chinese lab isn't far fetched given how much resources they have compared to their Chinese competitors.
- DCKing 5mo agoI think there was a leap around Opus 4/4.1 that hasn't quite been equalled by self hostable models yet. Perhaps full Kimi K2.6 and Deepseek V4 Pro can achieve Opus 4.1 levels (it's hard to compare anyway, benchmarks are largely a game nowadays), but both of these are also north of 1000B parameters and therefore really impractical to run at home for the foreseeable future. It's not yet obvious to me that you can achieve the breakthrough performance of say Opus 4.1/4.5 in a number of parameters you can swing at home.
- stymaar 5mo ago
- PunchTornado 5mo agoi would like to get a job like that. what can i study? I am mostly a ml engineer / researcher.
- wing-_-nuts 5mo agoThe fact that this is running on tpus is a huge point. Counting those against the other available datacenter hardware used by others, it puts google at a huge advantage, and compute > * while scaling is still working