6 ms·
How confident are you in the opus 4.6 model size? I've always assumed it was a beefier model with more active params that Qwen397B (17B active on the forward pa
by ymaws 7mo ago
How confident are you in the opus 4.6 model size? I've always assumed it was a beefier model with more active params that Qwen397B (17B active on the forward pass)
- codemog 7mo agoAlso curious if any experts can weigh in on this. I would guess in the 1 trillion to 2 trillion range.
- Chamix 7mo agoTry 10s of trillions. These days everyone is running 4-bit at inference (the flagship feature of Blackwell+), with the big flagship models running on recently installed Nvidia 72gpu rubin clusters (and equivalent-ish world size for those rented Ironwood TPUs Anthropic also uses). Let's see, Vera Rubin racks come standard with 20 TB (Blackwell NVL72 with 10 TB) of unified memory, and NVFP4 fits 2 parameters per btye... Of course, intense sparsification via MoE (and other techniques ;) ) lets total model size largely decouple from inference speed and cost (within the limit of world size via NVlink/TPU torrus caps) So the real mystery, as always, is the actual parameter count of the activated head(s). You can do various speed benchmarks and TPS tracking across likely hardware fleets, and while an exact number is hard to compute, let me tell you, it is not 17B or anywhere in that particular OOM :) Comparing Opus 4.6 or GPT 5.4 thinking or Gemini 3.1 pro to any sort Chinese model (on cost) is just totally disingenuous when China does NOT have Vera Rubin NVL72 GPUs or Ironwood V7 TPUs in any meaningful capacity, and is forced to target 8gpu Blackwell systems (and worse!) for deployment.
- aurareturn 7mo agoChina is targeting H20 because that's all they were officially allowed to buy.
- Chamix 7mo agoI generally agree, back of the napkin math shows H20 cluster of 8gpu * 96gb = 768gb = 768B parameters on FP8 (no NVFP4 on Hopper), which lines up pretty nicely with the sizes of recent open source Chinese models. However, I'd say its relatively well assumed in realpolitik land that Chinese labs managed to acquire plenty of H100/200 clusters and even meaningful numbers of B200 systems semi-illicitly before the regulations and anti-smuggling measures really started to crack down. This does somewhat beg the question of how nicely the closed source variants, of undisclosed parameter counts, fit within the 1.1tb of H200 or 1.5tb of B200 systems.
- aurareturn 7mo agoThey do not have enough H200 or Blackwell systems to server 1.6 billion people and the world so I doubt it's in any meaningful number.
- Chamix 7mo agoI assure you, the number of people paying to use Qwen3-Max or other similar proprietary endpoints is far less than 1.6 billion.
- aurareturn 7mo agoYou don't need to assure me. It's a theoretical maximum.
- jychang 7mo agoNobody is running 10s of trillion param models in 2026. That's ridiculous. Opus is 2T-3T in size at most.
- johndough 7mo agoDo you have any clues to guess the total model size? I do not see any limitations to making models ridiculously large (besides training), and the Scaling Law paper showed that more parameters = more better, so it would be a safe bet for companies that have more money than innovative spirit.
- magicalhippo 7mo ago> I do not see any limitations to making models ridiculously large (besides training) From my understanding, the "besides training" is a big issue. As I noted earlier[1], Qwen3 was much better than Qwen2.5, but the main difference was just more and better training data. The Qwen3.5-397B-A17B beat their 1T-parameter Qwen3-Max-Base, again a large change was more and better training data. [1]: https://news.ycombinator.com/item?id=47089780 https://news.ycombinator.com/item?id=47089780
- Chamix 7mo agoWhat do you think labs are doing with the minimum 10TB memory in NvLink 72 systems that were publicly reported to all start coming online in November/December of last year? And why would this 1 TB -> 10 TB jump matter so much for Anthropic previously being wholly dependent on running Opus 4x on TPUs, if the models were 2-3T at 4bit and could fit in 8x B200 (1.5 TB = 3T param) widely deployed during the Opus 4 era? You have presented a vibe-based rebuttal with no evidence or or logic to outline why you think labs are still stuck in the single trillions of parameters (GPT 4 was ~1 trillion params!). Though, you have successfully cunninghammed me into saying that while anything I publicly state is derived from public info, working in the industry itself is a helpful guide to point at the right public info to reference.
- johndough 7mo agoCould you point at some more public info about active parameter count? You said: > and while an exact number is hard to compute, let me tell you, it is not 17B or anywhere in that particular OOM :) I can see ~100B, but that would near the same order of magnitude. I find ~1000B active parameters hard to believe.
- daemonologist 7mo agoEven if it's larger, OpenRouter has DeepSeek v3.2 (685B/37B active) at $0.26/0.40 and Kimi K2.5 (1T/32B active) at $0.45/2.25 (mentioned in the post).
- johndough 7mo agoOpus 4.6 likely has in the order of 100B active parameters. OpenRouter lists the following throughput for Google Vertex: 42 tps for Claude Opus 4.6 https://openrouter.ai/anthropic/claude-opus-4.6 143 tps for GLM 4.7 (32B active parameters) https://openrouter.ai/z-ai/glm-4.7 70 tps for Llama 3.3 70B (dense model) https://openrouter.ai/meta-llama/llama-3.3-70b-instruct For GLM 4.7, that makes 143 * 32B = 4576B parameters per second, and for Llama 3.3, we get 70 * 70B = 4900B, which makes sense since denser models are easier to optimize. As a lower bound, we get 4576B / 42 ≈ 109B active parameters for Opus 4.6. (This makes the assumption that all three models use the same number of bits per parameter and run on the same hardware.)
- jychang 7mo agoYep, you can also get similar analysis from Amazon Bedrock, which serves Opus as well. I'd say Opus is roughly 2x to 3x the price of the top Chinese models to serve, in reality.
- Bolwin 7mo agoYeah that's a massive assumption they're making. I remember musk revealed Grok was multiple trillion parameters. I find it likely Opus is larger. I'm sure Anthropic is making money off the API but I highly doubt it's 90% profit margins.
- aurareturn 7mo agoAnthropic CEO said 50%+ margins in an interview. I'm guessing 50 - 60% right now.
- deleted 7mo ago[deleted]
- jychang 7mo ago> I find it likely Opus is larger. Unlikely. Amazon Bedrock serves Opus at 120tokens/sec. If you want to estimate "the actual price to serve Opus", a good rough estimate is to find the price max(Deepseek, Qwen, Kimi, GLM) and multiply it by 2-3. That would be a pretty close guess to actual inference cost for Opus. It's impossible for Opus to be something like 10x the active params as the chinese models. My guess is something around 50-100b active params, 800-1600b total params. I can be off by a factor of ~2, but I know I am not off by a factor of 10.
- simianwords 7mo agoAre you sure you can use tps as a proxy?
- jychang 7mo agoIn practice, tps is a reflection of vram memory bandwidth during inference. So the tps tells you a lot about the hardware you're running on. Comparing tps ratios- by saying a model is roughly 2x faster or slower than another model- can tell you a lot about the active param count. I won't say it'll tell you everything; I have no clue what optimizations Opus may have, which can range from native FP4 experts to spec decoding with MTP to whatever. But considering chinese models like Deepseek and GLM have MTP layers (no clue if Qwen 3.5 has MTP, I haven't checked since its release), and Kimi is native int4, I'm pretty confident that there is not a 10x difference between Opus and the chinese models. I would say there's roughly a 2x-3x difference between Opus 4.5/4.6 and the chinese models at most.