3 ms·
Is their hardware programming interface reasonable for implementing inference of frontier models: no quantization, several tera params? BTW, how many many para
by sylware 2mo ago
Is their hardware programming interface reasonable for implementing inference of frontier models: no quantization, several tera params?
BTW, how many many params open weight frontier models have? A few teras, 100s of teras?
- wren6991 2mo agoKimi-K3: 2.8T Qwen3.8-Max: 2.4T DeepSeek V4 Pro: 1.6T DeepSeek V4 Flash: 284B (all are total parameter counts, not active parameters)
- sylware 2mo agoRumors say chatgpt/claude/gemini/etc are in the 100s of teras. True?
- wren6991 2mo agoI'll ask my uncle (he works for Nintendo) and get back to you on that one
- sylware 2mo agoMy question is that wrong?
- wren6991 2mo agoSorry if the joke didn't land; I have heard a lot of different numbers for the size of US labs' models, but never seen any of them substantiated, so I think you're likely to just get more rumours in answer to this question. My personal take, with no sources: 100T sounds excessively high given they need to be able to actually serve these things on commercially available hardware. I would guess they are in the same order of magnitude as the Chinese frontier models. It's possible their edge is in RL training methods, training-time compute, and access to data (e.g. from customers' CC/Codex sessions), not in model size.
- sylware 2mo agoAllright! Well, the rumors I got are around 100T for those frontier models: reading stuff here and there on internet, on discussion forums with guys pretending running AI models, etc. This is so hard to sort the true from the false nowadays. If LLM becomes that good at coding, I'll have to run an open weight one locally.
- wmf 2mo agoNo, rumors say 5-10T.
- sylware 2mo agoYour rumors have one less 0 than my rumors. :) I have no clue anyway.
- happosai 2mo agoMaybe you could estimate frontier model sizes from AWS bedrock pricing?
- sylware 2mo agoHow???
- happosai 2mo agoBecause bedrock isn't subsided by anyone. Or at it would be financially insane go do so. Compare the of bedrock inference on known open source models to bedrock pricing of SOTA models. If the price of SOTA model is 10x qwen3 480B the the SOTA models is max 10x larger. Obviously anthropic Openai etc take a hefty cut from AWS for letting or run their models, so 10x more expensive model is probably only 5x heavy to run.
- aitchnyu 2mo agoMinimax 2.7 is 230 billion and 3.0 is 2.7 trillion, they had to jump into the bandwagon to improve performance.
- wmf 2mo agoYes, ROCm can be used to run frontier models and is being used by OpenAI, Anthropic, and Meta.
- sylware 2mo agoI would prefer direct hardware kernel interface. Like linux DMABUFs with userland hardware command ring buffers (I guess this hardware ring buffer instance would be specific to a VMID and a PASID).