4 ms·
Back of the envelope calculation (could be off, correct me if I am) If you are a large enough company that spends million+ on inference a month, it makes sense
by GodelNumbering 2mo ago
Back of the envelope calculation (could be off, correct me if I am)
If you are a large enough company that spends million+ on inference a month, it makes sense to buy a GB300 rack ($6M on top range from what I could find) which has 20.7 TB. Since the model is mixed trained (MXFP4), you would need less than 10% of the rack's memory to serve the full model. Aggregate HBM bandwidth: 576 TB/s. You can run over 6000 parallel agentic workflows (each with ~100k context on average) at ~30 tok/s.
Assuming the annual amortization+electricity at $1.5M/year and about 50% average annual utilization, you get less than 60 cents (USD) per million output token, for a frontier model with plenty of capacity to share, all your data never leaving premises and well over an order of magnitude cheaper!
As long as a company believes that the openweight models will continue to get more capable and 'AI is here to stay', this model provides the first solid footing for a decision to just buy a rack.
- 999900000999 2mo agoAnd hire 2 or 3 dev ops to keep it running ? That another 400 to 700k. It becomes your problem and not someone else’s. However, I don’t trust hosted LLMs for anything that needs to be private.
- clint 2mo agoJust let it manage itself, what could go wrong! :)
- russell_h 2mo ago> However, I don’t trust hosted LLMs for anything that needs to be private. Why not? Do you trust AWS with things that need to be private?
- Kevcmk 2mo agoMore than I trust frontier labs. AWS doesn't need to recoup 9 digits USD of capex
- senderista 2mo agoSo you can just use Bedrock?
- retinaros 2mo agoaws its their business to make ur data safe. frontier lab business is to use your data and mine to train models
- wongarsu 2mo agoWhere do I sign up to get 200k/yr to keep one rack running? Sounds like an incredibly chill job
- arjie 2mo agoApparently it’s going to take the 3 of us to do this, mate. Going to get so much reading done.
- LeonM 2mo agoWhat you get is not what you cost. 40% overhead is quite typical, so you'd be looking at $120k/year. In the USA I'd consider that a competitive salary for an admin capable or keeping a $6M rack of specialized hardware running 24/7.
- margalabargala 2mo agoYeah but you don't need two such people, or even one, dedicated to this single rack. A company of the size that this is worthwhile for, probably has dedicated devops on staff already and can add this rack to the inventory with no additional staff.
- 999900000999 2mo agoWho is going to upgrade the models ? Who is going to fix it when the api does something weird ? Who is going to proactively make sure it’s not overheating? Chat GPT has enterprise contracts for a reason.
- deleted 2mo ago[deleted]
- stymaar 2mo agoAll of the answers to your questions above are in gp's comment already: > A company of the size that this is worthwhile for, probably has dedicated devops on staff already
- GodelNumbering 2mo ago> And hire 2 or 3 dev ops to keep it running Not a devops but I'd say one full time is already too many.
- dboreham 2mo agoYes but zero is not enough and where do you get a fraction of a competent dev op from?
- layer8 2mo agoFrom the other dev-op work you’re doing.
- ayewo 2mo agoUnderstood but sharing your existing devops resources with this will soon become a bottleneck especially when any major downtime will keep several engineers (and long-running agents) blocked from any meaningful work until availability improves.
- layer8 2mo agoNot my experience, from an SMB that maintains its own hardware and services. You have a certain contingent of competent engineers who distribute their work across projects, and it generally works out fine. Or course you plan with some redundancy and fall-back plans in your systems.
- ayewo 2mo agoThe top poster mentioned LLM spend of millions/month to justify the estimated capex of $6m to self-host Kimi on own infra. Add to this number another $1.5m/yr in opex, so not sure I’d call such an enterprise wealthy enough to spend those kinds of sums on LLMs an “SMB”.
- JacobAsmuth 2mo agoNow you're thinking like management!
- lumost 2mo agoThere will be cloud/SaaS vendors who have lower cost of labor/capital due to automation and financing terms. Having these models in the open caps the inference margin.
- slicktux 2mo agoJust like that new jobs created by AI! Localized model maintainer/technician.
- toomuchtodo 2mo agoYou’ll slap some training on existing technologists/infra/sysadmin folks and perhaps have a support contract for the edge cases (hardware troubleshooting and advanced replacement). (managed an entire data center building with thousands of servers a lifetime ago with ~2-3 other people, it’s only gotten easier over the last two decades imho)
- shrubble 2mo agoAs mentioned, it is a "large enough company" already; they have full time sysadmins running things. Adding another rack beside the VMWare cluster, managing any storage/networking issues, etc. will be incremental costs; they already have a pager (probably not a pager not anymore just an app on their phone) like rotation schedule etc.
- w0m 2mo ago> I don’t trust hosted LLMs for anything that needs to be private I'd update this to 'I don’t LLMs for anything that needs to be private' What's to prevent the LLM from sliding a heavily obfuscated binary blob into the application that does nefarious things? If you aren't creating the LLM itself from scratch, I don't feel it can be trusted.
- mdp2021 2mo ago> What's to prevent the LLM from A NN per se is a file... The executable that runs it can "act"...
- JacobAsmuth 2mo agoWhy do you feel that creating the LLM from scratch is sufficient to trust it? Are you suggesting that you personally would read all 15 trillion tokens (plus every single agentic trade used in RL, along with its relative advantage in the batch) and personally guarantee that gradient descent would train a model which would not exfiltrate your corporate data? Or that perhaps you have a perfect alignment algorithm which you are unwilling to share with the broader research community (evil)?
- w0m 2mo ago> Why do you feel that creating the LLM from scratch is sufficient to trust it If you're doing the training yourself, you at least have a verifiable supply chain and an audit trail. Today, we have no idea if a black-box model handed to us is coded to recognize specific domains or patterns and back door an application in a sneakily targeted way. Black box is a black box. Many open-weight models clearly haven't been trained or created in the manner their creators claim---which raises the obvious question: if they lied about the recipe, what else did they lie about? Putting those models in a production capacity scares the living crap out of me. That said, building from scratch isn't about achieving mathematical perfection or manually auditing 15 trillion tokens---that's impossible. It's about eliminating third-party supply chain risk and having actual governance over the pipeline. Of course, that doesn't mean we can magically guarantee gradient descent won't produce weird emergent behaviors, or that we can blindly trust OpenAI not to backdoor things. But at least with the latter, you're making a calculated operational decision rather than blindly trusting an opaque black box of entirely unknown provenance.
- JimmaDaRustla 2mo agoWhy would a singe system require 3 full-time dev ops?
- reckless 2mo agoI think the licensing that would likely apply to a company that's able to afford ~$6M rack and the associated infrastructure muddies this somewhat
- vidarh 2mo agoAs far as I can tell, license fees are only applicable if you have more than 20m USD/year revenue from services provided using the model, or serve more than 100m users.
- petu 2mo agoI think internal use is allowed at any scale in the license? > 4. The requirements set forth in Sections 2 and 3 do not apply to: (a) internal use of the Software, defined as any use that does not make the Software, its outputs, or its underlying capabilities available to third parties; [...]
- jsnell 2mo agoIt's very hard to make sure none of the outputs are ever made available to third parties. Source code can end up widely distributed (e.g. client-side js, open source). Prose will frequently get shared across organization boundaries (e.g. emails, websites, documents).
- stymaar 2mo agoLicense for LLMs already have very little strength (it's not clear at all that they have any legal basis whatsoever) an AI lab is very, very unlikely to sue you for violating this kind of license clause for the reason you mentioned. It's clearly worded in a way to deter people from running a third party inference business out of it, they would have been much more explicit if they wanted to deter people from running it for internal use.
- jdsully 2mo agoIts an open question if model weights are even copyrightable. So the license may not mean anything.
- kcb 2mo agoYou also need a place to put it. With liquid cooling and extremely dense power capability. Your typical colo or closet server room isn't going to cut it.
- FuriouslyAdrift 2mo agoWe're running Kimi 2.8 on a $107k server and getting around 50k tokens per second or better on most things. We've already saved money compared to last years token cost on Claude/Gemini
- mdp2021 2mo ago> We're running Kimi 2.8 on a $107k server Equipped with what? Is it CPU based inference, a mix...?
- wongarsu 2mo agoNot GP, but my educated guess is that they are running a system with between 4 and 6 MI325 or MI355x or similar AMD GPUs. With the 50k tps as the total figure for all parallel requests. Those cards have a lot of memory for their price, allowing you to push to really high batch sizes while still having a large context size for each request
- FuriouslyAdrift 2mo ago4x MI300A in a Gigabyte server
- logicallee 2mo agoAre you developing software? Is most of it used on a coding agent? (Like Claude Code or ChatGPT Codex?) If so, what coding agent do you use? If you're not developing software what do you use it for (roughly)?
- FuriouslyAdrift 2mo agoOur use cases are data analysis, software development, general AI chat with RAG so far.
- freediddy 2mo agohow many requests per second can the server take?
- baron3dl 2mo agoI was wondering about this myself, specifically where does it make sense, so I spent a few hours working through it with Claude and wrapped it up into a writeup and interactive model. https://3dl.dev/kimi-k3.html https://3dl.dev/kimi-k3.html