3 ms·
What you get is not what you cost. 40% overhead is quite typical, so you'd be looking at $120k/year. In the USA I'd consider that a competitive salary for an a
by LeonM 2mo ago
What you get is not what you cost.
40% overhead is quite typical, so you'd be looking at $120k/year. In the USA I'd consider that a competitive salary for an admin capable or keeping a $6M rack of specialized hardware running 24/7.
- margalabargala 2mo agoYeah but you don't need two such people, or even one, dedicated to this single rack. A company of the size that this is worthwhile for, probably has dedicated devops on staff already and can add this rack to the inventory with no additional staff.
- 999900000999 2mo agoWho is going to upgrade the models ? Who is going to fix it when the api does something weird ? Who is going to proactively make sure it’s not overheating? Chat GPT has enterprise contracts for a reason.
- deleted 2mo ago[deleted]
- stymaar 2mo agoAll of the answers to your questions above are in gp's comment already: > A company of the size that this is worthwhile for, probably has dedicated devops on staff already
- margalabargala 2mo agoThe same person managing the company's email accounts and whatnot. I didn't say it's fire-and-forget. I'm saying all that is maybe a day of work every 3 months.
- DiskoHexyl 2mo agoIt doesn’t really work like that. The companies which have the will and the budget to host their own LLMs typically require a ton of other, much smaller, models as well. They have internal security requirements, guardrails, audit, critical workflows start depending on your onprem setup, downtime is now something that’s not even allowed. There’s going to be a zoo of tooling, lots of bespoke work with internal clients who have no clue about docker, but now want their vibe-coded app to access a model they downloaded yesterday, and this model better be served and monitored 24/7 because now C-levels use it. No one is going to budget several millions to then look at an email admin who hears ‘cuda’ for the first time in their life and ask them to just support the entire thing somehow. And god forbid it’s an AMD setup. This obviously isn’t relevant for a 10-person startup and their second-hand xeon with a single H100. From my experience only the companies which are REALLY interested in privacy and data security bother with hosting their own models
- margalabargala 2mo agoThat's some of those companies not all of them. I agree that what you've described is probably worth a dedicated person. The earlier comment describing one rack running one model, is not.
- niltecedu 2mo agoMy company is not even the same ballpark, but even we already have people for it. And doesn't include the fact that you can get a colo location and just a MSP or a contractor to do it for you