5 ms·
Hetzner is working on LLM Inference
- embedding-shape 2mo ago> The enable_thinking option is worth mentioning. Without it, the model can spend a surprising amount of the completion budget reasoning before it returns a visible answer. Straight up the opposite, which the name makes abundantly clear, with the option it does reasoning, without it it doesn't...
- jonas_scholz 2mo agoMy writing wasnt clear here I think, without the option it defaults to reasoning enabled. With "without it" I meant without enable_thinking=False!
- cyanydeez 2mo agollamacpp has a reasoning-budget and reasoning-message setting that can both be a global or per header setting. Using it, you can stop it's reasoning token count and insert a message at the point you stopped it. This allows both the client and server to customize it. I typically use a message that tells it to compact the conversation and use subagents. I find the reasoning gets bloated when it's failed to do whatever task it's doing and often times it either has too little context (subagents) or its context is bloated (compact). This works fairly well to get it to extend workable life up to ~1M on a local model.
- ano-ther 2mo agoGood to see more developments in this space. I quite like this service, which is a little further than Hetzner and has several models to choose from: https://www.infomaniak.com/en/hosting/ai-services https://www.infomaniak.com/en/hosting/ai-services
- jonas_scholz 2mo agoInteresting, didn’t know about them! Weird model selection though, no glm or deepseek?
- archerx 2mo agoInfomaniak is such a shitty company, I had to use them on a previous job I worked at and dealing with them was awful.
- MASNeo 2mo agoThat makes it sound like dealing with Hetzner is easy. Is the case?
- sam_lowry_ 2mo agoDepends on your mindset. If you are an Engineer that does not need hand-holding and you accept terse answers from support with the gratitude to the human on the other side, then go for it. You will at least not need months of training and tough exams to be an expert in Hetzner Cloud. It is simple, but that's the point.
- jonas_scholz 2mo agoI mostly use Hetzner baremetal servers and not cloud, but the support there is in my experience very competent, but also VERY german and direct. I can imagine if youre neither german or super technical that this can be intimidating or considered unfriendly
- drcongo 2mo agoBeing slightly on the spectrum, I find Hetzner's "VERY german and direct" support perfect.
- Saris 2mo agoI only have a couple of VPS plans with them, but my experience is good.
- sureglymop 2mo agoI highly agree. I once registered a domain through them that happened to be my last name but coincidentally also the name of a company. They restricted me from using the domain and never refunded me the yearly fee for it (citing trademark reasons, although under Swiss trademark law, someones legal name is not a trademark violation, especially as it was just my personal website and I wasn't working on a competing product/service). Only weeks later someone else successfully registered it without such issues. Worse even, in my account panel I still had control over the e-mail service of the domain, so there was most definitely a pretty bad security critical bug there.
- drcongo 2mo agoImportant note on Infomaniak's offering - when they first launched it, I gave it a try but despite claims that it's OpenAI API compatible, even basic things like sending a base64 encoded image were broken. I reached out to their support, who replied with (verbatim): "Unfortunately, we do not provide support for our AI Service, as the solution is highly unmanaged and uses our API." "highly unmanaged" didn't fill me with confidence, and "uses our API" is a very weird reason to give for not offering support. I wrote off ever using them for anything beyond email after that reply.
- Aldipower 2mo agolurus.ai and cortecs.ai are worth a try too!
- swiftcoder 2mo agoIt would certainly be interesting to have a highly respected EU-native inference provider, if only to make the regulatory gods happy
- jonas_scholz 2mo agoagree. I have a usecase for a bigger open-weight model hosted by an EU company and the selection isnt really great, hard to make the regulatory gods (and the devs) happy at the same time right now
- sofixa 2mo agoScaleway have two separate (one fully managed one a bit less) services for that: https://www.scaleway.com/en/generative-apis/ https://www.scaleway.com/en/generative-apis/ https://www.scaleway.com/en/inference/ https://www.scaleway.com/en/inference/
- jonas_scholz 2mo agoyea, but no prompt caching right? This makes it unusable for my usecase at least, the cost would be insane
- PaoloBarbolini 2mo agoThey confirmed on LinkedIn that they are working on it. Also, they have a feature request that has been getting many votes recently: https://feature-request.scaleway.com/posts/1251/prompt-caching-to-reduce-input-tokens-cost https://feature-request.scaleway.com/posts/1251/prompt-cachi...
- deleted 2mo ago[deleted]
- hoppp 2mo agoScaleway does that already but competition is always good.
- mark_l_watson 2mo agoThis seems like a smart move, given their ability to host efficiently. I approve of efforts to make the cost of inference for smaller useful models slowly approach 'close to zero' and there are many good paths for getting there. It is useful for companies to get fast hosting for the class of smaller models they may end up hosting in house.
- jonas_scholz 2mo agoI really hope they dont stop at the small models though! The bigger ones that dont fit on a single GPU are more interesting I think
- mips_avatar 2mo agoProblem is right now the biggest GPU boxes they have is single rtx pro 6000s.
- scoriiu 2mo ago[flagged]
- nubg 2mo agoPotentially interesting article ruined by AI slop hallucinations like > For now, the API is fast, free, and fun to try. The next hardware announcement will tell us much more than another small model would.
- jonas_scholz 2mo agowhat is the hallucination here? it is fast, free and fun to try. And I genuinely think that the hardware decision (if they get bigger gpus) decides if this will be a banger product or not?
- nubg 2mo agonobody writes like that. the article is based on a prompt, which i would much rather read, instead of the ironed out version that the llm produced
- jonas_scholz 2mo agofine, i agree that this doesnt sound like me. But the content is correct and not a hallucination!
- petesergeant 2mo ago> the article is based on a prompt, which i would much rather read I would much rather read someone's carefully written and LLM-assissted article than shallow and lazy dismissals like this one. What, precisely, do you feel you've added to the conversation here?
- rebelde 2mo agoHetzner is very efficient hosting servers Will this be the new division of labor? Americans - best proprietary models Chinese - best open weight models Europeans - best / most efficient inference service
- andsoitis 2mo agoWe need frontier open weight models from the West. Maybe Nvidia's Nemotron can get there. https://research.nvidia.com/labs/nemotron/Nemotron-3-Ultra/ https://research.nvidia.com/labs/nemotron/Nemotron-3-Ultra/
- ForHackernews 2mo agowhy?
- andsoitis 2mo agoReasons include, but are not limited to: a) broad competition is good b) jurisdiction diversification (you don't want to be dependent on the regulatory winds of a single jurisdiction)
- ForHackernews 2mo ago...but you're not dependent on anyone? If I'm running a Kimi model on my local cluster, the CCP can't shut it off. It's not like the Finnish government has veto power over Linux.
- andsoitis 2mo agoYou're not necessarily wrong for now, but that framing might be too narrow. It is not beyond imagination that a government can limit availability to future frontier versions of the open weight models produced within their jurisdiction. It pays to be a little paranoid and have a diversified portfolio of options to choose from. Anything that can have major geopolitical consequence is especially susceptible to such structural risk. This has nothing to do with any particular government/country/jurisdiction except insofar as said country is a leading power and hence has immense leverage across a very wide range of dimensions. Said more plainly, it is safe to assume that a country with immense power is unlikely to be willing to cede or distribute such power to others. If you're hung up on my choice of "the West", consider that I mean it geopolitically and economically, so broadly includes North America, Europe, Australia, New Zealand, Japan, and South Korea. I could have also said the Global South, but I think it is a fair statement that none of those countries have the means, but I'd be totally cool and happy if I'm proven wrong!
- Havoc 2mo agoInteresting. I could see them perhaps coming in competitive for models that fit into single cards? Less so playing in the big model serving league...climbing into that esp right now would be madness
- jonas_scholz 2mo agoWhy do you think this would be madness? It seems like, at least in the EU, there is barely competition for the big open-weight models?
- perelin 2mo agoWhats definitely missing: a solid (non Mistral) GDPR compliant coding plan / subscription. All offerings are either US or China based. With the newest open weights models this became really interesting imo.
- dk970 2mo ago[flagged]
- _pdp_ 2mo agoI mean yah... host glm and kimi and I am game.
- NetOpWibby 2mo agoThis is interesting because I thought Hetzner was anti-crypto? LLMs aren't the same but they're often lumped in with crypto as "things no one wants."
- satvikpendem 2mo agoWho lumps LLMs with crypto? The former has actual usage and purpose, the latter doesn't. Anyone who does must think any new technology is all the same.
- NetOpWibby 2mo agoA lot of people on Mastodon, for one.
- cousinbryce 2mo agoCrypto DAU: 10’s of k LLMaaS DAU: 100’s of k
- danlitt 2mo agoExcited to see the price of every other product they offer triple for no reason.
- jonas_scholz 2mo ago"no reason" like hardware prices going through the roof?
- danlitt 2mo agoHasn't that already happened?
- jonas_scholz 2mo agoyeah, but why would LLM inference now triple the price of the other products? Or am I misunderstanding what you're saying?
- danlitt 2mo agoI am implying that they will subsidise the LLM inference product by increasing the prices of other products.
- jonas_scholz 2mo agoThey wont, they are very public about never subsidising any products. I doubt they change that after 29 years now.
- toomuchtodo 2mo agogestures broadly at ~$1T+ in AI capex spend
- danlitt 2mo agoNot sure what you're getting at here. Hetzner has not spent 1 trillion dollars!
- pmg1991 2mo agoVery soon we will be having reseller programs for inference, this will be just like web hosting reseller. After big players, small players will also start entering in this field. I'm waiting for that day so that inference will be affordable just like web hosting. 200$ per month is in no way affordable by everyone.
- luciana1u 2mo ago[flagged]