4 ms·
Hi everyone. I'm one of the maintainers of this project. We're both excited and humbled to see it on Hacker News! We created this handbook to make LLM inferenc
by sherlockxu 1y ago
Hi everyone. I'm one of the maintainers of this project. We're both excited and humbled to see it on Hacker News!
We created this handbook to make LLM inference concepts more accessible, especially for developers building real-world LLM applications. The goal is to pull together scattered knowledge into something clear, practical, and easy to build on.
We’re continuing to improve it, so feedback is very welcome!
GitHub repo: https://github.com/bentoml/llm-inference-in-production https://github.com/bentoml/llm-inference-in-production
- armcat 1y agoAmazing work on this, beautifully put together and very useful!
- DiabloD3 1y agoI'm not going to open an issue on this, but you should consider expanding on the self-hosting part of the handbook and explicitly recommend llama.cpp for local self-hosted inference.
- leopoldj 1y agoThe self hosting section covers corporate use case using vLlm and sglang as well as personal desktop use using Ollama which is a wrapper over llama.cpp.
- DiabloD3 1y agoRecommending Ollama isn't useful for end users, its just a trap in a nice looking wrapper.
- nl 1y agoStrong disagree on this. Ollama is great for moderately technical users who aren't really programmers or proficient with the command line.
- DiabloD3 1y agoYou can disagree all you want, but Ollama does not keep their llama.cpp vendored copy up to date, and also ships, via their mirror, completely random badly labeled models claiming to be the upstream real ones, often misappropriated from major community members (Unsloth, et al). When you get a model offered by Ollama's service, you have no clue what you're getting, and normal people who have no experience aren't even aware of this. Ollama is an unrestricted footgun because of this.
- nl 1y agoI thought the models were like HuggingFace, where anyone can upload a model and you choose which you pull. The Unsloth ones look like this to me, eg: https://ollama.com/secfa/DeepSeek-R1-UD-IQ1_S https://ollama.com/secfa/DeepSeek-R1-UD-IQ1_S
- DiabloD3 1y agoOllama themselves upload models to the mirror, and often mislabel them. When R1 first came out, for example, their official copy of it was one of the distills labeled as "R1" instead of something like "R1-qwen-distill". They've done this more than once.
- ChromaticPanic 1y agoNot the footgun you think it is. Ollama comes with a few things that make it convenient for casual users.
- criemen 1y agoThanks a lot for putting this together! I have a question. In https://github.com/bentoml/llm-inference-in-production/blob/main/docs/llm-inference-basics/img/llm-inference-flow.png https://github.com/bentoml/llm-inference-in-production/blob/..., you have a single picture that defines TTFT and ITL. That does not match my understanding (but you guys know probably more than me): In the graphic, it looks like that the model is generating 4 tokens T0 to T3, before outputting a single output token. I'd have expected that picture for ITL (except that then the labeling of the last box is off), but for TTFT, I'd have expected that there's only a single token T0 from the decode step, that then immediately is handed to detokenization and arrives as first output token (if we assume a streaming setup, otherwise measuring TTFT makes little sense).
- sherlockxu 1y agoThanks. We have updated the image to make it more accurate.
- sethherr 1y agoThis seems useful and well put together, but splitting it into many small pages instead of a single page that can be scrolled through is frustrating - particularly on mobile where the table of contents isn't shown by default. I stopped reading after a few pages because it annoyed me. At the very least, the sections should be a single page each.