3 ms·
I run an AI platform and we need to tokenize fast and early to make a lot of decisions on the subsequent steps (things like routing, rate limiting and such). I
by scottcha 3mo ago
I run an AI platform and we need to tokenize fast and early to make a lot of decisions on the subsequent steps (things like routing, rate limiting and such). Its really important to do this efficiently even though its not a large % of total end to end time for the request.
- jaggederest 3mo agoTo concur it's "latency critical", not "performance critical", people often confuse those two - optimize it all, but especially the chained critical path latency!
- dataflow 2mo agoLatency isn't performance? Maybe you mean "not throughput-critical"?
- rockwotj 2mo agobut according to Little’s law, if you improve latency, you also improve throughput, right? If I have the same number of CPU cores and they all can do their work in half the time they can double the number of requests now
- scheme271 2mo agoI don't think that's accurate. If tokenization takes say 10ms and the rest of the inference steps take 50 ms then, improving tokenization will improve the time to first token but won't affect throughput much. After the first token, the inference steps effectively hide the tokenization time.
- ludston 2mo agoIt depends on whether or not latency is the constraint of throughput in your circumstances.
- jaggederest 2mo agoJust to clarify, latency is one form of performance, and a separate thing to optimize from total resource usage in more classic "performance critical" situations. That performance might be energy, space, or other dimensions besides latency. It might also be something like reliability, accuracy, precision, or even the very human factors like simplicity, modifiability, and visibility. Heck, even latency alone you can just reduce the standard deviation and get smoother flows. Little's law is a great callout here too, one of my favorite computer science principles.
- quietfox 3mo ago‘I run an AI platform’ I have so many genuine questions I don’t even know where to start.
- NuclearPM 3mo agoPick one. I would like to hear it.
- flockonus 2mo agoSame here, as we're sitting in the middle between requests and what budget constraints are allowed given a particular token allowance there can be 10 ~ 100 milliseconds improvement in the UX (TTFT) given such massive tokenization speed up.
- Xalutiono 2mo agoWhats the prefered LLM runtime to use? vLLM? Tips & Tricks on parameters/settings? What happens at peak? Do people have to wait now? Increase of latency?
- scottcha 2mo agoWe use vllm as it generally has the best ecosystem support. Parameters are largely dependent on what type of requests you are serving (concurrency, input/output ratios, cached hit patterns). We've never had a limitation at the tokenizer step. Limitations at peak tend to manifest more on slower time in vllm doing prefill or decode though we actively try and minimize this.