3 ms·
from what I understand this can come from the batching of requests.
by Wheaties466 8mo ago
from what I understand this can come from the batching of requests.
- chrisjj 8mo agoSo, a known bug?
- embedding-shape 8mo agoNo, basically, the requests are processed in batches, together, and the order they're listed in matters for the results, as the grid (tiles) that the GPU is ultimately processing, are different depending on what order they entered at. So if you want batching + determinism, you need the same batch with the same order which obviously don't work when there are N+1 clients instead of just one.
- chrisjj 8mo agoSure, but how can that lead to increased demand resulting in decreased intelligence? That is the effect we are discussing.
- embedding-shape 8mo agoSmall subtle errors that are only exposed at certain execution parts could be one. You might place things differently onto the GPU depending on how large the batch is, if you've found one way to be faster batch_size<1024, but another when batch_size>1024. As number of concurrent incoming requests goes up, you increase batch_size. Just one possibility, guess there could be a multitude of reasons, as it's really hard to reason about until you sit with the data in front of you. vLLM has had bugs with these sort of thing too, so wouldn't surprise me.
- chrisjj 8mo agoWouldn't you think that was as likely to increase as decrease intelligence, so average to nil in the benchmarks?
- embedding-shape 8mo agoNo, I'm not sure how that'd make sense. Either you're making the correct (expected) calculations, or you're getting it wrong. Depending the type of wrong or how wrong, could go from "used #2 in attention instead of #1" so "blue" instead of "Blue" or whatever, to completely incoherent text and garbled output.
- chrisjj 8mo agoI accept errors are more likely to decrease "intelligence". But I don't see how increased load, through batching, is any more likely to increase than decrease errors.