4 ms·
noob question: why would increased demand result in decreased intelligence?
by megabless123 8mo ago
noob question: why would increased demand result in decreased intelligence?
- vidarh 8mo agoIt would happen if they quietly decide to serve up more aggressively distilled / quantised / smaller models when under load.
- chrisjj 8mo agoThey advertise the Opus 4.5 model. Secretly substituting a cheaper one to save costs would be fraud.
- kingstnap 8mo agoOld school Gemini used to do this. It was super obvious because mid day the model would go from stupid to completely brain dead. I have a screenshot of Google's FAQ on my PC from 2024-09-13 that says this (I took it to post to discord): > How do I know which model Gemini is using in its responses? > We believe in using the right model for the right task. We use various models at hand for specific tasks based on what we think will provide the best experience.
- chrisjj 8mo ago> We use various models at hand for specific tasks based on what we think will provide the best experience ... for Google :)
- vidarh 8mo agoIf you use the API, you pay for a specific model, yes, but even then there are "workarounds" for them, such as someone else pointed out by reducing the amount of time they let it "think". If you use the subscriptions, the terms specifically says that beyond the caps they can limit your "model and feature usage, at our discretion".
- chrisjj 8mo agoSure. I was separating the model - which Anthropic promises not to downgrade - and the "thinking time" - which Anthropic doesn't promise not to downgrade. It seems the latter is very likely the culprit in this case.
- seunosewa 8mo agoOr just reducing the reasoning tokens.
- deleted 8mo ago[deleted]
- Wheaties466 8mo agofrom what I understand this can come from the batching of requests.
- chrisjj 8mo agoSo, a known bug?
- embedding-shape 8mo agoNo, basically, the requests are processed in batches, together, and the order they're listed in matters for the results, as the grid (tiles) that the GPU is ultimately processing, are different depending on what order they entered at. So if you want batching + determinism, you need the same batch with the same order which obviously don't work when there are N+1 clients instead of just one.
- chrisjj 8mo agoSure, but how can that lead to increased demand resulting in decreased intelligence? That is the effect we are discussing.
- embedding-shape 8mo agoSmall subtle errors that are only exposed at certain execution parts could be one. You might place things differently onto the GPU depending on how large the batch is, if you've found one way to be faster batch_size<1024, but another when batch_size>1024. As number of concurrent incoming requests goes up, you increase batch_size. Just one possibility, guess there could be a multitude of reasons, as it's really hard to reason about until you sit with the data in front of you. vLLM has had bugs with these sort of thing too, so wouldn't surprise me.
- chrisjj 8mo agoWouldn't you think that was as likely to increase as decrease intelligence, so average to nil in the benchmarks?
- awestroke 8mo agoI've seen some issues with garbage tokens (seemed to come from a completely different session, mentioned code I've never seen before, repeated lines over and over) during high load, suspect anthropic have some threading bugs or race conditions in their caching/inference code that only happen during very high load
- exitb 8mo agoAn operator at load capacity can either refuse requests, or move the knobs (quantization, thinking time) so requests process faster. Both of those things make customers unhappy, but only one is obvious.
- codeflo 8mo agoThis is intentional? I think delivering lower quality than what was advertised and benchmarked is borderline fraud, but YMMV.
- chrisjj 8mo agoThere is no level of quality advertised, as far as I can see.
- pseidemann 8mo agoWhat is "level of quality"? Doesn't this apply to any product?
- chrisjj 8mo agoIn this case, it is benchmark performance. See the root post.
- denysvitali 8mo agoIf there's no way to check, then how can you claim it's fraud? :)
- mcny 8mo agoPersonally, I'd rather get queued up on a long wait time I mean not ridiculously long but I am ok waiting five minutes to get correct it at least more correct responses. Sure, I'll take a cup of coffee while I wait (:
- lurking_swe 8mo ago