4 ms·
Why are people giving these n=1 comparisons like they mean anything? The worst offender is that pelican guy. These are non-deterministic systems and a single tr
by bigmadshoe 2mo ago
Why are people giving these n=1 comparisons like they mean anything? The worst offender is that pelican guy. These are non-deterministic systems and a single trial should not update your priors much at all.
Of course it's significant that your response had a bug and took four times longer, but if you're only going to try once, this isn't real science, it's just vibes.
- jklmnopqrstuvw 2mo agoMonths ago I start making this kind of test for my own reference. At beginning I I test each model multiple times, and results always same(pass or fail). Later I test only once for new models, I trust the results.
- kees99 2mo ago> multiple times, and results always same Not my experience at all. With smaller models, whenever I see a response that is going into wrong direction, I would just redo that step, and more often that not that brings improvement. This effect is less pronounced with SOTA, but still there.
- shunia_huang 2mo agoYes not my experience either. I've tried or sometimes be stupid to work on bugs/features and ask with almost identical prompts with same modal and harness set, and yes, they generate totally different results. Sometimes the output is unusable and even with extended guidance it will still drift away from what I was expecting. Sometimes the output is just one shot and follows almost whatever I want. I then be used to work like this, if the model and harness set does not work for one time, I just start a new session and do it again. And currently there is one of my task working like this.
- gnunez 2mo agoI don’t understand why people are calling these transformers models non-deterministic? Are you referring to the temperature parameter? I haven’t played with transformer internals in a while but my understanding is that if the temperature is fixed at a value where the top logit is always picked, then because they weights are fixed, the exact same input should produce the exact same output. Am I missing something?
- LeBit 2mo agoYou could answer your own question really, really quickly.
- gnunez 2mo agoI could, but then I would miss the chance to interact with such lovely people as yourself.
- gpm 2mo agoWell, yes and no. By non-deterministic I think people really mean "chaotic" in the chaos theory sense. Small perturbations in the input lead to wild and unpredictable changes in the output. Even with temperature parameters a fixed PRNG seed could mean an LLM was just chaotic and not technically non-deterministic. But more literally while LLMs are in theory deterministic (though perhaps not inference providers implementations if there's anything like a race condition affecting how things are rounded when added together) - we use the LLMs in harnesses that aren't. There are very likely races in the terminal outputs, dates both intentionally put in the context and accidentally leaked to the context, things like that.
- gnunez 2mo agoOk. I see. I guess people are not referring to the raw models themselves when they say non-deterministic, but are also including the harness used in conjunction with the model. Then, in that case, for the exact same input you could get a non-deterministic output. But the model itself and all the mathematical machinery around the model is still very much deterministic. I guess if we really needed to, we could construct a deterministic agent harness. But in most use cases we probably want some chaotic behavior to increase our chances of stumbling on the desired results. Thank you for the clarification
- nl 2mo ago> if the temperature is fixed at a value where the top logit is always picked, then because they weights are fixed, the exact same input should produce the exact same output. Am I missing something? Yes. Your input is part of a batch, and you don't know where in the batch it is. By default batches are not invariant and VLLM only supports invariance at all on some Huwaei Ascend hardware. See https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/batch_invariance.html https://docs.vllm.ai/projects/ascend/en/latest/user_guide/fe...