4 ms·
Is there some reference or explanation on why the model it is non deterministic at temperature 0?
by pietroppeter 3y ago
Is there some reference or explanation on why the model it is non deterministic at temperature 0?
- Frummy 3y agohttps://news.ycombinator.com/item?id=37006224 https://news.ycombinator.com/item?id=37006224
- pietroppeter 3y agoThanks!
- luc4sdreyer 3y agoI'm not aware of anything concrete by OpenAI, but others have offered possible explanations. One idea is that the cause is batched inference in sparse MoE (mixture of experts) models. https://152334h.github.io/blog/non-determinism-in-gpt-4/ https://152334h.github.io/blog/non-determinism-in-gpt-4/ HN discussion: https://news.ycombinator.com/item?id=37006224 https://news.ycombinator.com/item?id=37006224
- davrosthedalek 3y agoSo in some sense the spectre attack for AI?
- awestroke 3y agoNo
- rav 3y agoOne important source of non-determinism is from using massive parallelism together with floating point arithmetic. In real math, a sum of numbers has an exact value that doesn't change if you change which order the numbers are added up in, but floating point arithmetic addition is not associative in the same way as real math, and parallelism can cause numbers to be added in a different order from execution to execution, which is one cause of non-determinism.
- imtringued 3y ago>and parallelism can cause numbers to be added in a different order from execution to execution Parallelism doesn't magically add non-determinism of this kind unless you intentionally build it to be non deterministic. Nothing prevents you from processing an array in order in parallel.
- kykeonaut 3y agoHowever, the poster mentions parallelism in conjunction with floating point arithmetic, not parallelism by itself.
- tmearnest 3y agoNo. The problem is in a reduction op of some sort (sum or whatever). Since there no guarantee of the order you receive the terms for the reduction, the nondeterminism enters from order of terms reduced. Since float math isn't associative, there will be slight differences depending on the order and these can amplify quickly over a deep net. You would have to explicitly order the terms prior to reduction but you don't always have that level of control.
- fl7305 3y ago> Nothing prevents you from processing an array in order in parallel. 100% correct if you remove processing time from the equation. In reality, Nvidia Cuda calculations run much faster if you let it schedule the order of floating points operations itself. This makes the ordering different from run to run. This in turn causes the results to be non-deterministic.