3 ms·
Out of curiosity, which models are fully deterministic? I was under the impression that all LLMs were fundamentally probabilistic.
by antx 19d ago
Out of curiosity, which models are fully deterministic? I was under the impression that all LLMs were fundamentally probabilistic.
- Wowfunhappy 19d agoThe randomness is something we add on purpose; you can set an LLM's "temperature" to 0 to get deterministic output. This tends to make the quality of its responses worse for reasons I don't think anyone really understands, but it's still functional. I don't think the state of the art LLM providers let you do this anymore (?), but they certainly could if they wanted to, and you can do it yourself with a local model.
- antx 18d agoSetting the temperature to 0 mathematically tells the system to always choose the absolute highest-probability word (known as "greedy decoding"), but in no reality is this "deterministic". Output drift is still a thing.
- Wowfunhappy 18d agoBut that's from, like, floating point errors, right? If you used higher precision that wouldn't happen, it's just because we're cheap in how we do rounding. I can see how you'd nitpick this, but to me this is a deterministic algorithm that just happens to be running on nondeterministic hardware.
- pixl97 18d agoI mean, currently we'd have some difficultly proving our hardware isn't deterministic, just that we can't actually test it. But, I think you're tricking yourself on determinism. You'll say something like "I know if I ask an LLM what 1+1 is, it will answer 2", but the thing is, you don't. You have to run the LLM first to figure out it's output. And when you send in just a few bits of text, it's outputs are going to be rather limited. But this all breaks when it hits the real world. Inputs are unpredictable. Hence while LLM outputs, like humans, are probabilistic, you can't figure out what it's going to be until you ask. And in any high complexity data gathering environment you're going have a difficult time ensuring your entire systems conditions are the same. System consistency is very hard, once you start running thousands of processors in an agentic loop small errors accrue and timing starts differing and the system will take non-deterministic paths.
- throwaway63486 18d agoI think you're arguing a different thing than determinism. If I ask an llm to "add 2 and 2" is and it replies corectly, then I ask for "the sum of 2 and 2" and it replies "banana" that is a lack of predictability and consistency but not a lack of determinism. As long as it produces the same output for a given input, unhinged or not, it is deterministic. Your example at the end of different systems feeding data to each other is non-deterministic only at the system level, not the individual llm level.
- pixl97 18d ago>only at the system level, not the individual llm level. Which is why llms aren't agents and depend on harnesses. The llm itself doesn't have a continual loop built in, that would be very power hungry. The harness works as the orchestrator of memory and action. Now, I can't think of a reason why an LLM couldn't bootstrap its own harness, but in general it sounds like a very dumb idea to actually build that from an AI safety perspective. This discussion falls under the idea and refutation of the Chinese Room. The room may have no idea what Chinese characters are, but the system does.
- mitxela 18d agoYou can also use a seeded random generator to get the same random numbers each time
- piker 19d agoSame weights, same seed, same input tokens, same algorithm, same output tokens, probabilistic or not. Quantum effects have been de-noised, but I guess there are still random gamma rays.
- qarl 19d agoNaw - computers are really deterministic. It's hard to get them to behave otherwise. As I understand it, if you turn down the temperature to 0 you get repeatable behavior - EXCEPT - on large servers with lots of users - the GPU can sometimes produce slightly different results based on batch size.
- joe_the_user 19d agoUnless you have something exotic, the randomness that's adding to a computer is a combination of how it's configured combined with a pseudo-random number generator. I assume the system adds entropy to the generator regularly but all you need to do is fix the various supposedly random inputs and you can get full determinism even without zero temperature.
- qarl 19d agoYes, in theory, of course. In practice - on a multitasking OS with input from multiple human users - it's hard to get it deterministic because of that GPU scheduling thing I mentioned.
- fc417fc802 18d agoGPU scheduling only affects the result due to buggy optimizations. It's the exact same mechanism as fp rounding error on the CPU or updates to globally shared PRNG state. We use lots of buggy optimizations because they don't matter in practice in most situations (see ex -ffast-math).
- qarl 18d agoI don't know the details but apparently it has something to do with multiprocess contention for the GPU and batch sizing.
- fc417fc802 18d agoRight but regardless, it's a buggy optimization. The calculations are all fully deterministic when done "properly" and fully consistently but we almost never bother with that because it slows things down and the errors don't matter in practice 99% of the time.
- SkyBelow 18d agoBy default they are matrix multiplications. Temperature is added in as forced PRNG because testing found that correlated with better outputs. Given the same prompts and the same weights, one can get the same answer each time. In practice, there are a number of optimizations that makes the results dependent upon thing we give up control of to increase performance, meaning the results end up being effectively non-deterministic. But, if you are willing to run it in a slower mode so we don't do some steps out of order to speed things up and don't batch results (or if you consider the determinism of a given batch of requests rather than individual requests), then the same input gets the same output.