5 ms·
> After trial and error with different models As a mere occasional customer I've been scanning 4 to 5 pages of the same document layout every week in gemini fo
by makeitdouble 2y ago
> After trial and error with different models
As a mere occasional customer I've been scanning 4 to 5 pages of the same document layout every week in gemini for half a year, and every single week the results were slightly different.
To note the docs are bilingual so it could affect the results, but what stroke me is the lack of consistency, and even with the same model, running it two or three times in a row gives different results.
That's fine for my usage, but that sounds like a nightmare if everytime Google tweaks their model, companies have to reajust their whole process to deal with the discrepancies.
And sticking with the same model for multiple years also sound like a captive situation where you'd have to pay premium for Google to keep it available for your use.
- tomrod 2y agoConsider turning down the temperature in the configuration? LLMs have a bit of randomness in them. Gemini 2.0 Flash seems better than 1.5 - https://deepmind.google/technologies/gemini/flash/ https://deepmind.google/technologies/gemini/flash/
- iandanforth 2y agoAt temperature zero, if you're using the same API/model, this really should not be the case. None of the big players update their APIs without some name / version change.
- pigscantfly 2y agoThis isn't really true unfortunately -- mixture of experts routing seems to suffer from batch non-determinism. No one has stated publicly exactly why this is, but you can easily replicate the behavior yourself or find bug reports / discussion with a bit of searching. The outcome and observed behavior of the major closed-weight LLM APIs is that a temperature of zero no longer corresponds to deterministic greedy sampling.
- brookst 2y agoIf temperature is zero, and weights are weights, where is the non-deterministic behavior coming from?
- petesergeant 2y agoThe parent is suggesting that temperature only applies at the generation step, but the choice of backend “expert model” that a request is given to (and then performs the generation) is non-deterministic. Rather than being a single set of weights, there are a few different sets of weights that constitute the “expert” in MoE. I have no idea if that’s true, but that’s the assertion
- brookst 2y agoI don't think it makes sense? Somewhere there has to be a RNG for that to be true. MOE itself doesn't introduce randomness, and the routing to experts is part of the model weights, not (I think) a separate model.
- pigscantfly 2y agoThe samples your input is batched with on the provider's backend vary between calls and sparse mixture of experts routing when implemented for efficient utilization induces competition among tokens with either encouraged or enforced balance of expert usage among tokens in the same fixed-size group. I think it's unknown or at least undisclosed exactly why sequence non-determinism at zero temperature occurs in these proprietary implementations, but I think this is a good theory. [1] https://arxiv.org/abs/2308.00951 https://arxiv.org/abs/2308.00951 pg. 4 [2] https://152334h.github.io/blog/non-determinism-in-gpt-4/ https://152334h.github.io/blog/non-determinism-in-gpt-4/
- kettleballroll 2y agoI thought the temperature only affects randomness at the end of the network (when turning embeddings back I to words using the softmax). It cannot influence routing, which is inherently influenced by which examples get batched together (ie, it might depend on other users of the system)
- kiratp 2y agoQuantized floating point math can, under certain scenarios, be non-associative. When you combine that fact with being part of a diverse batch of requests over an MoE model, outputs are non-deterministic.
- mejutoco 2y ago> and every single week the results were slightly different. This is one of the reasons why open source offline models will always be part of the solution, if not the whole solution.
- rafaelmn 2y agoInconsistency comes from scaling - if you are optimizing your infra to be cos effective you will arrive at same tradeoffs. Not saying it's not nice to be able to make some of those decisions on your own - but if you're picking LLMs for simplicity - we are years away from running your own being in the same league for most people.
- mejutoco 2y agoAnd if you are not you wont. You can decide if you change your local setup or not. You cannot decide the same of a service. There is nothing inevitable about inconsistency in a local setup.
- bushbaba 2y agoThat’s why you have azure openAI APIs which give a lot more consistency