3 ms·
Would changing the seed affect generation much? Even though beam search depends on the seed, the llms woul still be generating good probability distributions on
by rdedev 3y ago
Would changing the seed affect generation much? Even though beam search depends on the seed, the llms woul still be generating good probability distributions on the next word to select. Maybe a few words would change but don't think the overall meaning would
- hansvm 3y agoOverall meaning can vary profoundly. As a toy example, consider the prompt "randomly generate the first word that comes to mind." The output is deterministic in the seed, so to get new results you need new seeds, but with new seeds you open up the 2k most common words in a language in a uniform-esque distribution. Building on that, instead of <imagining> words, suppose you <imagine> attack vectors. Many, many attacks exist and are known. Presumably, many more exist and are unknown. The distribution the LLM will produce in practice is extremely varied, and some of those variations probably won't work. If we're not just talking about a single prompt but rather a sequence of prompts with feedback, you're right that the seed matters less (when its errors are presented, it can self-correct a bit), but there are other factors at play. (1) You're resetting somehow eventually anyway. Details vary, but your context window isn't unlimited, and LLM perf drops with wider windows, even when you can afford the compute. You might be able to retain some state, but at some point you need something that says "this shit didn't work, what's next". A new seed definitely gives new ideas, whereas clever ways to summarize old information might yield fixed points and other undesirable behavior. (2) Seed selection, interestingly, matters a ton for model performance in other contexts. This is perhaps surprising when we tend to use random number generators which pass a battery of tests to prove they're halfway decent, but that's the reason you want to see (in reproducible papers) a fixed seed of 0 or 42 or something, and the authors maintaining that seed across all their papers (to help combat the fact that they might be cherry-picking across the many choices of "nice-looking" random seeds when they publish a result to embelish the impact). The gains can be huge. I haven't seen it demonstrated for LLMs, but most of the architecture shouldn't be special in that regard. And so on. If nothing else, picking a new seed is a dead-simple engineering decision to eliminate a ton of things which might go wrong.
- rdedev 3y agoI agree with you except for point 2. A well performing model should show such drastic changes wrt the seed value. Besides the huge amount of training data as well as test data should mitigate differences in data splitting. There would be difference but my hunch is it would be negligible. Of course as you said no one has tested this out so we can't say how the performance would change either way