3 ms·
What's interesting is that each run of the model tends to converge to a different "local maximum" in the solution space, and some of these local maxima correspo
by msoad 3y ago
What's interesting is that each run of the model tends to converge to a different "local maximum" in the solution space, and some of these local maxima correspond to better performance than others. By running the model multiple times, we increase the chances of finding a higher-quality local maximum or even the absolute best solution.
This got me thinking: why is this ensembling step implemented as a higher-level abstraction on top of the base LLM, rather than being built directly into the neural network architecture and training process itself?
- sp332 3y agoWell you’re right that LLM tooling is totally inadequate. At least we already have beam search. But the more boring answer (and why beam search is also uncommon) is that running the query multiple times is more expense.