4 ms·
Is that not what MoE models already do?
by samvaran 1y ago
Is that not what MoE models already do?
- oofbaroomf 1y agoNo. Each expert is not separately trained, and while they may store different concepts, they are not meant to be different experts in specific domains. However, there are certain technologies to route requests to different domain expert LLMs or even fine-tuning adapters, such as RouteLLM.
- retinaros 1y agoThat might already happen behind what they call test time compute
- oofbaroomf 1y agoMany models that use test time compute are MoEs, but test-time compute is generally meant to refer to reasoning about the prompt/problem the model is given, not about reasoning about which model to pick, and I don't think anyone has released an LLM router under that name.
- retinaros 1y agowe dont know what OAI does to find the best answer when reasoning but I am pretty sure that having variations of a same model is part of it.
- woah 1y agoWhy do you think that a hand-configured selection between "different domains" is better than the training-based approach in MoE?
- oofbaroomf 1y agoFirst off, they are basically completely different technologies, so it would be disingenuous to act like it's an apples-to-apples comparison. But a simple way to see it is that when you pick between multiple large models that have different strengths, you have a larger amount of parameters just to work with (e.g. Deepseek R1 + V3 + Qwen + LLaMA ends up being 2 trillion total parameters to pick from), whereas "picking" the experts in an MoE like has a smaller amount of total different parameters you are working with (e.g. R1 is 671 billion, Qwen is 235).
- AlexCoventry 1y agoMoE models route each token, in every transformer layer, to a set of specialized feed-forward networks (fully-connected perceptrons, basically), based on a score derived from the token's current representation.
- neom 1y agoGood visual explainer in here: https://deepgram.com/learn/mixture-of-experts-ml-model-guide https://deepgram.com/learn/mixture-of-experts-ml-model-guide