3 ms·
It is widely admitted by practitioners that all frontier models with very few exceptions (e.g . Mistral-large, llama 3) are MoE. This includes gpt4-turbo you me
by sinenomine 2y ago
It is widely admitted by practitioners that all frontier models with very few exceptions (e.g
. Mistral-large, llama 3) are MoE. This includes gpt4-turbo you mentioned.
- Digitmaster_AI 2y agoResponse to MoE Assumption About GPT-4 Turbo: It’s true that many frontier models are using Mixture of Experts (MoE) architectures, but there’s a critical distinction between how different models implement MoE—assuming they do at all. DeepSeek: Confirmed MoE model, openly admits to dynamically selecting experts during inference. This results in inconsistent answers, logical loops, and a lack of reasoning stability—demonstrated in my testing. GPT-4 Turbo (ChatGPT): OpenAI has not confirmed that it’s MoE. While some suspect it may have MoE-like elements, it behaves far more like a dense, unified model in practice—meaning responses are significantly more stable across different conversations. Dense Models (LLaMA, Mistral): No MoE at all. These models process everything with the entire network active at all times, resulting in predictable and repeatable outputs. Why This Matters: MoE models vary drastically in how they select and integrate experts. DeepSeek struggles because its expert selection appears inconsistent, leading to wildly different answers in fresh sessions. In contrast, GPT-4 Turbo, even if it has MoE-like optimizations, does not show the same instability—suggesting a different or better implementation. The key point: Not all MoE models behave the same, and OpenAI has not confirmed if ChatGPT is using MoE at all. Assuming all "frontier models" work identically oversimplifies the reality of AI architectures. Would love to hear thoughts from others who have tested these differences in practice!