4 ms·
>if you are VRAM constrained So this is a perfect model architecture for the alternate realities where nvidia decided to scale up VRAM instead of compute first
by Jackson__ 3y ago
>if you are VRAM constrained
So this is a perfect model architecture for the alternate realities where nvidia decided to scale up VRAM instead of compute first? I'll let them know over trans-dimensional text message.
Also if quantization scales similar per 7b expert as seen in dense LLMs, i.e. the bigger the model, the lower the perplexity loss, this could be the worst performing model at <=4bits compared to anything else currently available :(
-A very sad 24gb 3090 user.
- sebzim4500 3y agoMoE is a great architecture if you are running the model at scale. When you put different layers on different machines, the VRAM used for the parameters doesn't matter that much but the inference compute really does. That's why the SOTA proprietary models are probably all MoE (GPT-3.5/4, palm, gemini, etc.) but until recently no open models were.