3 ms·
Super shallow (24/36 layers) MoE with low active parameter counts (3.6B/5.1B), a tradeoff between inference speed and performance. Text only, which is okay. W
by RandyOrion 1y ago
Super shallow (24/36 layers) MoE with low active parameter counts (3.6B/5.1B), a tradeoff between inference speed and performance.
Text only, which is okay.
Weights partially in MXFP4, but no cuda kernel support for RTX 50 series (sm120). Why? This is a NO for me.
Safety alignment shifts from off the charts to off the rails really fast if you keep prompting. This is a NO for me.
In summary, a solid NO for me.