2 ms·
All inference providers (labs and neoclouds like TogetherAI or Fireworks) are incentivized to be efficient to be more competitive. For example MoE reduces compu
by brunaxLorax 2mo ago
All inference providers (labs and neoclouds like TogetherAI or Fireworks) are incentivized to be efficient to be more competitive. For example MoE reduces compute without reducing output quality. I would not be surprised if they end up implementing some kind of internal routing at some point.