3 ms·
> Then a US compute provider should be able to launch a similarly-priced competitor Right, you just need a few months to implement efficient inference for MLA
by rfoo 2y ago
> Then a US compute provider should be able to launch a similarly-priced competitor
Right, you just need a few months to implement efficient inference for MLA + their strangely looking MoE scheme + ..., easy!
Oh wait, the inference scheme described in their tech report is pretty much an exact fit for H800s. So if you run the recipe on H100s you are wasting the potential of your H100s. Otherwise have fun making variations to the serving architecture.
To be fair, we had chance. If someone decided to replicate the effort to serve their models back in May 2024 when DeepSeek-V2 was out we'd have it now. But nobody had interest as DS-V2 was pretty mediocre. They (and whoever realized the potential) made big bet and it is paying off.