3 ms·
Some more details on this: After realizing Hugging Face would be messy to work with to train Kimi-k2-thinking, we decided to do it ourselves. We started with P
by addiefoote8 7mo ago
Some more details on this:
After realizing Hugging Face would be messy to work with to train Kimi-k2-thinking, we decided to do it ourselves.
We started with PrimeRL and implemented Kimi in it, verifying it against the Moonshot API. The initial distributed training method, FSDP, is not ideal for memory bottlenecked MoEs, so we added support for Expert Parallel. This enabled faster training, but many optimizations remained. We discuss several in the post, and collectively, these efforts took us from training 125 tokens/s to 6,660 tokens/s on a single 8xH200 node! Per token, our codebase is cheaper than anything on the market, including training APIs like Tinker.
We plan to open source in the coming week or two, pending safety evals!