Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
addiefoote8
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
addiefoote8
7mo ago
I'm also excited about the research that could be enabled by having weight-level access and fine tuning access on frontier open source models. There's a lot of interesting behavior that just doesn't exist in 8B parameter mode
2.
▲
by
addiefoote8
7mo ago
Some more details on this: After realizing Hugging Face would be messy to work with to train Kimi-k2-thinking, we decided to do it ourselves. We started with PrimeRL and implemented Kimi in it, verifying it against the Moonshot API. The ini
3.
▲
50x Faster Post-Training
(workshoplabs.ai)
10 points
by
addiefoote8
7mo ago
|
4 comments
4.
▲
by
addiefoote8
7mo ago
Unsloth doesn't support distributed training well and doesn't support Kimi models.
5.
▲
by
addiefoote8
7mo ago
I agree full transparency on data adds several other challenges. Still, even releasing the software and infrastructure aspects would be a huge step from where we are now. Also, some recent work has shown pretraining filtering to be possible
6.
▲
by
addiefoote8
7mo ago
I'd also add training checkpoints to the list for active transparency. I think the Olmo models do a decent job, but it would be cool to see it for bigger models and for ones that are closer to state-of-the-art in terms of both architec
7.
▲
by
addiefoote8
7mo ago
yeah, the costs are definitely a factor and prohibitive in completely replicating an open source model. Still, there's a lot of useful things that can be done cheaply, including fine tuning, interpretability work, and other deeper inve
8.
▲
Open Weights isn't Open Training
(workshoplabs.ai)
121 points
by
addiefoote8
7mo ago
|
38 comments