3 ms·
But the source to train your own LLM equivalent is also released though (minus the data). Hence why there are so many variants of LLaMa. You also can’t fine tun
by syntaxing 3y ago
But the source to train your own LLM equivalent is also released though (minus the data). Hence why there are so many variants of LLaMa. You also can’t fine tune it without the original model structure. The weights give the community a starting point so they don’t need literally millions of dollar worth of compute power to get to the same step.
- monocasa 3y agoWould Mistral's engineers be satisfied with the release if they had to rebuild from scratch?
- syntaxing 3y agoBut they built a llama equivalent + some enhancements that gives better performance…I’m not sure if this would be possible at all without Meta releasing all the required code and paper for LLaMa to begin with.
- monocasa 3y agoMeta didn't release all of the required code to build LLaMa, just enough run inference with their weights.
- godelski 3y ago> Would Mistral's engineers be satisfied with the release if they had to rebuild from scratch? Yeah, probably. But depends on what you're asking. In the exact same method to get the exact same results down to epsilon error? (Again, ML models are not deterministic) Probably not. This honestly can even change with a different version of pytorch, but yes, knowing the HPs would help get closer. But to train another 7B model of the __exact__ same architecture? Yeah, definitely they've provided all you need for that. You can take this model and train it from scratch on any data you want and train it in any way you want.
- monocasa 3y agoI didn't ask if they'd be able to make do. I asked if they'd be satisfied. Also, wrt > Again, ML models are not deterministic) ML models are absolutely deterministic if you have the discipline to do so (which is necessary at higher scale ML work when hardware is stochastically flaky).