4 ms·
The disconnect here is that the model code isn’t the important part. If you read the paper, LLaMA is all about the training data and the training process (which
by throwaway1851 4y ago
The disconnect here is that the model code isn’t the important part. If you read the paper, LLaMA is all about the training data and the training process (which required a ”massive quantity” of compute, to quote the paper). Without releasing the significant part of their work, what are they really releasing?
- p1esk 4y agoExactly - massive amount of compute required to reproduce the training process, and that’s why almost no one is going to try training this model from scratch. Most people will use this model as is, and a small number of people will finetune it on their own data - those people most likely already have their own training scripts, because finetuning is much easier than training from scratch and requires different hyper parameters. I don’t know why they haven’t released the training code - I agree it would be nice if they did - but the important thing here is they created something valuable, open sourced it, and even provided model weights - for free. Let’s appreciate it. Even if they did publish their training code - that’s not enough to train the model. You also need their dataset. Would you still claim “the model is not open source” because the dataset is not available? Bottom line: the model has been open sourced. The training code hasn’t - and isn’t needed by most users of this model.
- kaoD 4y ago> You also need their dataset. With the training code you can fine tune the model with your own dataset, which exactly fits FSF's freedom 1.