3 ms·
pre training and fine tuning use the exact same method of next token prediction. the difference is in the quantity of data you have (& whether the model is pre
by make3 3y ago
pre training and fine tuning use the exact same method of next token prediction. the difference is in the quantity of data you have (& whether the model is pre trained).
you need to train the model on 1 trillion tokens (https://platform.openai.com/tokenizer https://platform.openai.com/tokenizer https://github.com/google/sentencepiece https://github.com/google/sentencepiece) anyways for it to get reasoning capacities, which it feels very unlikely that your data is that much.
I'm highly skeptical that you have enough data to pretrain if you don't have enough data to fine tune.
fine tuning + vector search + prompting of as much stuff as you can, on a LLM like palm2 or gpt4 is what I would do. otherwise you can use falcon 40B ofc.
maybe I should charge for this ahah
- weinzierl 3y agoThe data is not the problem. I could train on any of the public datasets combined with my own data. And here comes my point: The result I'd achieve by training on that combined dataset from scratch cannot be achieved cheaper by utilizing an already pre-trained model of the huge generic dataset plus whatever additional training with the large domain-specific dataset. From what I understand: If you fine-tune only with the domain specific data and just enough so that the model picks up that knowledge it will have forgotten most of its generic knowledge already. If you train on the combined dataset it will take as many epochs for the domain knowledge to shine through as for training the original model. It would cost the same as training from scratch. You need enough tokens but you also need to train your model the right amount on them. Too much or too little training is bad and the base models we have are just trained for the sweet spot and will tolerate a little bit of fine-tuning to adapt their behaviour but not nearly enough to teach them new facts. I'm not an expert, so I could be wrong. What makes me a bit confident is that I have not yet found a single project that reports to have used the approach you suggest successfully.
- make3 3y agoI'm an expert, mix your data with 50% random data. just do it.