3 ms·
The instruction tuning dataset is only 52,000 rows. It shouldn't be too hard to crowdsource high-quality human answers to this many questions and retrain the mo
by freedmand 4y ago
The instruction tuning dataset is only 52,000 rows. It shouldn't be too hard to crowdsource high-quality human answers to this many questions and retrain the model, nixing the dependency on OpenAI.
- Tiberium 4y agoSuch a thing already exists and there were some results - https://open-assistant.io https://open-assistant.io I'm not sure why the authors of Alpaca didn't try to train it on this dataset.
- deleted 4y ago[deleted]
- IanCal 4y agoThat dataset isn't released yet. > Can I download the data? > You will be able to, under CC BY 4.0, but it's not released yet. We want to remove spam and PII before releasing it.
- losteric 4y agoThere's the OIG dataset (https://laion.ai/blog/oig-dataset/ https://laion.ai/blog/oig-dataset/) which was used to train a NeoX 20B ChatBot (https://huggingface.co/togethercomputer/GPT-NeoXT-Chat-Base-20B https://huggingface.co/togethercomputer/GPT-NeoXT-Chat-Base-...). The dataset is larger and publicly available. I want to try finetuning LLaMa on this tonight.
- Jack5500 4y agohow did it go?
- ilaksh 4y agoWow.. I really hope someone will train this model with that dataset. Or maybe open assistant will pick it up. The results looks so promising.