4 ms·
Keep in mind that StarCoder(Base) is just a pretrained LM. The extra stuff that makes 3.5/4 like RLHF gets built on this.
by enum 3y ago
Keep in mind that StarCoder(Base) is just a pretrained LM. The extra stuff that makes 3.5/4 like RLHF gets built on this.
- manojlds 3y agoAren't GPT-3 etc base LM and ChatGPT the instruction tuned? Or am I wrong?
- dpf 3y agocode-davinci-002 is a base LM, and the other 3.5 models (text-davinci-{002,003}, gpt-3.5-turbo, and ChatGPT) use instruction tuning and/or RLHF. Source: https://platform.openai.com/docs/model-index-for-researchers https://platform.openai.com/docs/model-index-for-researchers