4 ms·
One thing I haven't seen in the comments so far is that Llama 2 is tuned with RLHF [0], which the original Llama work wasn't. In addition to all the other "upgr
by zora_goron 3y ago
One thing I haven't seen in the comments so far is that Llama 2 is tuned with RLHF [0], which the original Llama work wasn't. In addition to all the other "upgrades", seems like this will make it far easier to steer the model and get practical value.
[0] Training Llama-2-chat: Llama 2 is pretrained using publicly available online data. An initial version of Llama-2-chat is then created through the use of supervised fine-tuning. Next, Llama-2-chat is iteratively refined using Reinforcement Learning from Human Feedback (RLHF), which includes rejection sampling and proximal policy optimization (PPO).
https://ai.meta.com/resources/models-and-libraries/llama/ https://ai.meta.com/resources/models-and-libraries/llama/
- SparkyMcUnicorn 3y agoOn HF you'll see there's separate Llama-2-Xb and Llama-2-Xb-chat models, and more details on the model cards about -chat being the fine-tuned versions via SFT and RLHF.