3 ms·
Good point about the negative reinforcement training! Instruction tuning is on top of the base LLM and is often RLHF to train the base LLM to produce certain k
by dbmikus 3y ago
Good point about the negative reinforcement training!
Instruction tuning is on top of the base LLM and is often RLHF to train the base LLM to produce certain kinds of responses.