3 ms·Reinforcement Learning Finetunes Small Subnetworks in Large Language Models3 points by s-macke 1y ago