3 ms·Post-Training Layer Scaling Prevents Forgetting and Enhances Model Merging1 points by veryluckyxyz 2y ago