4 ms·CAD: Disaggregating Core Attention for Efficient Long-Context LLM Training6 points by ginda307 10mo agodeleted 10mo ago[deleted]