4 ms·Simple, zero overhead way to compress model, KV cache via Low-Rank Decomposition1 points by thw20 5mo ago