3 ms·
Conjecture: Does GLM only handle the prefill stage, compress its output hidden vectors (trained?), and then send them to Qwen for decoding?
by arpperzhao 28d ago
Conjecture: Does GLM only handle the prefill stage, compress its output hidden vectors (trained?), and then send them to Qwen for decoding?