3 ms·Transforming LLMs into parallel decoders boosts inference speed by up to 3.5x7 points by snyhlxde 2y ago