4 ms·Current Large Audio Language Models largely transcribe rather than listen4 points by earcar 8mo ago