3 ms·We used sparse autoencoders to explain LLM moderation flags of violent threats6 points by karinemellata 1y ago