3 ms·
The article itself explained that it was much easier to classify text as human or LLM generated than to have "human" as just a category along with all the diffe
by Tade0 3mo ago
The article itself explained that it was much easier to classify text as human or LLM generated than to have "human" as just a category along with all the different LLMs as it's likely the LLMs are distilled from each other, creating a unique footprint.
If a signal is weak, it might not even appear in every sentence, but that doesn't mean it doesn't exist. For instance, I don't recall ever consciously using an em dash, but you'll probably need an entire paragraph to find one in LLM-generated text.
My own sense of whether text is generated is partially based on its sheer length - humans typically don't bother writing so much.