3 ms·
I have no data to back up this claim but one hypothesis I have is it's gotten particularly acutely bad esp with newer gen models because of increased training o
by mvanveen 2mo ago
I have no data to back up this claim but one hypothesis I have is it's gotten particularly acutely bad esp with newer gen models because of increased training on reasoning and chain-of-thought.
To borrow a programming lang analogy the failure mode I see a lot is it invents its own jargon that is effectively like a compiler intermediate representation of the high level natural language you actually want a human to look at and then inserts it directly into what the human has to read.
It needs to stop doing that but it doesn't seem to do a good job at differentiating from what is or isn't chain of thought slop. I'm sure the jargon is useful during its reasoning but it's very unhelpful and not very nice to deliver it to a human.