3 ms·
> "Finding 1: Scale Buys Evaluation, Not Control" The attached paper's (https://arxiv.org/pdf/2604.16009 https://arxiv.org/pdf/2604.16009) title is "MEDLEY-BEN
by GodelNumbering 3mo ago
> "Finding 1: Scale Buys Evaluation, Not Control"
The attached paper's (https://arxiv.org/pdf/2604.16009 https://arxiv.org/pdf/2604.16009) title is "MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition"
This is the most blatant Claude line, or as Claude would put it, the smoking gun.
- stymaar 3mo agoBut is it load-bearing?
- malfist 3mo agoI thought it was a belt and suspenders conclusion
- edot 3mo agoHonestly? It’s the shape that makes it clearly AI. That’s the quiet admission at the heart of the problem.
- tonyarkles 3mo agoThis is one of the things that upsets me the most about LLM writing. “Load bearing” and “belt and suspenders” are two tropes I’ve used for a long, long time and now I have to be intentional about not using them lest I be accused of offloading my writing.
- _joel 3mo agoBelt and braces please Claude, I'm British.
- GodelNumbering 3mo agoThat was the first thing I Ctrl+F'd in the paper, no results haha Broadly, I keep thinking about this over last year or two: while LLMs have nearly eliminated the bar for slop and coding slop, the reviewers are still expected to perform their job diligently. The asymmetry here is extremely taxing for reviewers of all AI generated content. And this is one thing that AI can't help with (as with any statistical process that lacks world understanding and grasp of logical inference). That's why I fully support Arxiv's tough stance on the AI use responsibility.
- usui 3mo agoYou're absolutely right to push back on that. Bottom line⸻it's not load-bearing, it's structural. And honestly⸻that's not nothing. No loads. No bears. Just structure.
- jihadjihad 3mo agoYou're asking the right questions.
- x313 3mo agoThis is jaw-droppingly lazy slop. The authors really didn't put in even an ounce of thought or effort.