4 ms·
> The implicit claim is worthless. Failure to navigate a synthetic graph == failure to solve real world problems. False. This statement is the dictionary defin
by wavemode 11mo ago
> The implicit claim is worthless. Failure to navigate a synthetic graph == failure to solve real world problems. False.
This statement is the dictionary definition of attacking a strawman.
Every new model that is sold to us, is sold on the basis that it performs better than the old model on synthetic benchmarks. This paper presents a different benchmark that those same LLMs perform much worse on.
You can certainly criticize the methodology if the authors have erred in some way, but I'm not sure why it's hard to understand the relevance of the topic itself. If benchmarks are so worthless then go tell that to the LLM companies.