3 ms·
This is condescending and wrong at the same time (best combo). LLMs do stumble into long prediction chains that don’t lead the inference in any useful directio
by kubb 6mo ago
This is condescending and wrong at the same time (best combo).
LLMs do stumble into long prediction chains that don’t lead the inference in any useful direction, wasting tokens and compute.
- prodigycorp 6mo agoAre you sure about that? Chain of thought does not need to be semantically useful to improve LLM performance. https://arxiv.org/abs/2404.15758 https://arxiv.org/abs/2404.15758
- davidguetta 6mo agostill doesn't mean all tokens are useful. it's the point of benchmarks
- prodigycorp 6mo agoCare to share the benchmarks backing the claims in this repo?
- kubb 6mo agoIf you're misusing LLMs to solve TC^0 problems, which is what the paper is about, then... you also don't need the slop lavine. You can just inject a bunch of filler tokens yourself.