3 ms·
I was under the impression that “chain of thought” was getting the LLM itself to plan out the steps before it solves the problem. Is that not true?
by superb_dev 2y ago
I was under the impression that “chain of thought” was getting the LLM itself to plan out the steps before it solves the problem. Is that not true?
- kromem 2y agoIt is true, but when it's not zero shot there's a possibility that you are introducing additional information with the example. As well, depending on the study there have been issues with effectively a 'halting' assistance where as long as it isn't getting it right it keeps on going until getting it right, which is effectively a post-selection bias. But frankly this post reads like someone that doesn't understand much about CoT in the first place, let alone the various methods that have improved upon it since. It reads like one of the ad nauseum "look at me use this tool poorly, clearly it's a poor tool" examples. In general, I've noticed mathematicians, computer scientists, and engineers tend to be very poor at evaluating LLMs because they just aren't very good at correctly identifying the scope and depth of what was modeled in the training data in the first place. It's getting boring watching people foolishly try to evaluate things like "stack these clear blocks" (because that's something I regularly saw in social media posts) while glossing over or actively sabotaging the unbelievable modeling/simulating of much more complex and higher order critical reasoning behind various applications of things like empathy or psychological modeling. For anyone reading this who wants to have a fun project, try creating two versions of a set of word puzzles with the same underlying logic structure. One where the problem and solution are using engineering-ish language like "clear blocks" and another using emotional/social language like "grieving friends." Even as we wait for the next generation of models, there's a lot of people criminally underestimating and underutilizing the current models because they can't look beyond their own specialized domain languages.