14 ms·
Chain-of-thought can hurt performance on tasks where thinking makes humans worse
- oatsandsugar 2y agoTasks were thinking makes human worse > Three such cases are implicit statistical learning, visual recognition, and classifying with patterns containing exceptions. Fascinating that our lizard brains are better at implicit statistical reasoning
- Dilettante_ 2y agoWell, by definition, thinking is always explicit reasoning, no? And I'd hazard a guess that a well-thought through Fermi Estimation beats lizard-brain eyeballing every time, it's just that in the inbetween space the two interfere unfavourably.
- YetAnotherNick 2y agoMy guess would be no. I have terrible face recognition ability and I can look into face for hour and still other people could easily beat me in less than a second.(I am assuming "well-thought through Fermi Estimation" would be similar for me and others in this case).
- mjcohen 2y agoLook into a disease called faceblindness (there is a fancy name I forget).
- Terr_ 2y ago> Well, by definition, thinking is always explicit reasoning, no? That doesn't feel right to me. (Heh, accidentally appropriate word choice.) There are a lot of tasks we do that are arguably "thinking" yet don't involve an internal "Oh, hey, I'm gonna solve this problem, I'm thinking right now." For example, imagine you're at a park, and someone is feeding the ducks. Another person walks up behind them and sucker-punches them into the pond. It should be almost a reflex [0] that you'll conclude "the puncher is bad" and "the person in the water needs help" without explicitly reasoning out. I think that task qualifies as "thinking", especially since it involves some kind of theory-of-mind about those other humans. [0] An exception might be someone with a sociopathic disability, who would have to think more-explicitly to realize what reaction is expected of them.
- brewii 2y agoThink about how fast you’re able to determine the exact trajectory of a ball and location to place your hand to catch a ball using your lizard brain.
- asah 2y agoyou mean like pingpong? https://arstechnica.com/information-technology/2024/08/man-vs-machine-deepminds-new-robot-serves-up-a-table-tennis-triumph/ https://arstechnica.com/information-technology/2024/08/man-v...
- dools 2y agoBender: Now Wireless Joe Jackson, there was a blern-hitting machine! Leela: Exactly! He was a machine designed to hit blerns!
- taeric 2y agoThis isn't some innate ability that people have. As evidenced by how bad my kids are at catching things. :D That said, I think this is a good example. We call it "muscle memory" in that you are good at what you have trained at. Change a parameter in it, though, and your execution will almost certainly suffer.
- skrtskrt 2y agoI mean even people that are "bad at catching things" are still getting ridiculously close to catching it - getting hands to the right area probably within well under a second of the right timing - without being taught anything in particular about how a ball moves through the air.
- taeric 2y agoUh.... have you been around kids? It will take several absurd misses before they even start to respond to a ball in flight.
- daft_pink 2y agothis is exactly what I was looking for. tasks where I should not think and just trust my gut.
- m3kw9 2y agowould be slow to use COT on simple requests like 1+1
- ryoshu 2y ago95% * 95% = 90.25%
- npunt 2y ago"Don't overthink it" is sometimes good advice!
- marviel 2y agoI love backpropagating ideas from ML back into psychology :) I think it shows great promise as a way to sidestep the ethical concerns (and the reproducibility issues) associated with traditional psychology research. One idea in this space I think a lot about is from the Google paper on curiosity and procrastination in reinforcement learning: https://research.google/blog/curiosity-and-procrastination-in-reinforcement-learning/ https://research.google/blog/curiosity-and-procrastination-i... Basically the idea is that you can model curiosity as a reward signal proportional to your prediction error. They do an experiment where they train an ML system to explore a maze using curiosity, and it performs the task more efficiently -- UNTIL they add a "screen" in the maze that shows random images. In this case, the agent maximizes the curiosity reward by just staring at the screen. Feels a little too relatable sometimes, as a highly curious person with procrastination issues :)
- npunt 2y ago"...in AI" will be the psychology equivalent of biology's "...in Mice"
- marviel 2y agoIt will! Not 1:1, has issues, but gives hints. Also much more scalable.
- miningape 2y ago> Not 1:1, has issues, but gives hints. > Also much more scalable. This same description could be applied to lab mice
- Terr_ 2y agoIt'll probably be a ways before we start making shrines to their unwilling participation though. https://en.wikipedia.org/wiki/Monument_to_the_laboratory_mouse https://en.wikipedia.org/wiki/Monument_to_the_laboratory_mou...
- gpsx 2y agoI saw an LLM having this kind of problem when I was doing some testing a ways back. I asked it to order three fruits from largest to smallest. I think it was orange, blueberry and grapefruit. It could do that easily with a simple prompt. When the prompting included something to the effect of “think step by step”, it would try to talk through the problem and it would usually get it wrong.
- spockz 2y agoHow much does this align with how we learn math? We kind of instinctively learn the answers to simple math questions. We can even at some point develop an intuition for things like integrating and differentials. But the moment we are asked to explain why, or worse provide a proof, things become a lot harder. Even though the initial answer may be correct.
- larodi 2y agoI definitely don’t learn math by means of gradient descents. We can possibly say math is not learned, but a mental models of abstractions are developed. How? We dunno, but what we do know is we don’t learn by figuring the common features between all previously seen equations only to guess them later… Mind operates on higher and higher levels of abstractions building on each other in a much fascinating way, very often not with words, but with structure and images. Of course there are people with aphantasia, but i really fail to see how any reasoning happens in purely language level. Someone on this forum also noted - in order to reason one needs an ontology to facilitate the reasoning process. LLMs don’t do ontologies… And finally, not least though, LLM and ML people in general seem to equate intuition to some sort biased.random(). Well intuition is not random, and is hard to describe in words. So are awe and inspiration. And these ARE part of (precondition to, fuel for) humanity’s thought process more that we like to admit.
- shotnothing 2y ago> I definitely don’t learn math by means of gradient descents. https://physoc.onlinelibrary.wiley.com/doi/10.1113/JP282747 https://physoc.onlinelibrary.wiley.com/doi/10.1113/JP282747
- veryfancy 2y agoSo like dating?
- Terr_ 2y agoAlternate framing: A powerful autocomplete algorithm is being used to iteratively extend an existing document based on its training set. Sometimes you get a less-desirable end-result when you intervene to change the style of the document away from question-and-answer to something less common.
- youoy 2y agoThat's what one half of HN think. The other half: Artificial brains in the verge of singularity show another sign of approaching consciousness. The chain of thought of process performance is exactly human, showing yet another proof of the arrival of AGI before 2030.
- lazide 2y agoPfft, 2030?!? It’s already in the middle of manipulating the election! (/s, kinda)
- fiso64 2y agoA framing that is longer, far harder to parse, and carries less information.
- grain-o-salt 2y agoLet me give it a try...um...what about 'Star Trek' vs.: A delivering-service called Galaxyray?galaxyray brings wares and hot tasty meals galaxywide to recipients, even while they are 'traveling' with more-than-lightspeed in hyperspace? > ..ordered by Imperium just to troll the retros!? Sounds "less comon"...hu...?! P-: Ok! Ok! let me try to explain it a bit more, the whole Universe projected as a beam, say... scalable, 100m, placed in a storage depot, a 'parralaxy' ...So delivery agents are grabbing the ordered stuff and...no? Not? As reasonable like your answer is, do that sound very 'uncommon' while 'phrasing that with many questionmarks'? ?? Enjoying my day off... (-: regards,
- Y_Y 2y agoReminds me of a mantra from chess class: long think = wrong think
- TZubiri 2y agoWas that perhaps a speed chess class?
- hackable_sand 2y agoI prefer to call it Kung fu Because you feel like a martial artist.
- Y_Y 2y agoNope, just vanilla otb slow chess
- spongebobism 2y agoThe original by Bent Larsen is "Long variation, wrong variation"
- meowster 2y agoThink long; think wrong ( Flows off the tongue better ¯\_(ツ)_/¯ )
- TZubiri 2y agoSo, LLMs face a regression on their latest proposed improvement. It's not surprising considering their functional requirements are: 1) Everything For the purpose of AGI, LLM are starting to look like a local maximum.
- rjbwork 2y ago>For the purpose of AGI, LLM are starting to look like a local maximum. I've been saying it since they started popping off last year and everyone was getting euphoric about them. I'm basically a layman - a pretty good programmer and software engineer, and took a statistics and AI class 13 years ago in university. That said, it just seems so extremely obvious to me that these things are likely not the way to AGI. They're not reasoning systems. They don't work with axioms. They don't model reality. They don't really do anything. They just generate stochastic output from the probabilities of symbols appearing in a particular order in a given corpus. It continues to astound me how much money is being dumped into these things.
- ChadNauseam 2y agoHow do you know that they don’t do these things? Seems hard to say for sure since it’s hard to explain in human terms what a neural network is doing.
- nephy 2y agoIf you give an LLM a word problem that involves the same math and change the names of the people in the word problem the LLM will likely generate different mathematical results. Without any knowledge of how any of this works, that seems pretty damning of the fact that LLMs do not reason. They are predictive text models. That’s it.
- alexwebb2 2y agoDemonstrably false. https://chatgpt.com/share/6722ca8a-6c80-800d-89b9-be40874c5b65 https://chatgpt.com/share/6722ca8a-6c80-800d-89b9-be40874c5b... https://chatgpt.com/share/6722ca97-4974-800d-99c2-bb58c60ea632 https://chatgpt.com/share/6722ca97-4974-800d-99c2-bb58c60ea6...
- mitko 2y agoThis is so uncannily close to the problems we're encountering at Pioneer, trying to make human+LLM workflows in high stakes / high complexity situations. Humans are so smart and do so many decisions and calculations on the subconscious/implicit level and take a lot of mental shortcuts, so that as we try to automate this by following exactly what the process is, we bring a lot of the implicit thinking out on the surface, and that slows everything down. So we've had to be creative about how we build LLM workflows.
- lolinder 2y agoThis is a regression in the model's accuracy at certain tasks when using COT, not its speed: > In extensive experiments across all three settings, we find that a diverse collection of state-of-the-art models exhibit significant drop-offs in performance (e.g., up to 36.3% absolute accuracy for OpenAI o1-preview compared to GPT-4o) when using inference-time reasoning compared to zero-shot counterparts. In other words, the issue they're identifying is that COT is an less effective model for some tasks compared to unmodified chat completion, not just that it slows everything down.
- mitko 2y agoYeah! That's the danger with any kind of "model" whether it is CoT, CrewAI, or other ways to outsmart it. It is betting that a programmer/operator can break a large tasks up in a better way than an LLM can keep attention (assuming it can fit the info in the context window). ChatGPT's o1 model could make a lot of those programming techniques less effective, but they may still be around as they are more manageable, and constrained.
- haccount 2y agoLanguage seems to be confused with logic or common sense. We've observed it previously in psychiatry(and modern journalism, but here I digress) but LLMs have made it obvious that grammatically correct, naturally flowing language requires a "world" model of the language and close to nothing of reality, spatial understanding? social clues? common sense logic? or mathematical logic? All optional. I'd suggest we call the LLM language fundament a "Word Model"(not a typo). Trying to distil a world model out of the word model. A suitable starting point for a modern remake of Plato's cave.
- deleted 2y ago[deleted]
- nisten 2y agoThis sounds about right from my experience getting nerdsniped by new samplers along with trying to reproduce the API middleware for the whole reflection thing, and using 4400 questions for a new benchmark is not bad given that even the well-regarded gpqa benchmark is only 3000-something questions. What's ... mildly infuriating here is the lack of any kind of data, code, 0 mention of github in the paper, and nothing for anyone to reproduce or find any reason in my opinion to even recommend anyone to read this thing at all. If you think that whatever you're doing in the field of LLMs won't be obsolete in 6 months you're being delusional. Anyway, back to the paper, it says all questions culminated to a yes or no answer... meaning theres a 50/50 chance of getting right, so does that mean the 8% drop in performance you got from testing llama 3 8b this way is more like 4% which would make it statistically insignificant? And given that the only other scientifically usueful & reproducible (non-api walled models which no one knows on how many actual llms and retrieval systems are composing that solution you're testing)models were less than that leads me to the opinion that this whole thing was just useless slop. So please, if you're writing a paper in LLMs, and want to seem credible, either have some type of demo thing or show the actual god damn trash code and top secret garbage data you wrote for it so people can make some kind of use of it before it goes obsolete otherwise you're just wasting everyones time. TL:DR. It's trash.
- alexchantavy 2y agoThis seems to support how thinking out loud during a coding test might make you do worse.
- why-el 2y agoI like this analogy a lot. It's possible that forced externalization of thoughts accidentally causes the omission of crucial data. That is, much more goes on in your head, you probably laid out the whole algorithm, but being asked to state it on the spot and in clear, serial words is causing you to bork it by taking shortcuts.
- jwpapi 2y agoThis is so interesting. What are even the tasks where thinking makes humans worse?
- sigmoid10 2y agoThe answer is in the article. One example they give is grammar. Lots of people apparently do worse once they try to verbalize it.
- XCSme 2y ago> What are even the tasks where thinking makes humans worse? Not really related, but athletes perform A LOT worse when they are thinking about their movements/strategies/tactics. A top performing athlete does best when they are in a flow state, where they don't think about anything and just let their body/muscle memory do the work. Once you start thinking about micro-adjustments (e.g. I should lift my elbow higher), you start controlling your body in a conscious way, which is a magnitude slower and less coordinated than the automatic/subconscious way. Also, same happens for creativity/new ideas. If you intentionally think about something, step by step, you won't likely find new, innovative solutions. There is a reason why the "a-ha!" moments come in the shower, your subconscious mind is thinking about the problem instead of trying to force your thinking on a specific path. I would guess this happens in many other areas, where channelling the thought process through a specific template hinders the ability to use all the available resources/brain power.
- naasking 2y ago> What are even the tasks where thinking makes humans worse? Talking about religion and politics.
- sowbug 2y agoI can think myself into forgetting strong passwords if I try to spell each character out in my head. But then I sit at a keyboard, relax, and automatically type it perfectly.
- lucianbr 2y agoMuscle memory or something like it hardly seems a step towards AGI. Or towards solving any difficult problems.
- wg0 2y agoNot to mention that chain of thought is computationally very expensive. Prohibitively expensive for sure to be served free like previous generation of Web 2.0 products. Seems like repeated promoting can't juice AGI out of token probabilities. Retrospectively, if you can pin point one paper that led to the bust and pop of the AI bubble, this would be it.
- varelse 2y ago[dead]
- dev1ycan 2y agoStop dumping billions of your own money (if you are a VC) in LLMs, you are going to regret it in the long run. You are funding con-artist's salaries...
- cainxinth 2y agoThis says something fascinating about information processing in both biological and AI systems. Both systems compress information: the brain creates efficient neural patterns through experience and AI develops internal representations through training. Forcing verbalization "decompresses" this efficient encoding, potentially losing subtle patterns. Hence, for a task like visual recognition, which is optimized to occur almost instantly in a parallel process, you will only degrade performance by running it in a serial chain of thought sequence.