6 ms·
Reading through the Chain of Thought for the provided Cipher example (go to the example, click "Show Chain of Thought") is kind of crazy...it literally spells o
by Hansenq 2y ago
Reading through the Chain of Thought for the provided Cipher example (go to the example, click "Show Chain of Thought") is kind of crazy...it literally spells out every thinking step that someone would go through mentally in their head to figure out the cipher (even useless ones like "Hmm"!). It really seems like slowing down and writing down the logic it's using and reasoning over that makes it better at logic, similar to how you're taught to do so in school.
- Jasper_ 2y ago> Average:18/2=9 > 9 corresponds to 'i'(9='i') > But 'i' is 9, so that seems off by 1. Still seems bad at counting, as ever.
- deleted 2y ago[deleted]
- dymk 2y agoThe next line is it catching its own mistake, and noting i = 9.
- PoignardAzur 2y agoIt's interesting that it makes that mistake, but then catches it a few lines later. A common complaint about LLMs is that once they make a mistake, they will keep making it and write the rest of their completion under the assumption that everything before was correct. Even if they've been RLHF to take human feedback into account and the human points out the mistake, their answer is "Certainly! Here's the corrected version" and then they write something that makes the same mistake. So it's interesting that this model does something that appears to be self-correction.
- afro88 2y agoSeeing the "hmmm", "perfect!" etc. one can easily imagine the kind of training data that humans created for this. Being told to literally speak their mind as they work out complex problems.
- seydor 2y agolooks a bit like 'code', using keywords 'Hmm', 'Alternatively', 'Perfect'
- thomasahle 2y agoRight, these are not mere "filler words", but initialize specific reasoning paths.
- legel 2y agoAs a technical engineer, I’ve learned the value of starting sentences with “basically”, even when I’m facing technical uncertainty. Basically, “basically” forces me to be simple. Being trained to say words like “Alternatively”, “But…”, “Wait!”, “So,” … based on some metric of value in focusing / switching elsewhere / … is basically brilliant.
- impossiblefork 2y agoEven though there's of course no guarantee of people getting these chain of thought traces, or whatever one is to call them, I can imagine these being very useful for people learning competitive mathematics, because it must in fact give the full reasoning, and transformers in themselves aren't really that smart, usually, so it's probably feasible for a person with very normal intellectual abilities to reproduce these traces with practice.
- Salgat 2y agoIt's interesting how it basically generates a larger sample size to create a regression against. The larger the input, the larger the surface area it can compare against existing training data (implicitly through regression of course).
- crazygringo 2y agoSeriously. I actually feel as impressed by the chain of thought, as I was when ChatGPT first came out. This isn't "just" autocompletion anymore, this is actual step-by-step reasoning full of ideas and dead ends and refinement, just like humans do when solving problems. Even if it is still ultimately being powered by "autocompletion". But then it makes me wonder about human reasoning, and what if it's similar? Just following basic patterns of "thinking steps" that ultimately aren't any different from "English language grammar steps"? This is truly making me wonder if LLM's are actually far more powerful than we thought at first, and if it's just a matter of figuring out how to plug them together in the right configurations, like "making them think".
- AndyKelley 2y agoYou ever see that scene from Westworld? (spoiler) https://www.youtube.com/watch?v=ZnxJRYit44k https://www.youtube.com/watch?v=ZnxJRYit44k
- dsign 2y agoYou are just catching up to this idea, probably after hearing 2^n explanations about why we humans are superiors to <<fill in here our latest creation>>. I'm not the kind of scientist that can say how good an LLM is for human reasoning, but I know that we humans are very incentivized and kind of good at scaling, composing and perfecting things. If there is money to pay for human effort, we will play God no-problem, and maybe outdo the divine. Which makes me wonder, isn't there any other problem in our bucket list to dump ginormous amounts of effort at... maybe something more worth-while than engineering the thing that will replace Homo Sapiens?
- Nadya 2y agoWhen an AI makes a silly math mistake we say it is bad at math and laugh at how dumb it is. Some people extrapolate this to "they'll never get any better and will always be a dumb toy that gets things wrong". When I forget to carry a 1 when doing a math problem we call it "human error" even if I make that mistake an embarrassing number of times throughout my lifetime. Do I think LLM's are alive/close to ASI? No. Will they get there? If it's even at all possible - almost certainly one day. Do I think people severely underestimate AI's ability to solve problems while significantly overestimating their own? Absolutely 10,000%. If there is one thing I've learned from watching the AI discussion over the past 10-20 years its that people have overinflated egos and a crazy amount of hubris. "Today is the worst that it will ever be." applies to an awful large number of things that people work on creating and improving.
- cowsaymoo 2y ago> THERE ARE THREE R'S IN STRAWBERRY hilarious
- evilfred 2y agowhich makes it even funnier when the Chain is just... wrong https://x.com/colin_fraser/status/1834336440819614036 https://x.com/colin_fraser/status/1834336440819614036
- davesque 2y agoYes and apparently we won't have access to that chain of thought in the release version: "after weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring, we have decided not to show the raw chains of thought to users"