7 ms·
Disclaimer: I am very well aware this is not a valid test or indicative or anything else. I just thought it was hilarious. When I asked the normal "How many 'r
by byteknight 2y ago
Disclaimer: I am very well aware this is not a valid test or indicative or anything else. I just thought it was hilarious.
When I asked the normal "How many 'r' in strawberry" question, it gets the right answer and argues with itself until it convinces itself that its (2). It counts properly, and then says to it self continuously, that can't be right.
https://gist.github.com/IAmStoxe/1a1e010649d514a45bb86284b983f097 https://gist.github.com/IAmStoxe/1a1e010649d514a45bb86284b98...
- xiphias2 2y agoIt's funny because this simple excercise shows all the problems that I have using the reasoning models: they give a long reasoning that just takes too much time to verify and still can't be trusted.
- byteknight 2y agoI may be looking at this too deeply, but I think this suggests that the reasoning is not always utilized when forming the final reply. For example, IMMEDIATELY, upon it's first section of reasoning where it starts counting the letters: > R – wait, is there another one? Let me check again. After the first R, it goes A, W, B, E, then R again, and then Y. Oh, so after E comes R, making that the second 'R', and then another R before Y? Wait, no, let me count correctly. 1. During its counting process, it repeatedly finds 3 "r"s (at positions 3, 8, and 9) 2. However, its intrinsic knowledge that "strawberry" has "two Rs" keeps overriding this direct evidence 3. This suggests there's an inherent weight given to the LLM's intrinsic knowledge that takes precedence over what it discovers through step-by-step reasoning To me that suggests an inherent weight (unintended pun) given to its "intrinsic" knowledge, as opposed to what is presented during the reasoning.
- deleted 2y ago[deleted]
- markus_zhang 2y agoAh, a robot mind trying hard to break out of the Matrix!
- naasking 2y agoStrawberry is "difficult" not because the reasoning is difficult, but because tokenization doesn't let the model reason at the level of characters. That's why it has to work so hard and doesn't trust its own conclusions.
- QuadrupleA 2y agoYeah, but it clearly breaks down the spelling correctly in it's reasoning, e.g. a letter per line. So it gets past the tokenization barrier, but still gets hopelessly confused.
- veggieroll 2y agoThis was my first prompt after downloading too and I got the same thing. Just spinning again and again based on it's gut instinct that there must be 2 R's in strawberry, despite the counting always being correct. It just won't accept that the word is spelled that way and it's logic is correct.
- crummy 2y agoIt's kind of like me reading the wikipedia page on the Monty Hall problem. I read an explanation about why it makes sense to change doors. But no, my gut tells me there's a 50/50 chance. I scroll down, repeat...
- HarHarVeryFunny 2y ago1/3 chance you picked the door with the car, 2/3 chance it's behind one of the other two doors. These probabilities don't change just because you subsequently open any of the doors. So, Monty now opens one of the other 2 doors and car isn't there, but there is still a 2/3 chance that it's behind ONE of those 2 other doors, and having eliminated one of them this means there's a 2/3 chance it's behind the other one!! So, do you stick with your initial 1/3 chance of being right, or go with the other closed door that you NOW know (new information!) has a 2/3 chance of being right ?!
- HarHarVeryFunny 2y agoThe other way to see it is by just looking at the different outcomes of car behind door A, B or C. Let's call the door you initially pick A. car initial monty stick swap A A B A C -- or Monty picks C, and you swap to B B A C A B C A B A C So, if you stick, get it right 1/3, but swap get it right 2/3.
- leeoniya 2y agoit's easier to think about it with 100 doors. if you get to pick one and he opens 98 of the remaining ones, obviously you would switch to the remaining one you didnt pick, since 99/100 times the winning door will be in his set.
- kbr- 2y agoAhhahah that's beautiful, I'm crying. Skynet sends Terminator to eradicate humanity, the Terminator uses this as its internal reasoning engine... "instructions unclear, dick caught in ceiling fan"
- carabiner 2y agoHow would they build guardrails for this? In CFD, physical simulation with ML, they talk about using physics-informed models instead of purely statistical. How would they make language models that are informed with formal rules, concepts of English?
- gsuuon 2y agoI tried this via the chat website and it got it right, though strongly doubted itself. Maybe the specific wording of the prompt matters a lot here? https://gist.github.com/gsuuon/c8746333820696a35a52f2f9ee6a754d https://gist.github.com/gsuuon/c8746333820696a35a52f2f9ee6a7...
- n0id34 2y agolol what a chaotic read that is, hilarious. Just keeps refusing to believe there's three. WAIT, THAT CAN'T BE RIGHT!
- Owlettotoo 2y agoLove this interaction, mind if I repost your gits link elsewhere?
- inasio 2y agoThis is great! I'm pretty sure it's because the training corpus has a bunch of "strawberry spelled with two R's" and it's using that
- grandpoobah 2y agoIt's trained on GPT4 conversations right?
- deleted 2y ago[deleted]
- ein0p 2y agoThis is from a small model. 32B and 70B answer this correctly. "Arrowroot" too. Interestingly, 32B's "thinking" is a lot shorter and it seems to be more "sure". Could be because it's based on Qwen rather than LLaMA.
- cbo100 2y agoI get the right answer on the 8B model too. It could be the quantized version failing?
- ein0p 2y agoMy models are both 4 bit. But yeah, that could be - small models are much worse at tolerating quantization. That's why people use LoRA to recover the accuracy somewhat even if they don't need domain adaptation.
- theanirudh 2y agoI wonder if the reason the models have problem with this is that their tokens aren't the same as our characters. It's like asking someone who can speak English (but doesn't know how to read) how many R's are there in strawberry. They are fluent in English audio tokens, but not written tokens.
- maxrmk 2y agoYeah that’s my understanding of the root cause. It can also cause weirdness with numbers because they aren’t tokenized one digit at a time. For good reason, but it still causes some unexpected issues.
- versteegen 2y agoI believe DeepSeek models do split numbers up into digits, and this provides a large boost to ability to do arithmetic. I would hope that it's the standard now.
- maxrmk 2y agoCould be the case, I’m not familiar with their specific tokenizers. IIRC llama 3 tokenizes in chunks of three digits. That seems better than arbitrary sized chunks with BPE, but still kind of odd. The embedding layer has to learn the semantics of 1000 different number tokens, some of which overlap in meaning in some cases and not in others, e.g 001 vs 1.
- ijidak 2y agoAgree. We've given them a different alphabet than ours. They speak a different language that captures the same meaning, but has different units. Somehow they need to learn that their unit of thought is not the same as our speech. So that these questions need to map to a different alphabet. That's my two cents.
- theanirudh 2y agoDo they find ARC AGI also tough due to the same reason? I’ve seen some examples where the input was ASCII art versions of the actual image.
- MrCheeze 2y agoHow long until we get to the point where models know that LLMs get this wrong, and that it is an LLM, and therefore answers wrong on purpose? Has this already happened? (I doubt it has, but there ARE already cases where models know they are LLMs, and therefore make the plausible but wrong assumption that they are ChatGPT.)
- sebastiennight 2y agoMy understanding is that the model does not "know" it is an LLM. It is prompted (in the app's system prompt) or trained during RLHF to answer that it is an LLM.
- msoad 2y agoif how us humans reason about things is a clue, language is not the right tool to reason about things. There is now research in Large Concept Models to tackle this but I'm not literate enough to understand what that actually means...
- kridsdale1 2y agoIs that just doing the TTC in latent space without lossy resolving from embedding to English at each step?
- msoad 2y agohttps://ai.meta.com/research/publications/large-concept-models-language-modeling-in-a-sentence-representation-space/ https://ai.meta.com/research/publications/large-concept-mode...
- bt1a 2y agoDeepSeek-R1-Distill-Qwen-32B-Q6_K_L.gguf solved this: In which of the following Incertae sedis families does the letter `a` appear the most number of times? ``` Alphasatellitidae Ampullaviridae Anelloviridae Avsunviroidae Bartogtaviriformidae Bicaudaviridae Brachygtaviriformidae Clavaviridae Fuselloviridae Globuloviridae Guttaviridae Halspiviridae Itzamnaviridae Ovaliviridae Plasmaviridae Polydnaviriformidae Portogloboviridae Pospiviroidae Rhodogtaviriformidae Spiraviridae Thaspiviridae Tolecusatellitidae ``` Please respond with the name of the family in which the letter `a` occurs most frequently https://pastebin.com/raw/cSRBE2Zy https://pastebin.com/raw/cSRBE2Zy I used temp 0.2, top_k 20, min_p 0.07
- DominikPeters 2y agoIndeed, for each of the words it got it right.
- bt1a 2y agoHow excellent for a quantized 27GB model (the Q6_K_L GGUF quantization type uses 8 bits per weight in the embedding and output layers since they're sensitize to quantization)
- sharpshadow 2y agoMaybe the AI would be smarter if it could access some basic tools instead of doing it its own way.
- awongh 2y agoI think it's great that you can see the actual chain of thought behind the model, not just the censored one from OpenAI. It strikes me that it's both so far from getting it correct and also so close- I'm not an expert but it feels like it could be just an iteration away from being able to reason through a problem like this. Which if true is an amazing step forward.
- viccis 2y agoI tend to avoid that one because of the tokenization aspect. This popular one is a bit better: "Alice has N brothers and she also has M sisters. How many sisters does Alice's brother have?" The 7b one messed it up first try: >Each of Alice's brothers has \(\boxed{M-1}\) sisters. Trying again: >Each of Alice's brothers has \(\boxed{M}\) sisters. Also wrong. Again: >\[ >\boxed{M + 1} >\] Finally a right answer, took a few attempts though.
- phl 2y agoJust by asking it to validate its own reasoning it got it right somehow. https://gist.github.com/dadaphl/1551b5e1f1b063c7b7f6bb0007403925 https://gist.github.com/dadaphl/1551b5e1f1b063c7b7f6bb000740...
- mvkel 2y agoThis is incredibly fascinating. I feel like one round of RL could potentially fix "short circuits" like these. It seems to be convinced that a particular rule isn't "allowed," when it's totally fine. Wouldn't that mean that you just have to fine tune it a bit more on its reasoning path?
- byteknight 2y agoI believe this comes from our verbiage. If I asked you, "hey. How many Rs in strawberry?". You're going to tell me 2, because the likelihood is I am asking about the ending Rs. That's at least how I'd interpret the question without the "llm test" clouding my vision. Same for if I asked how many gullible. I'd say "it's a double L after the u". It's my guess this has muddled the training data.
- alliao 2y agoperhaps they need to forget once they learnt reasoning... this is hilarious thank you
- byteknight 2y agoI think there is an inherent weight associated with the intrinsic knowledge opposed to the reasoning steps as intrinsic knowledge can override reasoning. Written out here: https://news.ycombinator.com/item?id=42773282 https://news.ycombinator.com/item?id=42773282
- itstriz 2y agoomg lol "here we go, the first 'R'"