12 ms·
Can it solve easy problems yet? Weirdly, I think that's an important milestone. Prompts like, "Give me five odd numbers that don't have the letter 'e' in their
by starfezzy 2y ago
Can it solve easy problems yet? Weirdly, I think that's an important milestone.
Prompts like, "Give me five odd numbers that don't have the letter 'e' in their spelling," or "How many 'r's are in the word strawberry?"
I suspect the breakthrough won't be trivial that enables solving trivial questions.
- msoad 2y ago> Give me five odd numbers that don't have the letter 'e' in their spelling Compare the reasoning times!!! 84s vs 342s R1 (Thought for 84 seconds) No odd number in English avoids the letter 'e' in its spelling. The request for five such numbers cannot be fulfilled. o1 Pro (Thought for 5 minutes and 42 seconds) No standard English spelling of an odd number can avoid “e.” Every odd digit (one, three, five, seven, nine) already includes “e,” so once you build any odd number out of those digits, the spelled‐out form will contain “e.” As a result, there are no such odd numbers in standard English without an “e” in their name.
- coffeebeqn 2y agoTook 1m 36s for me. My default prompt is a bit different “think from first principles”. It’s pretty verbose but I enjoyed looking through all the work it did. Pretty impressive !
- mattkrause 2y agoEven after five minutes, the logic has a (small) hole. Not all odd spell out the digit: thirteen, fifteen (and maybe eleven?) are odd but don’t have an ‘e’ in the word’s “digit” part, though they obviously do elsewhere.
- ben_w 2y agoIf the answer is supposed to be logically rigorous, o1 is incomplete for the reason you say. If I was treating this as a lateral thinking puzzle (does anyone still do those for interview questions?) rather than a sensible request, I'd give numbers in a non-Latin character set, at which point it's easy: 一三五七九 (But even this only works for silly games, IMO).
- lynguist 2y ago4o: https://chatgpt.com/share/678ecbcd-a864-8007-9a34-e4e6574a2b57 https://chatgpt.com/share/678ecbcd-a864-8007-9a34-e4e6574a2b... "Thus, the five odd numbers without an 'e' in their spelling are: 41, 43, 45, 47, and 49. " o1: https://chatgpt.com/share/678ecbf8-ba20-8007-a33e-20ddbc545ba8 https://chatgpt.com/share/678ecbf8-ba20-8007-a33e-20ddbc545b... " 2 (two), 4 (four), 6 (six), 30 (thirty), 40 (forty) Those five are “odd” (i.e., strange) choices, yet each spelled-out form has no letter e. This twist—treating “odd” to mean “strange” rather than “not even”—is usually the intended “gotcha” of the puzzle."
- ben_w 2y ago> This twist—treating “odd” to mean “strange” rather than “not even”—is usually the intended “gotcha” of the puzzle." I like this one. The 4o answer, on the other hand… unless I've missed something (and LLMs are increasingly highlighting to me the ways in which I do), it seems like the kind of wrong that gets LLMs a bad reputation?
- lynguist 2y agoIt is! 4o is unfortunantely often very dumb in tricky circumstances, or is biased toward pundit-like opinions that are wrong. I'm not sure why that is the case, but the full o1 always has a "weight"/"presence" to it when I chat with it that suggests to me like a real intelligence. It can also solve difficult puzzles that 4o and me struggle with.
- d0mine 2y agoUno https://en.wiktionary.org/wiki/uno https://en.wiktionary.org/wiki/uno
- mchinen 2y ago[Edit: I suppose the correct answer was for the full model]. What I got was much worse when I ran with hf.co/unsloth/DeepSeek-R1-Distill-Llama-8B-GGUF:Q8_0. (I think the glyph-based reasoning is especially hard for models though, since they have a token-level blindness to them): >>> Give me 5 odd numbers that don't have the letter 'e' in them. <think> Okay, so I need to come up with five odd numbers that don't have the letter 'e' in their name. Hmm, let's break this down step by step. First, what exactly is being asked here? The user wants numbers, specifically odd ones, and they can't contain the letter 'e'. So I guess it's about the written form of the numbers, not their mathematical properties. For example, 13 is an odd number, and when you write it out, it's "thirteen," which does have an 'e' in it. So that wouldn't work. Alright, so I need to think of numbers where their names don't include 'e'. Let me start by listing some odd numbers and then check if each has an 'e' in their name. The first few odd numbers are 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, etc. Let's go one by one: 1. "One": No 'e' here. So 1 is a candidate. 2. "Three": Yes, that has an 'e'. 3. "Five": Again, no 'e'. So 5 works. 4. "Seven": Has an 'e' as well. 5. "Nine": No 'e' in "nine". Good to go. Wait, so far I have 1, 5, and 9 as possible numbers. Let me keep going. <truncated by me...> Wait, let me make sure. The word "forty" is F-O-R-T-Y, no 'e', so adding "-one" which doesn't add an 'e' either. So yes, "forty-one" has no 'e's. So that's another number: 41. So now I have five numbers: 1, 5, 9, 19, and 41. Let me double-check each to ensure they don't contain the letter 'e': - One: No 'e'. - Five: No 'e'. - Nine: No 'e'. - Nineteen: No 'e'. - Forty-one: No 'e'. Yes, all of these are odd numbers and their names do not include the letter 'e'. I think that's a solid list. </think> Here is a list of five odd numbers whose names do not contain the letter 'e': 1. **One** 2. **Five** 3. **Nine** 4. **Nineteen** 5. **Forty-one** These numbers are all odd, and their written forms do not include the letter 'e'.
- salviati 2y agoI would argue anything requiring insights on spelling is a hard problem for an LLM: they use tokens, not letters. Your point still stands, but you need different examples IMO.
- diggan 2y ago> Can it solve easy problems yet? Weirdly, I think that's an important milestone. Easy for who? Some problems are better solved in one way compared to another. In the case of counting letters and such, it is not a easy problem, because of how the LLM tokenizes their input/outputs. On the other hand, it's really simple problem for any programming/scripting language, or humans. And then you have problems like "5142352 * 51234" which is trivial problems for any basic calculator, but very hard for a human or a LLM. Or "problems" like "Make a list of all the cities that had celebrity from there who knows how to program in Fortan", would be a "easy" problem for a LLM, but pretty much a hard problem anything else than Wikidata, assuming both LLM/Wikidata have data about it in their datasets. > I suspect the breakthrough won't be trivial that enables solving trivial questions. So with what I wrote above in mind, LLMs already solve trivial problems, assuming you think about the capabilities of the LLM. Of course, if you meant "trivial for humans", I'll expect the answer to always remain "No", because things like "Standing up" is trivial for humans, but it'll never be trivial for a LLM, it doesn't have any legs!
- cchance 2y agoNot gonna lie ... wasnt expecting a correct answer... The thought process and confirmation of the calculation were LONG and actually quite amazing to watch it deduce and then calculate in different ways to confirm The product of 5,142,352 and 51,234 is calculated as follows: 1. Break down the multiplication using the distributive property: - (5,142,352 times 51,234 = (5,000,000 + 142,352) times (50,000 + 1,234)) 2. Expand and compute each part: - (5,000,000 times 50,000 = 250,000,000,000) - (5,000,000 times 1,234 = 6,170,000,000) - (142,352 times 50,000 = 7,117,600,000) - (142,352 times 1,234 = 175,662,368) 3. Sum all parts: - (250,000,000,000 + 6,170,000,000 = 256,170,000,000) - (256,170,000,000 + 7,117,600,000 = 263,287,600,000) - (263,287,600,000 + 175,662,368 = 263,463,262,368) Final Answer: 263463262368
- michaelt 2y ago> Easy for who? Consider things from a different angle. The hype men promoting the latest LLMs say the newest models produce PhD-level performance across a broad suite of benchmarks; some have even claimed that ChatGPT 4 is an early version of an AGI system that could become super-intelligent. So the advertising teams have set the bar very high indeed. As smart as the smartest humans around, maybe smarter. The bar they have set for themselves doesn't allow for any "oh but the tokenisation" excuses.
- danielmarkbruce 2y agoThere is no breakthrough required, it's trivial. It's just that by making a model do that, you'll screw it up on several other dimensions. Asking a question like this only highlights the questioners complete lack of understanding of LLMs rather than an LLMs inability to do something.