15 ms·
Links to chat with models that released this week: Large 2 - https://chat.mistral.ai/chat https://chat.mistral.ai/chat Llama 3.1 405b - https://www.llama2.ai/
by tikkun 2y ago
Links to chat with models that released this week:
Large 2 - https://chat.mistral.ai/chat https://chat.mistral.ai/chat
Llama 3.1 405b - https://www.llama2.ai/ https://www.llama2.ai/
I just tested Mistral Large 2 and Llama 3.1 405b on 5 prompts from my Claude history.
I'd rank as:
1. Sonnet 3.5
2. Large 2 and Llama 405b (similar, no clear winner between the two)
If you're using Claude, stick with it.
My Claude wishlist:
1. Smarter (yes, it's the most intelligent, and yes, I wish it was far smarter still)
2. Longer context window (1M+)
3. Native audio input including tone understanding
4. Fewer refusals and less moralizing when refusing
5. Faster
6. More tokens in output
- drewnick 2y agoAll 3 models you ranked cannot get "how many r's are in strawberry?" correct. They all claim 2 r's unless you press them. With all the training data I'm surprised none of them fixed this yet.
- vorticalbox 2y agoI just tried llama 3.1 8 b this is its reply. According to multiple sources, including linguistic analysis and word breakdowns, there are 3 Rs in the word "strawberry".
- tikkun 2y agoWhen using a prompt that involves thinking first, all three get it correct. "Count how many rs are in the word strawberry. First, list each letter and indicate whether it's an r and tally as you go, and then give a count at the end." Llama 405b: correct Mistral Large 2: correct Claude 3.5 Sonnet: correct
- layer8 2y agoIt’s not impressive that one has to go to that length though.
- unshavedyak 2y agoImo it's impressive that any of this even remotely works. Especially when you consider all the hacks like tokenization that i'd assume add layers of obfuscation. There's definitely tons of weaknesses with LLMs for sure, but i continue to be impressed at what they do right - not upset at what they do wrong.
- Spivak 2y agoTo me it's just a limitation based on the world as seen by these models. They know there's a letter called 'r', they even know that some words start with 'r' or have r's in them, and they know what the spelling of some words is. But they've never actually seen one in as their world is made up entirely of tokens. The word 'red' isn't r-e-d but is instead like a pictogram to them. But they know the spelling of strawberry and can identify an 'r' when it's on its own and count those despite not being able to see the r's in the word itself.
- layer8 2y agoThe great-parent demonstrates that they are nevertheless capable of doing so, but not without special instructions. Your elaboration doesn’t explain why the special instructions are needed.
- emmelaich 2y agoI think it's more that the question is not unlike "is there a double r in strawberry?' or 'is the r in strawberry doubled?' Even some people will make this association, it's no surprise that LLMs do.
- asadm 2y agothis can be automated.
- grumbel 2y agoGPT4o already does that, for problems involving math it will write small Python programs to handle the calculations instead of doing it with the LLM itself.
- jedberg 2y agoThis reminds me of when I had to supervise outsourced developers. I wanted to say "build a function that does X and returns Y". But instead I had to say "build a function that takes these inputs, loops over them and does A or B based on condition C, and then return Y by applying Z transformation" At that point it was easier to do it myself.
- mratsim 2y agoExact instruction challenge https://www.youtube.com/watch?v=cDA3_5982h8 https://www.youtube.com/watch?v=cDA3_5982h8
- HPsquared 2y ago"What programming computers is really like." EDIT: Although perhaps it's even more important when dealing with humans and contracts. Someone could deliberately interpret the words in a way that's to their advantage.
- hansworst 2y agoCan’t you just instruct your llm of choice to transform your prompts like this for you? Basically feed it with a bunch of heuristics that will help it better understand the thing you tell it. Maybe the various chat interfaces already do this behind the scenes?
- tcgv 2y agoChain-of-Thought (CoT) prompting to the rescue! We should always put some effort into prompt engineering before dismissing the potential of generative AI.
- johntb86 2y agoBy this point, instruction tuning should include tuning the model to use chain of thought in the appropriate circumstances.
- IncreasePosts 2y agoWhy doesn't the model prompt engineer itself?
- tcgv 2y agoBecause it is a challenging task, you would need to define a prompt (or a set of prompts) that can precisely generate chain-of-thought prompts for the various generic problems the model encounters. And sometimes CoT may not be the best approach. Depending on the problem other prompt engineering techniques will perform better.
- pegasus 2y agoAppending "Think step-by-step" is enough to fix it for both Sonnet and LLama 3.1 70B. For example, the latter model answered with: To count the number of Rs in the word "strawberry", I'll break it down step by step: Start with the individual letters: S-T-R-A-W-B-E-R-R-Y Identify the letters that are "R": R (first one), R (second one), and R (third one) Count the total number of Rs: 1 + 1 + 1 = 3 There are 3 Rs in the word "strawberry".
- doctoboggan 2y agoDue to the fact that LLMs work on tokens and not characters, these sort of questions will always be hard for them.
- ChikkaChiChi 2y ago4o will get the answer right on the first go if you ask it "Search the Internet to determine how many R's are in strawberry?" which I find fascinating
- paulcole 2y agoI didn't even need to do that. 4o got it right straight away with just: "how many r's are in strawberry?" The funny thing is, I replied, "Are you sure?" and got back, "I apologize for the mistake. There are actually two 'r's in the word strawberry."
- jcheng 2y agoGPT-4o-mini consistently gives me this: > How many times does the letter “r” appear in the word “strawberry”? > The letter "r" appears 2 times in the word "strawberry." But also: > How many occurrences of the letter “r” appear in the word “strawberry”? > The word "strawberry" contains three occurrences of the letter "r."
- brandall10 2y agoNeither phrase is causing the LLM to evaluate the word itself, it just helps focus toward parts of the training data. Using more 'erudite' speech is a good technique to help focus an LLM on training data from folks with a higher education level. Using simpler speech opens up the floodgates more toward the general populous.
- ofrzeta 2y agoI kind of tried to replicate your experiment (in German where "Erdbeere" has 4 E) that went the same way. The interesting thing was that after I pointed out the error I couldn't get it to doubt the result again. It stuck to the correct answer that seemed kind of "reinforced". It was also interesting to observe how GPT (4o) even tried to prove/illustrate the result typographically by placing the same word four times and putting the respective letter in bold font (without being prompted to do that).
- 2y ago
- Kuinox 2y agoTokenization make it hard for it to count the letters, that's also why if you ask it to do maths, writing the number in letters will yield better results. for strawberry, it see it as [496, 675, 15717], which is str aw berry. If you insert characters to breaks the tokens down, it find the correct result: how many r's are in "s"t"r"a"w"b"e"r"r"y" ? > There are 3 'r's in "s"t"r"a"w"b"e"r"r"y".
- GenerWork 2y ago>If you insert characters to breaks the tokens down, it find the correct result: how many r's are in "s"t"r"a"w"b"e"r"r"y" ? The issue is that humans don't talk like this. I don't ask someone how many r's there are in strawberry by spelling out strawberry, I just say the word.
- soneca 2y agoThis is only an issue if you send commands to a LLM as you were communicating to a human.
- antisthenes 2y ago> This is only an issue if you send commands to a LLM as you were communicating to a human. Yes, it's an issue. We want the convenience of sending human-legible commands to LLMs and getting back human-readable responses. That's the entire value proposition lol.
- pegasus 2y agoFar from the entire value proposition. Chatbots are just one use of LLMs, and not the most useful one at that. But sure, the one "the public" is most aware of. As opposed to "the hackers" that are supposed to frequent this forum. LOL
- bhelkey 2y agoIt's not a human. I imagine if you have a use case where counting characters is critical, it would be trivial to programmatically transform prompts into lists of letters. A token is roughly four letters [1], so, among other probable regressions, this would significantly reduce the effective context window. [1] https://help.openai.com/en/articles/4936856-what-are-tokens-and-how-to-count-them https://help.openai.com/en/articles/4936856-what-are-tokens-...
- Tepix 2y agoLLMs think in tokens, not letters. It's like asking someone who is dyslexic about spelling. Not their strong suit. In practice, it doesn't matter much, does it?
- recursive 2y agoSometimes it does, sometimes it doesn't. It is evidence that LLMs aren't appropriate for everything, and that there could exist something that works better for some tasks.
- Zambyte 2y agoLanguage models are best treated like consciousness. Our consciousness does a lot less than people like to attribute to it. It is mostly a function of introspection and making connections, rather than being the part of the brain where higher level reasoning and the functions of the brain that tell your body how to stay alive (like beating your heart). By allowing a language model to do function calling, you are essentially allowing it to do specialized "subconscious" thought. The language model becomes a natural language interface to the capabilities of its "subconsciousness". A specific human analogy could be: I tell you to pick up a pen off of the table, and then you do it. Most of your mental activity would be subconscious, orienting your arm and hand properly to pick up the pen, actually grabbing the pen, and picking it up. The linguistic representation of the action would exist in your concious mind (pick up the pen), but not much else. A language model could very easily call out to a text processing function to correctly do things like count the number of r's in the word strawberry. That is a job that your concious mind can dispatch to your subconciousness.
- joshstrange 2y agoLots of replies mention tokens as the root cause and I’m not well versed in this stuff at the low level but to me the answer is simple: When this question is asked (from what the models trained on) the question is NOT “count the number of times r appears in the word strawberry” but instead (effectively) “I’ve written ‘strawbe’, now how many r’s are in strawberry again? Is it 1 or 2?”. I think most humans would probably answer “there are 2” if we saw someone was writing and they asked that question, even without seeing what they have written down. Especially if someone said “does strawberry have 1 or 2 r’s in it?”. You could be a jerk and say “it actually has 3” or answer the question they are actually asking. It’s an answer that is _technically_ incorrect but the answer people want in reality.
- Der_Einzige 2y agoI wrote and published a paper at COLING 2022 on why LLMs in general won't solve this without either 1. radically increasing vocab size, 2. rethinking how tokenizers are done, or 3. forcing it with constraints: https://aclanthology.org/2022.cai-1.2/ https://aclanthology.org/2022.cai-1.2/
- generalizations 2y agoTesting models on their tokenization has always struck me as kinda odd. Like, that has nothing to do with their intelligence.
- b3nny 2y ago[dead]
- swatcoder 2y agoSurfacing and underscoring obvious failure cases for general "helpful chatbot" use is always going to be valuable because it highlights how the "helpful chatbot" product is not really intuitively robust. Meanwhile, it helps make sure engineers and product designers who want to build a more targeted product around LLM technology know that it's not suited to tasks that may trigger those kinds of failures. This may be obvious to you as an engaged enthusiast or cutting edge engineer or whatever you are, but it's always going to be new information to somebody as the field grows.
- wruza 2y agoIt doesn’t test “on tokenization” though. What happens when an answer is generated is few abstraction levels deeper than tokens. A “thinking” “slice” of an llm is completely unaware of tokens as an immediate part of its reasoning. The question just shows lack of systemic knowledge about strawberry as a word (which isn’t surprising, tbh).
- qeternity 2y agoIt is. Strawberry is one token in many tokenziers. The model doesn't have a concept that there are letters there.
- guywhocodes 2y agoThis is pretty much equivalent to the statement "multicharacter tokens are a dead end for understanding text". Which I agree with.
- Stumbling 2y agoClaude 3 Opus gave correct answer.
- taf2 2y agosonate 3.5 thinks 2
- stitched2gethr 2y agoInterestingly enough much simpler models can write an accurate function to give you the answer. I think it will be a while before we get there. An LLM can lookup knowledge but can't actually perform calculations itself, without some external processor.
- stanleydrew 2y agoWhy do we have to "get there?" Humans use calculators all the time, so why not have every LLM hooked up to a calculator or code interpreter as a tool to use in these exact situations?
- medmunds 2y agoHow much do threads like this provide the training data to convince future generations that—despite all appearances to the contrary—strawberry is in fact spelled with only two R's? I just researched "how many r's are in strawberry?" in a search engine, and based solely on the results it found, I would have to conclude there is substantial disagreement on whether the correct answer is two or three.
- fluoridation 2y agoSpeaking as a 100% human, my vote goes to the compromise position that "strawberry" has in fact four Rs.
- eschneider 2y agoThe models are text generators. They don't "understand" the question.
- m2024 2y agoDoes anyone have input on the feasibility of running an LLM locally and providing an interface to some language runtime and storage space, possibly via a virtual machine or container? No idea if there's any sense to this, but an LLM could be instructed to formulate and continually test mathematical assumptions by writing / running code and fine-tuning accordingly.
- killthebuddha 2y agoFWIW this (approximately) is what everybody (approximately) is trying to do.
- stanleydrew 2y agoYes, we are doing this at Riza[0] (via WASM). I'd love to have folks try our downloadable CLI which wraps isolated Python/JS runtimes (also Ruby/PHP but LLMs don't seem to write those very well). Shoot me an email[1] or say hi in Discord[1]. [0]:https://riza.io https://riza.io [1]:mailto:andrew@riza.io [2]:https://discord.gg/4P6PUeJFW5 https://discord.gg/4P6PUeJFW5
- mirekrusin 2y agoHow many "r"s are in [496, 675, 15717]?
- stanleydrew 2y agoPlug in a code interpreter as a tool and the model will write Python or JavaScript to solve this and get it right 100% of the time. (Full disclosure: I work on a product called Riza that you can use as a code interpreter tool for LLMs)
- kremi 2y agoYour question sounds like you want to know how the word is spelled, and no one would put two r’s at straw, so the model could be assuming that you’re asking whether it’s strawbery or strawberry. What happens if you ask the total number of occurrences of the letter r in the word? Does it still not get it right?
- exebook 2y agoTokenization can be bypassed like this; (mistral large 2) How many letters R are in the word "s-t-r-a-w-b-e-r-r-y"? The word "s-t-r-a-w-b-e-r-r-y" contains three instances of the letter "R." How many letters R contain the word strawberry? The word "strawberry" contains two instances of the letter "R."
- jbritton 2y agoI just tried Copilot 2 ‘r’ in strawberry, 3 ‘r’ in strawberries.
- groby_b 2y agoI mean, yes? Letter-based gymnastics are difficult in a token-based environment, and they're not a particularly valuable use case (outside of online dunking), so it's not really a surprise nobody cares much about those in terms of model training.
- ipython 2y agoI really wanted to believe this could be simply fixed by more prompt engineering and chain of thought, so I tried it with Claude 3.5 and … well unfortunately still no luck. the reply is amusing - I’ll include it verbatim below. Of note, this happened because earlier I had asked it without CoT and it incorrectly stated two “r”s. Once I started a new conversation (so fresh context window), it gave the right answer (3) and no longer claimed the “still the second r” as shown below. Let's think through this step-by-step: 1. First, let's spell out "strawberry" letter by letter: s-t-r-a-w-b-e-r-r-y 2. Now, let's count the "r"s as we go through each letter: s (0) t (0) r (1) - First "r" found a (1) w (1) b (1) e (1) r (2) - Second "r" found r (2) - This is still the second "r" y (2) 3. After going through all the letters, we've counted 2 "r"s. Therefore, the word "strawberry" contains 2 "r"s.
- ryanpanwar 2y agoWorks for me with Claude 3.5: Me: How many Rs are in the word strawberry? Think step by step Claude: Let's approach this step-by-step: 1. First, let's write out the word: strawberry 2. Now, let's go through each letter one by one: s - not an R t - not an R r - this is an R (count: 1) a - not an R w - not an R b - not an R e - not an R r - this is another R (count: 2) r - this is a third R (count: 3) y - not an R 3. After going through all the letters, we've counted 3 Rs. Therefore, there are 3 Rs in the word "strawberry".
- ipython 2y agoYeah, I figured out that it didn’t work at first because I still had the previous attempt (where it got it wrong) in my conversation history. Starting with a fresh conversation gave me the correct answer. It was still funny to see it “rationalize” the wrong answer tho.
- takumif 2y agoFor these classes of problems that LLMs struggle with, a more reliable way to go about them seems to be to ask them to solve them using tools, e.g. writing and executing a Python script to count the "R"s.
- Terr_ 2y agoI'm not surprised, because it's an issue with the fundamental design of the "pick words that tend to flow after the other words" machine. Training data will only "fix" it in the shallow sense that it will have seen a comment like yours before. (As opposed to the deeper sense of "learning to count.")
- 0x1ceb00da 2y ago> how many r's are in strawberry How many thoughts go through your brain when you read this comment? You can give me a number but it will be a guess at best.
- sashank_1509 2y agoWhile strawberry can be attributed to tokenization here are some other basic stuff I’ve seen language models fail at: 1. Play tic tac toe such that you never lose 2. Which is bigger 9.11 or 9.9 3. 4 digit multiplication even with CoT prompting
- rkwz 2y ago> Longer context window (1M+) What's your use case for this? Uploading multiple documents/books?
- tikkun 2y agoCorrect
- freediver 2y agoThat would make each API call cost at least $3 ($3 is price per million input tokens). And if you have a 10 message interaction you are looking at $30+ for the interaction. Is that what you would expect?
- rkwz 2y agoMaybe they're summarizing/processing the documents in a specific format instead of chatting? If they needed chat, might be easier to build using RAG?
- tr4656 2y agoThis might be when it's better to not use the API and just pay for the flat-rate subscription.
- coder543 2y agoGemini 1.5 Pro charges $0.35/million tokens up to the first million tokens or $0.70/million tokens for prompts longer than one million tokens, and it supports a multi-million token context window. Substantially cheaper than $3/million, but I guess Anthropic’s prices are higher.
- msp26 2y agoLarge 2 is significantly smaller at 123B so it being comparable to llama 3 405B would be crazy.
- qwertox 2y agoClaude needs to fix their text input box. It tries to be so advanced that code in backticks gets reformatted, and when you copy it, the formatting is lost (even the backticks).
- nickthesick 2y agoThey are using Tiptap for their input and just a couple of days ago we called them out on some perf improvements that could be had in their editor: https://news.ycombinator.com/item?id=41036078 https://news.ycombinator.com/item?id=41036078 I am curious what you mean by the formatting is lost though?
- qwertox 2y agoOdd, multiline backtick code works very good, I don't know why I thought that it was also broken. When you type "test `foo` done" in the editor, it immediately changes `foo` into a wrapped element. When you then copy the text without submitting it, then the backticks are lost, losing the inline-code formatting. I thought that this could also happen to multiline code. Somehow it does. Type the following: Test: ``` def foo(): return bar ``` Delete that and type Test: ``` def foo(): return bar ``` done In the first case, the ``` in the line "Test: ```" does not open the code block, this happens with the second backtics. Maybe that's the way markdown works. In the second case, all behaves normally, until you try to copy what you just wrote into the clipboard. Then you end up with Test: def foo(): return bar done Ok, only the backticks are lost but the formatting is preserved. I think I have been trained by OpenAI to always copy what I submit before submitting, because it sometimes loses the submitted content, forcing me to re-submit.
- cpursley 2y agoClaude is truly incredible but I'm so tired of the JavaScript bloat everywhere. Just why. Both theirs and ChatGPTs UIs are hot garbage when it comes to performance (I constantly have to clear my cache and have even relegated them to a different browser entirely). Not everyone has an M4, and if we did - we'd probably just run our own models.
- campers 2y agoIt seems to be the way with these releases, sticking with Claude, at least for the 'hard' tasks. In my agent platform I have LLMs assigned for easy/medium/hard categorised tasks, which was somewhat inspired from the Claude 3 release with Haiku/Sonnet/Opus. GPT4-mini has bumped Haiku for the easy category for now. Sonnet 3.5 bumped Opus for the hard category, so I could possibly downgrade the medium tasks from Sonnet 3.5 to Mistral Large 2 if the price is right on the platforms with only 123b params compared to 405b. I was surprised how much Llama3 405b was on together.ai $5/mil for input/output! I'll stick to Sonnet 3.5. Then I was also surprised how much cheaper Fireworks was at $3/mil Gemini has two aces up its sleeve now with the long context and now the context caching for 75% reduced input token cost. I was looking at the "Improved Facuality and Reasoning in Language Models through Multi-agent debate" paper the other days, and thought Gemini would have a big cost advantage implementing this technique with the context caching. If only Google could get their model up to the level of Anthropic.
- scosman 2y ago> Native audio input including tone understanding I can't seem to find docs on this. Have a link?
- GaggiX 2y agoIt's a wishlist of the parent comment.
- 0x1ceb00da 2y ago> Native audio input including tone understanding Is there any other LLM that can do this? Even chatgpt voice chat is a speech to text program that feeds the text into the llm.
- jacooper 2y agoDoes Claude support plug-ins like GPTs? Chatgpt with Wolfram alpha is amazing, it doesn't look like Claude has anything like it.