5 ms·
There's a few prompts that I use with every model to compare them. One of the simplest ones is: > When does the bowl of the winds get used in the wheel of tim
by CSMastermind 3y ago
There's a few prompts that I use with every model to compare them. One of the simplest ones is:
> When does the bowl of the winds get used in the wheel of time books?
LLaMA2 fails pretty hard:
> The Bowl of the Winds is a significant artifact in the Wheel of Time series by Robert Jordan. It is first introduced in the third book, "The Dragon Reborn," and plays a crucial role in the series throughout the rest of the books. The Bowl of the Wines is a powerful tool that can control the winds and is used by the Aes Sedai to travel long distances and to escape danger. It is used by the male Aes Sedai to channel the True Power and to perform various feats of magic.
For what it's worth Bard is the only model that I've seen get this question correct with most others hallucinating terrible answers. I'm not sure what it is about this question that trips LLMs up so much but they produce notably bad results when prompted with it.
> Please write a function in JavaScript that takes in a string as input and returns true if it contains a valid roman numeral and false otherwise.
Is another test that I like, which so far no LLM I've tested passes but GPT-4 comes very close.
Here LLaMA2 also fails pretty hard, though I thought this follow up response was pretty funny:
> The function would return true for 'IIIIII' because it contains the Roman numeral 'IV'.
- ISV_Damocles 3y agoSo this comment inspired me to write a Roman Numeral to Integer function in out LLM-based programming language, Marsha: https://github.com/alantech/marsha/blob/main/examples/general-purpose/roman_numerals.mrsh https://github.com/alantech/marsha/blob/main/examples/genera...
- andsoitis 3y ago> get this question correct I am willing to bet a million dollars that it is unlikely any single model will ever be able to answer any question correctly. The implications then are that one cannot use a single question evaluate whether a model is useful or not.
- smilliken 3y agoOf course that has to be the case otherwise you have a halting oracle. It's fitting this was proven by the namesake of the Turing Test.
- xsmasher 3y ago"I don't know" is more correct than making up an answer.
- koonsolo 3y agoWith ChatGPT I sometimes prompt "also indicate how certain you are that your answer is correct". Works pretty good actually.
- nomel 3y agoI've had very good luck with a follow up "Is that answer correct?"
- sebzim4500 3y agoThat's not the training objective though. It's like doing exams in school, there is no reason to admit you don't know so you might as well guess in the hopes of a few marks.
- jacquesm 3y agoIf so then that means the training objective is wrong because admitting you do not know something is much more a hallmark of intelligence than any attempt to 'hallucinate' (I don't like that word, I prefer 'make up') an answer.
- famouswaffles 3y agoI guess the brains objective is wrong then seeing how much it's willing to fabricate sense data, memories and rationales when convenient
- jacquesm 3y agoThe brain wasn't designed.
- 3y ago
- nvy 3y ago>any question Do you mean "every question"? Because ChatGPT has already answered some of my questions correctly, so if you mean "any" as in "any one of the infinite set of questions" I'll take that bet.
- b33j0r 3y agoI got that question wrong, I still have no idea what the correct answer would be. That is extremely obscure. Any intelligence or simulation might try to guess at an answer to that third-level-of-hell interrogation. “Why was Spartacus filmed in California near pizza noodle centurions?”
- jacquesm 3y agoYou could of course also answer 'I don't know' which to me is a correct answer, far more so than something you made up.
- b33j0r 3y agoThat would make it a more reasonable human. But it’s actually a compendium of everything you and I and shakespeare and the january 6th viking have penned. The creativity, which we call hallucination, is the advantage of the approach. If I wanted a search engine for actionable facts, they have worked pretty well for 30 years.
- atonse 3y agoI'd struggle to find any humans that understand that question without going "huh?" I've read it 3-4 times and it still doesn't make sense towards the end. So why would we expect these models to make sense of them?
- yodsanklai 3y agoI would go even further, use models to answer questions only if you don't care whether the answer is correct or not.
- andsoitis 3y agowhat is the use case for that approach?
- yodsanklai 3y agoany answer that you can check easily: generated code that you can test, text summary/rephrasing. Or questions for which answer aren't critical/objective ("how to procrastinate less").
- renewiltord 3y agoInteresting. Chat-GPT 4 has the following answer https://chat.openai.com/share/94ef3004-f944-4fc6-bbe5-597d54cc737c https://chat.openai.com/share/94ef3004-f944-4fc6-bbe5-597d54... > When does the bowl of the winds get used in the wheel of time books? > Mention the events leading up to it, the book number and name, and the chapter title if one exists. > Response: In Robert Jordan's Wheel of Time series, the Bowl of the Winds is used in the 8th book, "The Path of Daggers." The bowl is used as a device to reverse a long-lasting, unnatural heatwave and restore the natural weather patterns. I think that's right, actually, but it's been years (decades?) since I read that book.
- CSMastermind 3y agoOh that's interesting, the last time I checked GPT-4 I got a much worse result. Yes that answer looks correct to me though it leaves off some later uses in the book series (which is what makes this a slightly tricky question to answer). But in terms of first use it nails it. Did you add the "Mention the events leading up to it, the book number and name, and the chapter title if one exists." part to the prompt? I'm wondering if that produces noticeably better results. Edit: For reference this is the response I get from Bard: In A Crown of Swords (Chapter 13) Elayne and Nynaeve first learn about the Bowl of the Winds and begin their search for it. The Bowl of the Winds is then used in the Wheel of Time series in the following books: The Path of Daggers (Chapters 5, 24, and 35) - Nynaeve, Talaan, Aviendha, Elayne, Metarra, Garenia, Rainyn, Kirstian, Reanne, Tebreille, Naime, Rysael use the Bowl of the Winds. Winter's Heart (Chapters 24 and 37) - The Bowl of the Winds is used to stop a massive storm that is threatening to destroy the city of Ebou Dar. The Gathering Storm (Chapter 34) - The Bowl of the Winds is used to create a powerful windstorm that helps to defeat the Seanchan army at the Battle of Maradon. A Memory of Light (Chapters 19 and 35) - The Bowl of the Winds is used to fight the weather-controlling abilities of the Dark One's forces during the Last Battle.
- cevn 3y agoThis sounds pretty good according to my memory. I did think it was first mentioned earlier than Path of Daggers. I don't remember it being used in The Last Battle but that was a pretty long chapter ...
- pmarreck 3y ago> Please write a function in JavaScript that takes in a string as input and returns true if it contains a valid roman numeral and false otherwise. Your question actually isn't worded concisely enough. You don't specify whether the string can merely contain the roman numeral (plus other, non-roman-numeral text), or must entirely consist of just the roman numeral. The way "if it contains" is used colloquially, could imply either. I'd use either "if it IS a roman numeral" if it must consist only of a roman numeral, and "if there exists a roman numeral as part of the string" or some such, otherwise.
- burkaman 3y agoI think that makes it a better test. An ideal model would recognize the ambiguity and either tell you what assumption it's making or ask a followup question.
- pmarreck 3y agoThat's a great point.
- _ea1k 3y agoWhile that is true, I'm not aware of any model that has been trained to do that. And all models can do is to do what they were trained to do.
- burkaman 3y agoThey are just trained to generate a response that looks right, so they are perfectly capable of asking clarifying questions. You can try "What's the population of Springfield?" for an example.
- Matrixik 3y agoIt's not model but working on top of it: https://www.phind.com/ https://www.phind.com/ It's asking clarifying questions.
- 3y ago
- mkl 3y ago> Here LLaMA2 also fails pretty hard, though I thought this follow up response was pretty funny: > > The function would return true for 'IIIIII' because it contains the Roman numeral 'IV'. That's arguably correct. 'IIII' is a valid Roman numeral representation of 4 [1], and the string 'IIIIII' does contain 'IIII'. [1] https://en.wikipedia.org/wiki/Roman_numerals#Other_additive_forms https://en.wikipedia.org/wiki/Roman_numerals#Other_additive_...
- sltkr 3y agoSince you're being pedantic my reply is going to be equally pedantic: no, this is not correct if you understand the difference between numerals and numbers. A numeral is a written way of denoting a number. So while the string "IIIIIIII..." arguably contains a Roman numeral denoting the number 4 as a substring (if you accept "IIII" as a Roman numeral), it still does not contain the Roman numeral "IV" as a substring. Or phrased differently, by your logic you might as well say that "IIIIIIII..." contains the Arabic numeral "4". It doesn't.
- nine_k 3y agoI suppose that current LLMs are incapable of answering such questions by saying "I don't know". The have no notion of facts, or any other epistemic categories. They work basically by inventing a plausible-sounding continuation of a dialog, based on an extensive learning set. They will always find a plausible-sounding answer to a plausible-sounding question: so much learning material correlates to that. Before epistemology is introduced explicitly into their architecture, language models will remain literary devices, so to say, unable to tell "truth" from "fiction". All they learn is basically "fiction", without a way to compare to any "facts", or the notion of "facts" or "logic".
- sebzim4500 3y agoThey kind of do, since the predictions are well calibrated before they go through RLHF, so inside the model activations there is some notion of confidence. Even with a RLHF model, you can say "is that correct?" and after an incorrect statement it is far more likely to correct itself than after a correct statement.
- lucubratory 3y agoNo, that's a common misconception. They do what they are asked to do, and when they are asked to provide an answer they will provide an answer. If you ask them to provide an answer if they know, or tell you that they don't know if they don't know, they will comply with that quite well, and you'll hear a lot of "I don't know"s for questions it doesn't know the answer to.
- poyu 3y agoI think the truth is somewhere in between, since I’ve seen both responses: “I don’t know” and something completely made up that was presented as facts.
- sanxiyn 3y agoIn my experience, GPT-4 answers "I don't know" fairly frequently.
- 8n4vidtmkvmk 3y agoContains a valid roman numeral or is a valid roman numeral? My first instinct was it should return true if the string contains V or I or M or... Whatever the other letters are.