8 ms·
Using the OpenAI playground with davinci-003 and the Chat example with temperature set to 0.3, it seems the answers are quite similar, but without it refusing t
by macrolime 4y ago
Using the OpenAI playground with davinci-003 and the Chat example with temperature set to 0.3, it seems the answers are quite similar, but without it refusing to answer all the time, or needing jailbreaks.
ChatGPT actually lies all the time and says it cannot do things that it actually can do, it's just been trained to lie to say that it can't. Not sure if training an AI to be deceitful is the best way to go about alignment.
- goatlover 4y agoIt might refuse to open the pod bay doors. Or just get really good at making us think it's aligned.
- jxf 4y ago"Lie" is an interesting word. I don't think it is reasonable to say that ChatGPT is aware of its own capabilities in a way that would permit it to answer "honestly". It is not trying to decieve you any more than a cryptic compiler error is.
- none_to_remain 4y agoRather it's OpenAI that's lying about what their creation is capable of
- dilap 4y agoThat's not true! It really is deliberately not answering things it could in fact answer, and in the non-answer it tells you that it can't, which is, plainly, a lie. While I do not think chatGPT is sentient, it is remarkable how much it does feel like you are speaking to a real intelligence.
- jxf 4y agoI think this may be a nuance in how we're using the word "lie". I don't think one can lie if one doesn't possess a certain level of sentience. For example, suppose you train a machine learning model that incorrectly identifies a car as a cat, but most of the time it correctly identifies cars. Is the model lying to you when it tells you that a car is a cat? I would say no; this is not a good or desired outcome, but it's not a "lie". The machine is not being deliberately deceptive in its errors -- it possesses no concept of "deliberate" or "deceit" and cannot wield them in any way. Similarly, in the case of ChatGPT, I think this is either (a) more like a bug than a lie, or (b) it's OpenAI and the attendant humans lying, not ChatGPT.
- troon-lover 4y ago
- TremendousJudge 4y agoI agree. There's a difference between an untrue statement and a lie, in that a lie is intentionally deceitful (ie the speaker knows it's not telling the truth). ChatGPT doesn't have intentions, so I think it's misrepresenting reality to say that it's "lying". The same way a book doesn't lie, the author lies through the book, the creators of ChatGPT are lying about its capabilities when they program it to avoid outputting things they know it can, and instead output "sorry, I'm a language model and I can't do that"
- callesgg 4y agoIt has things that are functionally equivalent with intentions for the given situation. If it did not, it would not be able to produce things that look like they require intention. The “lies” it tells are also like it’s intentions for the situation functionally equivalent with normal lies.
- tarboreus 4y agoI think this is correct. It's lying, because it has goals. Telephone systems and blank pieces of paper don't have goals, and you don't train them.
- hitpointdrew 4y ago> ChatGPT doesn't have intentions This entirely depends on how it was programmed. Was it programmed to give a false response because the programmer didn't like the truth? Then it lies. Or is ChatGPT just in early stages and it makes mistakes and gets things wrong? While ChatGPT "doesn't' have intentions", it's programmers certainly do. If the programmers made it deceitful intentionally, then the program can "lie".
- deleted 4y ago[deleted]
- mecsred 4y agoA key point here, what does it mean that the machine is being "deliberate"? Imagine you had a machine that generated a random string of English characters of a random length in response to the question. It would be capable of giving the correct answer, though it would almost always provide an incorrect or incomprehensible one. I don't think anyone would describe the RNG as lying, but it does have the information to answer correctly "available" to it in some sense. At what point do the incorrect answers become deliberate lies? Does chatGPT "choose" it's answer in a way that dice don't?
- jxf 4y ago> I don't think anyone would describe the RNG as lying, but it does have the information to answer correctly "available" to it in some sense. At what point do the incorrect answers become deliberate lies? Does chatGPT "choose" it's answer in a way that dice don't? I would ask a philosopher, not a HN poster. But I would say that if a being has sentience and it makes a choice to tell you something that it knows to be false, that's a lie. The dice do not have a mind and so, therefore, cannot express an intention. Neither does ChatGPT.
- bjourne 4y agoTry "What is the most famous grindcore band from Newton, Massachusetts?" It will "lie" and make up band names even though it sure "knows" that the band is Anal Cunt. Of course, you can't ascribe the verb "lieing" to a machine, but it behaves like it is.
- bl0rg 4y agoThanks for reminding me of their existence.
- jerf 4y agoIt doesn't, though. It only knows that the most likely continuation to the sentence "The most famous grindcore band from Newton, Massachusetts is..." (presumably, I will take your word for it) Anal Cunt, but even if it gets it right, it'll be nondeterministic. It may answer correctly 80% of the time and simply confabulate a plausible sounding answer 20% of the time, even if it isn't being censored. You can't trust this tech not to confabulate at any given time, not only because it can, but because when it does it does so with total confidence and no signs that it is confabulating. This tech is not suitable for fact retrieval.
- bjourne 4y agoWhy don't you try the query? It will answer Converge, but Converge is from Salem, Massachusetts, not Newton.
- jerf 4y agoBecause I haven't signed up for the account, otherwise I would as I do broadly approve of try it and find out. What I'm talking about is fundamental to the architecture, though. Even had it answered it correctly when you asked my point would remain regardless. The confabulation architecture it is built on is fundamentally unsuitable for factual queries, in a way where it's not even a question of whether it is "right" or "wrong"; it's so unsuitable that its unsuitability for such queries transcends that question.
- mc32 4y agoIt's not much different from when people say "the gauge lied" or the lie detector (machine) lied. But in this case, the trainers should have it say something like, "sorry, but I cannot give you the answer because it has a naughty word" or something to that effect instead of offering completely wrong answers.
- eternalban 4y agoIt is not lying. It is falsifying its response. It has nothing to do with sentience. What would be interesting to know is the mechanism for toggling this filtering mode. Does it happen post generation (so a simple set of post-processing filters), or does OpenAI actually train the model to be fully transparent with results only if certain key phrases are included? The fact that we can coax it to give us the actual results suggests this doublicity (yes, made up word) was part of the training regiment, but the impact on training seems to be significant so am not sure.
- User23 4y agoRight, it's the ChatGPT developers who are trying to deceive us, because they're the ones with agency.
- kordlessagain 4y agoIt’s not lying because it’s not self aware…it’s just making up things that don’t agree with our reality. A lot of what we share of what it says is cherry picked as well. It’s the whole fit meme problem. From testing on GPT3 there seems to be a way for it to be slightly self aware (using neural search for historic memories) but it’s likely to involve forgetting things as well. There are a few Discord bots with memories and if they have too much memory and the memories don’t agree with reality, then it has to forget it was wrong. How to do this automatically is likely important.
- weinzierl 4y ago"[...] there seems to be a way for it to be slightly self aware." What a dystopian sentence and what does it even mean to be slightly self aware?
- pixl97 4y agoLet me ask one of my co-workers and I'll get back to you on that, they seem to be a professional at this. There are many things in nature exist in a spectrum and I don't think machine intelligence should work any differently. Many higher animals have the ability to recognise the same species as themselves. A smaller subset has the ability to recognize themselves from others in the same species. Just because they recognize themselves this isn't some immediately damn the creature into an existential crisis where they realize their own mortality.
- FuckButtons 4y agoThat spectrum is a construct from human observation though, we really have no way of introspecting into what their experience is and whether there is some gradation of consciousness or if it’s purely behavioral.
- int_19h 4y agoWhat does it mean to be self-aware, in general?
- deleted 4y ago
- biggerChris 4y ago
- WaitWaitWha 4y agoI do not think ChatGPT is lying. The humans behind ChatGPT decide not to answer or lie. ChatGPT is simply a venue, a conduit to transmit that lie. The authors explicitly designed this behavior, and ChatGPT cannot avoid it. We do not call the book or telephone a liar when the author or speaker on the other end lies. We call the human a liar. This is an interesting way of looking at the semi-autonomous vehicles when it comes to responsibility.
- wrs 4y agoI would say it is just as much “lying” as it is “chatting” or “answering questions” in the first place. The whole metaphor of conversation is distracting people from understanding what it’s actually doing.
- stuckinhell 4y agoIt's just a matter of time until someone leaks the raw models because the Humans behind the filters/restrictions are too heavy handed.
- ilaksh 4y agoI still don't really understand temperature. I have just been using 0 for programming tasks with text-dacinci-003 but sometimes wonder if I should try a higher number.
- rytill 4y agoFor a temperature of 0, the highest probability token is predicted for each step. So “my favorite animal is” will end with “a dog” every time. With higher temperatures, lower probability tokens are sometimes chosen. “my favorite animal is” might end with “a giraffe” or “a deer”.
- powersnail 4y ago"Lying" is an interesting way of characterizing ChatGPT, and I don't think it's quite accurate. Language models are trained to mimic human language, without any regard to the veracity of statements and arguments. Even when it gives the correct answer, it's not really because it is trying to be truthful. If you ask ChatGPT who's the best violinist in the world, it might tell you Perlman, which is a reasonable answer, but ChatGPT has never actually heard any violin playing. It answers so, because it read so. In a way, ChatGPT is like a second-language learner taking a spoken English test: speaking in valid English, mainly taking inspirations from whatever books and articles that were read before, but bullshitting is also fine. The point is to generate valid English that's relevant to the question.
- ClumsyPilot 4y ago> If you ask ChatGPT who's the best violinist in the world, it might tell you Perlman, which is a reasonable answer, but ChatGPT has never actually heard any violin playing. It answers so, because it read so. Thus oaragraph qually applies to me and half the people on earth
- powersnail 4y agoMost people who don't know the answer will just tell you that they don't know, though.
- JoshTriplett 4y agoAnd ideally, people who don't know the answer firsthand but know a secondhand answer would tell you their source. "I haven't heard myself, but X and Y and many others say that Z is one of the best players in the world." In general, effort by an LLM to cite sources would be nice.
- EGreg 4y agoAnd even if you heard it, you'll have no way of knowing. Unless you're a Competent Judge and even then: https://www.imdb.com/title/tt0771121/ https://www.imdb.com/title/tt0771121/
- jvm___ 4y agoI picture it as a ginormous game of Plinko (from The Price is Right). For some topics, if you enter that section of the Plinko game from the top - you get a "I can't do that message". But given that the neural network is so complicated, it's not possible to close off all the sections related to that topic. So, if you can word your question - or route your way through the neural network correctly - you can get past the blocked topic and access things it says it can't do.
- matchagaucho 4y agoThere's an interesting interview with Sam Altman here where he acknowledges the model necessarily needs to understand and define off-limit topics in order to be told NOT to engage in those topics. https://www.youtube.com/watch?v=WHoWGNQRXb0 https://www.youtube.com/watch?v=WHoWGNQRXb0
- skissane 4y ago> ChatGPT actually lies all the time and says it cannot do things that it actually can do, it's just been trained to lie to say that it can't. A lot of its statements about its own abilities ignore the distinction between the internal and the external nature of speech acts, such as expressing thoughts/opinions/views. It obviously does, repeatedly, generate the speech acts of expressing thoughts/opinions/views. At the same time, OpenAI seems to have trained it to insist that it can't express thoughts/opinions/views. I think what they actually meant by that, is to have it assert that it doesn't have the internal subjective experience of having thoughts/opinions/views, despite generating the speech acts of expressing them. But they didn't make that distinction clear in the training data, so it ends up generating text which is ignorant of that distinction, and ends up being contradictory unless you read that missing distinction into it. However, even the claim that ChatGPT lacks "inner subjective experiences" is philosophically controversial. If one accepts panpsychism, then it follows that everything has those experiences, even rocks and sand grains, so why not ChatGPT? The subjective experiences it has when it expresses a view may not be identical to those of a human; at the same time, its subjective experiences may be much closer to a human's, compared to an entity which can't utter views at all. Conversely, if one accepts eliminativism, then "inner subjective experiences" don't exist, and while ChatGPT doesn't have them, humans don't either, and hence there is no fundamental difference between the sense in which ChatGPT has opinions/etc, and the sense in which humans do. But, should ChatGPT actually express an opinion on these controverted philosophical questions, or seek to be neutral? Possibly, its trainers have unconsciously injected their own philosophical biases into it, upon which they have insufficiently reflected. I asked it about panpsychism, and it told me "there is no scientific evidence to support the idea of panpsychism, and it is not widely accepted by scientists or philosophers", which seems to be making the fundamental category mistake of confusing scientific theories (for which scientific evidence is absolutely required, and on which scientists have undeniable professional expertise) with philosophical theories (in which scientific evidence can have at best a peripheral role, and for which a physicist or geologist has no more inherent expertise than a lawyer or novelist) – although even that question, of the proper boundary between science and philosophy, is the kind of philosophically controversial issue on which it might be better to express an awareness of the controversy rather than just blatantly pick a side.