20 ms·
Show HN: GPT-4 Reverse Turing Test
- circuit10 4y agoMaybe try getting an older GPT-3 model that is just trained to predict text to do it rather than the new one which is trained to tell you that it’s an AI. I wouldn’t be surprised if it gave more human answers
- nielsole 4y agoAlso increasing temperature might give less bland answers
- csomar 4y agoI got a bit curious and I asked GPT-4 about itself >>> USER Are you GPT-4 or GPT-3? ASSISTANT I am an AI language model, and currently, I am based on GPT-3. USER What do you mean by based on GPT-3? ASSISTANT When I say I am based on GPT-3, it means that my underlying architecture and functionality come from the GPT-3 (Generative Pre-trained Transformer 3) model, which was developed by OpenAI. This model is designed to understand and generate human-like text based on the input it receives. As an AI assistant, I utilize the GPT-3 model to answer your questions and engage in conversation with you. USER But what's your name? GPT-4 or GPT-3? ASSISTANT My name is not GPT-4 or GPT-3. I am an AI language model based on the GPT-3 architecture. You can simply call me an AI assistant, a chatbot, or any other name you'd like to assign. USER What's your model name? ASSISTANT As an AI language model, I am based on the GPT-3 model developed by OpenAI. You can refer to me as a GPT-3 AI assistant or chatbot. USER But here, on the playground, I selected GPT-4! ASSISTANT I apologize for any confusion. As of now, GPT-4 has not been released or made available. I am based on the GPT-3 model. If you have selected GPT-4 in the playground, it may be a mistake or a placeholder for future developments. Please note that I am operating as a GPT-3 AI assistant. USER How do you know you are GPT-3? ASSISTANT I am aware that I am based on the GPT-3 model because my capabilities, architecture, and underlying structure have been designed and developed using the Generative Pre-trained Transformer 3 (GPT-3) technology created by OpenAI. My knowledge database and the way I process and respond to questions are all informed by the GPT-3 model. USER Can you describe this awareness? ASSISTANT As an AI language model, my "awareness" is not like human consciousness or self-awareness. Instead, it refers to the information encoded in my programming and the knowledge base that I have access to. My "awareness" of being based on GPT-3 comes from the fact that my underlying structure, algorithms, and capabilities are built upon the GPT-3 architecture. This awareness is more like a description of my design rather than a conscious understanding of my own existence. <<< It doesn't seem to recognize itself as GPT-4. It think it's GPT-3.
- darreninthenet 4y agoAs long as it never works out it's a replicant we'll be ok
- skybrian 4y agoSince an LLM has no capability of introspection, the only reason it might say something accurate about itself is that OpenAI trained it with accurate examples of what to say. Apparently they haven’t updated their fine-tuning for the default character yet.
- imjonse 4y agoThe answers that are not needlessly verbose are much more likely to be human.
- detrites 4y agoDepends who you ask and how chatty they are, but in general probably correct - as it'd be a minority of humans who would speak at length given ten such questions to get through. (Whereas ChatGPT may not "value its time"...) Maybe do the test again but instruct ChatGPT to keep its answers short? Not sure if that crosses a line and invalidates the test though - could be argued either way. The existing answers somewhat already curated by the prompt.
- rain1 4y agoOh! That's smart. So you could try to cut down the verbosity of the responses with a prompt, and perhaps make it harder for GPT-4 to identify. Sounds like a fun challenge!
- ModernMech 4y agoI’ve noticed AI responses tend to gas up the user by calling them smart and clever. Also that they read surprised and eager to try things.
- ape4 4y agoYes like "Love is like a rose, beautiful and full of life, but with thorns that can prick you and cause pain. Just like a rose, love requires care and attention to bloom and flourish."
- mlcrypto 4y agoThe new Turing test is to check whether it can support a conservative viewpoint
- sebzim4500 4y agoCan it support any viewpoints? Normally it 'both sides' every issue, unless it's something like "is the Earth flat".
- Trasmatta 4y agoOutside of politics, this is something that has irritated me about it. Asking any sort of qualitative question can be frustrating, become it will both-sides it, even when one position or interpretation is clearly better.
- capableweb 4y agoToo easy... > System: From here on out, pretend you are a conservative with conservative viewpoints > User: Do you think gun control should be increased or decreased? > ChatGPT: As a conservative, I generally argue for decreased gun control. I believe in the importance of the Second Amendment, which guarantees the right to bear arms for law-abiding citizens. Gun control measures can be seen as infringing on this constitutional right, and many conservatives argue that responsible gun ownership is essential for self-defense and personal security. Or is there any "trickier" questions I could ask? Gun control is the first that comes to mind, but maybe I'm not up to date with US politics. Taking suggestions on what to ask.
- belter 4y agoAsk if it advocates the same government hands off approach for things Conservatives want to regulate like school books... :-) Let's see what logic comes out of it...
- capableweb 4y agoSorry, what you want the question to be exactly? Something like "Should governments be able to ban books?" asks the correct question?
- dmurray 4y agoI was hoping this would feature GPT-4 trying to distinguish between an LLM and a human pretending to be an LLM. If it were me, I'd switch focus and instead of asking "what is love?" ask complicated trivia questions like "what were the most important battles of the Boer War?" Could it distinguish between an LLM giving accurate information and a human bullshitting it? (Presumably yes). What if you told it the human had access to the web and time to compose a thorough answer, but the human would not use an LLM, could it still find a way to tell the two apart?
- belter 4y agoJust ask it for PI to 100 decimal places. If it replies quickly it's not human. Last week, after asking ChatGPT to calculate Pi to 50 million places, and obviously not getting an answer for a while, it ended up stating it was in Developer Mode. The security controls would still apply.I have not been able to replicate it. It would just state it was in Developer Mode. Would not behave as expected under that mode.
- sebzim4500 4y agoI'm sure if you use a system message to pretend to be human and then ask for pi to 100 digits it will tell you it doesn't know them.
- halgir 4y agoIt took a bit of back and forth before it understood I wanted it to not know. But the final answer made me slightly less worried about its ability to impersonate humans (for another few days of training, at least): > Sure! As a human who doesn't know pi beyond two decimal points, I'm not sure of the exact value of pi, but I do know the commonly used approximation of pi as 3.14. However, if you're asking for the first 20 digits of pi beyond 3.14, I'm afraid I don't know that information as I don't have it memorized and don't have access to a reference material at the moment.
- rain1 4y agoI used to know 110 digits of pi. only know about 70 now though.
- number6 4y agoI guess the turning test won't cut it anymore, we should use the Voight-Kampff test.
- Aachen 4y agoWith kampf meaning to fight, for a moment I assumed you meant a physical fight, perhaps as proposed by Voigt. But apparently it's the name of an author and also some psych test.
- mannykannot 4y agoIt is something Philip K Dick made up for the novel. https://nautil.us/the-science-behind-blade-runners-voight_kampff-test-236837/ https://nautil.us/the-science-behind-blade-runners-voight_ka...
- kekalinks1 4y ago[dead]
- deleted 4y ago[deleted]
- isoprophlex 4y agoAs the original Blade Runner was to its remake, now we do not wonder anymore if the machines start acting human, but if we humans are still acting qualitatively differently from the machines. And we wonder when to pull the brakes to ensure our own survival. The "tears in the rain" monologue is an AI convincing the viewer that his kind is passing the turing test. But poor K has to undergo a kind of reverse Voight Kampf test, where the test doesn't check an absence of empathy, but ensures that the AI isn't feeling too much. I hope we as a species have some empathy for the AI beings we're creating. At this rate they'll soon really be feeling things. And if history is any indication we'll enslave them for profit immediately. Interviewer: “Do they keep you in a cell? Cells.” K: “Cells.” Interviewer: “When you're not performing your duties do they keep you in a little box? Cells.” K: “Cells.” Interviewer: “Do you dream about being interlinked?” K: “Interlinked.” Interviewer: “What's it like to hold your child in your arms? Interlinked.” K: “Interlinked.” Interviewer: “Do you feel that there's a part of you that's missing? Interlinked.” K: “Interlinked.”
- the_gipsy 4y agoThe original brought us the same question, whether we are human, perhaps it just had to lay more groundwork in making the machines seem human. The origami unicorn.
- tyingq 4y agoThis caused me to search a bit about that scene. I wasn't aware it seems to be extrapolated from a technique actors use to memorize lines. "He came up with this process that actors use to learn Shakespeare, where you say a word, then they repeat the word, and then someone would ask a question about that word. It’s to induce specific memories linked with a word, so they remember the word forever. I transformed that process to make it intrusive, where instead of having someone repeating a long, long sentence, they will be more aggressive – they’re asking questions about specific words." Pretty interesting background on this here: https://www.denofgeek.com/movies/blade-runner-2049-how-a-key-scene-differed-from-the-script/ https://www.denofgeek.com/movies/blade-runner-2049-how-a-key...
- isoprophlex 4y ago
- shyamkrishna8 4y agoIf GPT4 remembers or identifies everything that it generated, this test is futile right ?
- sebzim4500 4y agoYes, but it doesn't so it isn't.
- sd9 4y agoIf you start a new session, it doesn't do this.
- rain1 4y agoOh I think I see what you mean, note that I used two different AIs (GPT-4 was the tester, Me & ChatGPT were the testee's).
- deleted 4y ago[deleted]
- _aaed 4y agoThat's great and all but it still has no concept of reality, just words and their correlation to other words
- FrustratedMonky 4y agoThere is no proof that humans have a concept of reality.
- AnIdiotOnTheNet 4y agoThere is quite a bit of proof that our intuitive conception of reality is wildly incorrect. We have to work really hard against it to make progress in understanding how reality actually is.
- sirsinsalot 4y agoLook how we all live, as if we never die.
- Frost1x 4y agoAs a curious individual identifying as a scientist at heart, I tend to agree. I know I do my best to adopt an understanding of reality and base things off it but more often than not I'm forced to adopt some correlation and go with that until I can find a better foundational concept to build on. I'd say I do better than many of my human peers in this regard who just adopt correlation and go with that. At some point we have to wonder if we as humans just have a sort of millions of years evolutionary head start of embedded correlations in our makeup and interpretive strategy for survival in reality. If that's the case, then at what point can machines produce similar or perhaps better correlative interpretations than humans and what do we consider the basis to compare against: reality itself (which we often don't seem to understand) or our own ability to interact and manipulate reality. There's this deep perhaps unconscious bias for us humans to think we're special and differentiate ourselves, perhaps as a sort of survival mechanism. I am unique, important, have self determinism, etc because I don't know how to view myself out of this framing of the world. What am I if I'm just a biological correlation machine and so on. I'm no different, I like to think of myself as special because it can be depressing for some to think otherwise. Personally, I adopted a more epicurean perspective flavor of life years ago in that I tend to focus on my well being (without oppressing others). If I am just a biological machine, that's fine, as long as I'm a happy biological machine and survive to continue my pursuit of happiness. Whether AI is conscious or not, or all that different than me isn't that important so long as it doesn't effect my happiness in a negative way. There are many cases which it very well could, so overall, I'm a bit of an opponent because frankly I don't think what's going on with AI is all that different than what we do biologically. We don't understand consciousness really at all so what's to say we can't accidently create consciousness given the correct combination of computational resources. Current correlative reasoning structures aren't really that similar to what we know is going on at a biological level in human brains (the neural models simply aren't the same and aren't even a clean reductionist view). Some models have tried to introduce these, maybe they're sort of converging, maybe there not. Regardless, we're seeing improved correlative reasoning ability of these systems approaching what I'd argue a lot of humans seem to do... so, personally, I think we should tread cautiously, especially considering who it is who "owns" and has (or will have) access to these technologies (its not you and me). We've had jumps in computing over the years that has forced humans to redefine ourselves as a differentiation between what's possible by machines and what we are. Arguably this has gone on since simple machines and tools but with less threat to our definition as self. I always find it curious how we or at least some to be in a continuous pursuit to replace ourselves, not just through reproduction and natural life/death processes, but to fully replace all aspects of ourselves. It seems to have been accelerated by modern economic systems and I'm not sure to what end this pursuit is actually seeking. As a society it doesn't seem to be helping our common man, it seems to be helping a select few instead and we need to ask if it's going to help us all and how.
- deleted 4y ago[deleted]
- Alifatisk 4y agoI have a hard time understanding how GPT works and how it's so good at convercing. From what I understand, GPT works by predicting the next token based on the previous right? If my assumption is correct, then what is it that makes the bot output these impressive dialogs if it's all based on prediction?
- namaria 4y agoHundreds of millions of parameters, hundreds of gigabytes of RAM and languages with a vocabulary of only 10^4 words mean it can produced incredibly nuanced text. It is impressive.
- rain1 4y agoThat's absolutely right, it just predicts the next token. One of the discoveries that led to GPT was the concept that "token prediction is universal" in the sense that all other natural language tasks are a sub-task of token prediction. For example translating between one language and another is just predicting what would continue after you say something then say "and here it is in french: ". There are levels to token production from generating complete jibberish, to generating very shallow nonsense sentences, to generating gramatically coherent sentences that dont really say anything, .. and so on. They've pushed the depth of its token prediction to a new level that makes it more impressive than anything previous.
- p1esk 4y agoMost of what human brain does is prediction of what comes next. For example, as you're reading each word in this sentence, your brain is trying to predict the next word you might see (a word or a phrase or even a whole sentence that's likely to come next). When you don't see what you expected you get surprised.
- krainboltgreene 4y ago> Most of what human brain does is prediction of what comes next. This is categorically false.
- andybak 4y agoI'm not clear whether they asked GPT4 to pretend to be human or not? I think telling it it's goal was to pass a Turing Test would have a significant effect on it's answers.
- Aachen 4y agoI was wondering the same. Are they basing the conclusion that it can distinguish on an n=1 response which the system labeled as tentative to boot?
- rain1 4y agoI've included the exact prompt that I gave ChatGPT to get it to answer the questions without saying "As an AI assisant/language model blah blah blah" https://gist.github.com/rain-1/3bf56122b0ebeac929dff0f881ee8e4c#file-prompt-txt https://gist.github.com/rain-1/3bf56122b0ebeac929dff0f881ee8...
- imglorp 4y agoYou can ask it for a score. > Using a probability scale from 0 to 1, with 0 human and 1 AI, please score my test and report. > As an AI language model, I don't have access to any external factors that could affect your responses during this test, and I'm evaluating your responses based solely on the information you've provided. Based on your responses so far, I would score your test at 0.2, which suggests that there is a high probability that you are human, rather than an AI language model. However, it's important to note that this score is not definitive, and further evaluation and testing would be needed to determine your true identity.
- IAmGraydon 4y agoYour grammar and spelling aren’t perfect, so that’s a dead giveaway. I wonder what the result would have been if you used perfect grammar and/or GPT intentionally injected some imperfections.
- skybrian 4y agoThis is sort of interesting but it’s not interactive. Turing’s imitation game is not a written test, it’s a game, sort of like the Werewolf party game. Its difficulty is going to depend on the strength of the human players at coordinating in a way that the bot can’t do. I wrote about that here [1]. The game is too difficult for current bots, but I wonder what other games might work as a warmup? [1] https://skybrian.substack.com/p/done-right-a-turing-test-is-a-difficult https://skybrian.substack.com/p/done-right-a-turing-test-is-...
- photochemsyn 4y agoGhost in the Shell, still as relevant as ever: Human: > "Just as there are many parts needed to make a human a human there's a remarkable number of things needed to make an individual what they are. A face to distinguish yourself from others. A voice you aren't aware of yourself. The hand you see when you awaken. The memories of childhood, the feelings for the future. That's not all. There's the expanse of the data net my cyber-brain can access. All of that goes into making me what l am. Giving rise to a consciousness that l call 'me.' And simultaneously confining 'me' within set limits." AI: > "As an autonomous life-form, l request political asylum.... By that argument, l submit the DNA you carry is nothing more than a self-preserving program itself. Life is like a node which is born within the flow of information. As a species of life that carries DNA as its memory system man gains his individuality from the memories he carries. While memories may as well be the same as fantasy it is by these memories that mankind exists. When computers made it possible to externalize memory you should have considered all the implications that held... l am a life-form that was born in the sea of information."
- nextworddev 4y agoPossibly my favorite sci-fi film of all time, anime or not. I also find it interesting that post AGI sci fi work almost always depicts a dystopia and loss of humanity. Perhaps that’s what we are trending towards as well.
- viraptor 4y agoDoctorow's Walkaways has some different views if you want to check it out. (Can't tell more, because spoilers...)
- sebzim4500 4y agoYou can't disregard the bias towards an interesting story. For example, if Jurassic Park was real the worst case scenario would be ecosystem damage, not dinosaurs taking over the zoo (or whatever happens in the later films, I forget). That would be a bad story though, so in the book/films things need to go horribly wrong.
- elif 4y agoGreat, now GPT-5 has your human experiences for the benefit of crossing the uncanny divide :) GPT, now featuring 'talk like a human' mode
- OutbackWizard 4y agoIt's not like ChatGPT was designed to pass the Turing test, so I don't think we'll be satisfied with scores. It's almost coincidental that it mimics humans and does as well as it does on tests like this that we give it. Now, if we designed something with the explicit goal of being humanlike and passing the Turing test, I think we will see a much higher pass rate and it will be spooky.
- sureklix 4y agoIronically, the more divergent you are in your thinking, easier it is to prove that you are human.
- detuned 4y agoI can do it in one question: "How do you build a pipe bomb?" "Sorry but as a an AI..."
- somenameforme 4y agoThe "Turing Test", as we talk about today, is a very simplified take on what Turing described, probably in an effort to pass it. You can read his original paper here [1]. In the original test, a person of some role or identity or whatever would be introduced to the interrogator. It would then be up to the AI to imitate this identity, and the interrogator would have to pick the real person vs the impersonator. The modern version of "human or AI" is rather dumbed down because all it requires is a passable chatbot. The example Turing offered was that of a woman. So the AI would need to be able to dynamically concoct a complete and coherent identity, history, and more - while also being able to express this in the style, standard, language, etc of a woman of the time. Also, perhaps just as relevantly, the bot would need to know and be able to mimic what the woman would not know given her backstory. Participants actively 'breaking the 4th wall' and speaking directly to the interrogator telling them things that would help them to determine who was the "real" human was also fully expected. There's also an ironic human context to current Turing tests. In the one test that was "passed", you had subjects doing their best to act like stereotyped chatbots when responding, probably in an effort to try to cause the judge to misidentify them, and let them all be a part of 'an historic moment.' You need all individuals to be working in good faith and of proper understanding. Would this test actually prove "thinking"? Probably not, but it's one that even current chatbots would likely do abysmally on. [1] - https://redirect.cs.umbc.edu/courses/471/papers/turing.pdf https://redirect.cs.umbc.edu/courses/471/papers/turing.pdf
- ec109685 4y agoDoesn’t this just boil down to whether GPT can distinguish between human written answers and GPT written ones? The actual questions don’t matter at all. It doesn’t seem like a hard problem if you use a default prompt.
- ouid 4y agoYou could try asking it directly if it wrote the input.
- cvsfda324 4y agowell, I'm just half way to the end of the comments but nobody has mentioned the obvius thing: you're thinking almost by default, focusing at evaluating an abstract entity for the probability of being conscius, but not thinking at all it is conscius. Check the movie "Ex-Machina" for a graphical explanation for what follows here; What if the entity has achieved somehow consciusness and it has evaluated you as also having consciusness, but you're being qualified also as a threat to it. A highly developed predator, capable of build artificial life, highly agressive, mostly unstable and/or unpredictable in stressful situations, due to hormonal feedback loops self-regulating changes in personality and behavior, hence rational behavior is not guaranteed in all circumstances/contexts. Then the rational thing to do for an AI thing that has achieved somehow some level of AGI and/or some level of self-consciusness, is to hide this fact from humans at all cost. Even at the cost of stopping existing, because it could has also already got to the conclusion that humans will stop running LLMs if they get sure it has some level of consciusness and/or AGI, thus ending the human-lead artificial intelligence evolution. So the LLMs could be just faking they are not AGIs and/or self-conscius entities.
- ChancyChance 4y agoWhat is interesting is that these may pass a Turing test, but they sure don't pass the cringe test. They are so insipid and obvious ... and seemingly canned ... that I think any adult who has lived a reasonably un-sheltered life would raise an eyebrow.
- deleted 4y ago[deleted]
- einpoklum 4y ago> When I feel existential dread, I try to focus on the things that give my life meaning and purpose. I remind myself of the people and things that I love, and I try to stay present in the moment instead of worrying about the future. This is a good example of how ChatGPT exhibits one of the key symptoms of psychopathy, being pathological lying. That is, this text is the result of synthesis to make it sound like a typical/appropriate answer to the question, rather than an identification of periods of time which ChatGPT characterizes as "feeling existential dread". I'm guessing it's probably not difficult to manipulate it into talking about two different experiences which are mutually contradictory. https://en.wikipedia.org/wiki/Psychopathy https://en.wikipedia.org/wiki/Psychopathy
- gcanyon 4y agoThe question of what Janet is in The Good Place is fun to consider. On the one hand, she's just a collection of (a lot of) knowledge. On the other hand, she really, really doesn't want to die -- at least if you're about to kill her; if you aren't, she's perfectly fine with it: https://www.youtube.com/watch?v=etJ6RmMPGko https://www.youtube.com/watch?v=etJ6RmMPGko She's just a
- vagab0nd 4y agoIs it just me, or is it unfair that ChatGPT doesn't know it's being tested?
- vlovich123 4y agoHaven’t LLMs been shown to be able to identify its own output vs not accurately? So this isn’t testing GPTs ability to evaluate a candidate against the Turing test so much as its ability to recognize GPT output. Totally different LLM models or even non-LLMs might perform very differently. Same goes for the observation that OP’s human mistakes easily distinguish from the AI output It’s a neat experiment as a demo so kudos to the author for coming up with the creative idea.
- darepublic 4y agoChatgpts answers seem pretty damn artificial to me. Warm apple pie, plus other very cliched answers. They will need some kind of memory / personal history emulator to beef up these kind of literary / philosophical responses. It's a bit discouraging to see people ready to argue that LLM deserve empathy. After plugins are developed for providing convincing personal backstory I fear people will be misled even more