5 ms·
When I first watched the XPrize flight in 2004 (!!) on a tiny pixelated video over an ISDN connection, I thought we would only be a couple of years away from re
by biofox 3y ago
When I first watched the XPrize flight in 2004 (!!) on a tiny pixelated video over an ISDN connection, I thought we would only be a couple of years away from regular space tourism. I never would have dreamt that chatbots would pass the Turing test before Virgin Galactic became operational. Predicting technological advancements is hard.
- TaylorAlexander 3y agoI really don’t think our current crop of chat bots can pass a Turing test. https://news.ycombinator.com/item?id=37057043 https://news.ycombinator.com/item?id=37057043
- justrealist 3y agoNone of those examples are GPT-4
- TaylorAlexander 3y agoI paid for GPT-4 and noticed some weird behavior that seemed reproducible, so you could test for those things to reveal if you’re talking to a person or a chatbot. Biggest thing is you can ask it for a story of someone’s childhood (you can’t ask it for its childhood) and it will provide a plausible sounding but not internally consistent story.
- 93po 3y agoYou can also just ask it what time it is, or what day it is, or how much time has elapsed since your last message. It doesn't know any of these.
- justrealist 3y agoThat's an incredibly stupid Turing test, and you know it.
- TaylorAlexander 3y agoIf the robot cannot answer that question it is simple and effective test. The whole point of a turing test is that you cannot distinguish between a person and a robot. If you can wait one minute and then ask it how long it has been since your last request this sounds like a simple and effective test.
- justrealist 3y agoDo you seriously think it would be hard for OpenAI to give GPT-4 a timer? The point of a Turing test is a passive evaluation of how intelligent a bot is. It's not a test that OpenAI is trying to pass.
- TaylorAlexander 3y agoYou misunderstand me. To me, a Turing test is passed when the interviewer cannot distinguish between a robot and a human. If OpenAI were to add a special timer function to GPT-4, then that specific question would no longer work to tell the difference, but that specific question is just one example of many types of questions one can use to tell that GPT-4 is not a human. GPT-4 with a clock does not change the fact that there are many ways to tell it is not human. > The point of a Turing test is a passive evaluation of how intelligent a bot is. I don't know what "the point" of the test is (that is subjective) but I know how the test works. You create a system where a tester and a subject can pass messages to each other but cannot see each other (so a chat system is great for this) and then the tester tries to determine if the subject is a robot or a person. If the tester cannot tell if the subject is a human or a robot, the robot has passed the test. Many people have pointed out the test is not necessarily useful for evaluating AI! An AGI may be extremely intelligent and useful and still not pass the Turing test, but it would nonetheless be true that it did not pass. However a highly intelligent AGI would probably be good at pretending to be a human. > It's not a test that OpenAI is trying to pass. We are in agreement on this. They are not trying to pass this test and their current system can not pass it. This is why I say the current system cannot pass a Turing test. I am not talking about a theoretical future system!
- 3y ago
- Workaccount2 3y agoCutting edge bots fail because they are too perfect to be human.
- TaylorAlexander 3y agoActually I think they make weird choices sometimes that could be tested for systematically if someone knew which model it was and had some time with it before the Turing test took place.
- AbrahamParangi 3y agoThe main issues with gpt and the Turing test are not that it isn’t smart enough but rather that fine tuning for helpfulness and harmlessness has turned it into a Quokka. Completely harmless and naive in its attitudes towards harmful or controversial things, and this is trivially detectable.
- dTal 3y agoThis is a bit like saying that none of our current crop of computers are Turing complete, because their tape isn't infinite. We will never achieve a theoretically perfect Turing machine, or a program indistinguishable from a real human in every single respect including funny edge cases. But we can get close enough that it's meaningful to use the theoretical ideal in a colloquial sense that means "crosses the threshold in most contexts". GPT models can actually keep up a half-decent facade against a suspicious interlocutor for a while, something no program has ever done before.
- TaylorAlexander 3y agoI agree the software is doing better than ever before! However it still has serious shortcomings which is an active area of research. An AI researcher familiar with these issues could probably make an accurate detection in under one minute, so I don’t think this is an “infinite tape” issue this is simply an issue of performance. They are pretty good and they will get better, but I believe a true AGI would be able to mimic a human to properly pass the test.
- dTal 3y agoThat's entirely fair. It definitely feels like some sort of threshold has been crossed, and the only term we have for this performance axis is "the Turing test", but you're absolutely right that LLMs are very far from perfect. It feels more like the beginning of a journey rather than the end. We should probably get in the habit of treating Turing Test performance as more like a continuous scale rather than a pass/fail condition.
- satvikpendem 3y agoAtoms are harder than bits. I'm no longer surprised anymore when software advances break through while hardware advances like self driving cars and humanoid robotics languish, because the real world really is much messier than moving data around.
- e-master 3y agoSoftware scales better than hardware, so most investment (capital and intellectual) is in software.
- orbital-decay 3y agoML advancements wouldn't have been possible without the extreme progress in hardware. It's not a hardware vs software thing, rather it's a function of demand. The business model of space tourism in general was always dubious. Although in case of SS2 in particular, it also got delayed by the experimental nature of its hybrid engine (they changed fuels and redesigned it several times) as well as their approach to safety (2014 disaster).
- ylee 3y agoI had just joined a large investment bank and was in NYC for training when I heard that the XPrize flight had succeeded. Although finance has been my career, I grew up wanting to be an astronaut and closely follow space science. I also expected that suborbital manned tourism flights would become routine, and that in time SpaceShipOne and similar designs would be scaled up to reach orbit. Although the latter didn't happen, I am glad that SpaceX came along and made reusing (!) rockets that land at their own launch pads (!!) routine. It is not impossible that I will someday ride Starship or a descendant to orbit.