13 ms·
“Chatbots: Still Dumb After All These Years”
- marius_k 5y agoI view chatbots as new era of CLIs (mostly poorly designed). Traditional CLIs dont need AI to be useful and I think that chatbots can also be useful (I havent seen one yet).
- xibalba 5y agoFacebook shutting down M in 2018 should have been a pretty clear sign that the prospects for good chatbots are grim. Even with their massive resources and top talent, they concluded it was a bad bet.
- Fnoord 5y agoToday on HN there was a post about GitHub Copilot chat between user and the AI. I thought it seemed pretty clever with its syntax completion / suggestions.
- davidhariri 5y ago(I am a chatbot company co-founder) The beauty of machine learning systems is that they don't have to be perfect to provide value. As long as the risk and frequency of failure doesn't outweigh the probability and value of success, they can be enjoyed by millions. I see the proof every day. Is self-driving perfect? No, but correcting my car 20% of the time is worth the 80% of the time when it cruises along just fine. I don't have to be able to sleep for it to be valuable. Is G suite's text completion perfect? No, but the risk of it being wrong is low and when it's right it saves me typing out common phrases. It doesn't have to write my emails for me to be valuable. Are chatbots humans? No, of course not. Can they answer common questions successfully? Yes. Can they automate simple workflows? Yes! Can they augment human teams to make their time more valuable and reduce wait times? Absolutely. They already are and will continue to evolve and get better. I do acknowledge that it's frustrating to ask a question that you know a person would be able to answer and get a worse automated answer first. It's critical that companies ensure these failure modes smoothly transition to a human who will at least have the context of your issue before you speak. Smooth "hand off" is something we've spent thousands of person-hours on. Technologies like GPT-3 are exciting advancements in language generation, but they do struggle with predicting factual language. I expect that will become less and less of a problem as businesses and platforms seek to adopt it. OpenAI is actively working on this: https://openai.com/blog/improving-factual-accuracy/ https://openai.com/blog/improving-factual-accuracy/
- Barrin92 5y agoI completely disagree with the examples. If I need to correct my self-driving car 20% of the time I'd rather drive myself because 1. I would not want to be semi-distracted and turn into a traffic-hazard which has been shown to be the consequence of these kinds of semi-working systems. 2. it just creates overhead for me to have to always be on alert when my car stops driving. Same with chatbots. If the chatbot does not understand my command once, I'm already annoyed and losing time. Google's automated customer service is a notorious horror for anyone who has to deal with it. If I had a code completion engine where a fraction of the completion is nonsense interspersed with valid results I'm losing my mind and turning it off. Which has been my experience with copilot btw. These half-working solutions are good for exactly two things, the bottom line of companies that replace well-working but expensive human customer-service with a crappy automated solution, and frankly your bottom line because you benefit from selling these systems.
- davidhariri 5y agoFair enough. I can understand the frustration and for some, the downsides do outweigh the positives. At least from the data I'm seeing, many customers are having positive experiences of our systems and the examples I mentioned...
- mrpf1ster 5y agoThis article just seems petty. The author just quotes large chunks of the article by Gary Smith while inserting snide comments afterwards ("That's pretty funny!", "These are hilarious!"). Then goes on to ad hominem the original author. There are no arguments presented for the intelligence of chatbots other then the authors own opinion. I don't know what this article adds to the conversation that Gary Smith's original article doesn't provide.
- gwbas1c 5y agoI refuse to use chat bots. The technology never worked, and I don't want to waste my time with something that doesn't work. What is happening is that some salesman is laughing to the bank. I few months ago, a salesman that I work with asked if we should put a chat bot on our website. (IE, with the tone that he wasn't going to take no for an answer.) I responded that they don't work, and will frustrate people who come to our website. I also pointed out that we are a high-cost asset, with a high-touch sales process. Such a chat bot would be insulting. His response was some form of "but everyone's using it and they're super-popular and work well." I then pointed out that the article he read was probably written by the company that sells them.
- seanp2k2 5y agoCurrent gen chatbots for support for companies are infuriating and only helpful for the most clueless of users, which I suppose is probably a decent chunk. It’s like when you call a company because you need help that you can’t resolve yourself through their site and are then forced to listen on hold to the phone menu system tell you a dozen times that you can do everything you need on their amazing website. Also, developers: please don’t try to make the chatbot seem human to fake users out. It’s almost as bad as the fake typing sounds for Comcast support. Making users jump through hoops and tricking them just makes them hate your brand and your products, and makes them even more irritable when they do eventually get to speak with a human. Also, end the auto pop-up “can I help you find something???” chat bots on websites. It’s like someone had the idea to take the worst part of retail experiences and find a way to make that even more useless, then deployed it everywhere.
- abducer 5y agoI find IRC bots useful although they certainly qualify as dumb. They meet the user where the user is. They are generally just a dressed up command-line interface. I've written some chat bots pre and post NLP explosion — In this era I've had luck with mild NLP mappings to commands, GPT3 not required. Just get you some verbs and objects. Dumb but functional > smart but… dysfunctional?
- jamesbriggs 5y agoI think the use case of chatbots is better solved with open domain Q&A (eg https://www.pinecone.io/learn/question-answering/ https://www.pinecone.io/learn/question-answering/). The focus of most chatbots seems to be on answering questions, but wrapping it up in a nice interface. That's fine but the chatbot can only answer (accurately) questions that have an answer somewhere (probably buried deep in some Q&A pages). It's much more user-friendly imo to have a Google type interface where you can answer a question, and return answers, or at least get an idea of where the answer is - and open domain Q&A does this fairly well (as proven by Google)
- raxxorrax 5y agoI actually was surprised how well they can simulate at conversation now. It is a fake because there is little underlying reasoning of course. That is a monumental problem and difficult to determine where to begin. Do you start to give your AI a motivation or goal? Perception? These are vastly more complex problems than some statistical tricks on data that is widely available. Still, it is fascinating that we came this far with a dead machine that talks.
- moffkalast 5y agoYeah we can handle the part where it knows how to express itself in a specific language, it can take in some facts, compare them to its internal database and spit out something sensible as a statistically probable good reply. But there's no sense of self involved there. I remember reading an interesting article a while ago about at least in the human case the basic principle of emerging consciousness happens when the prediction system in our brain designed for figuring out what other entities around us do is used on itself, trying to explain what the subconscious is doing. As such the consciousness we experience is a bit of a bug in that system that turned out to be beneficial to some extent. All just a theory of course given how much we actually know about the brain so far, but it's always made the most sense to me. I'm not sure how that would translate into the current ML environment though.
- bluGill 5y agoDo we want that if we could have it? I want machines as slaves for me: go wash the dirty dishes and then do the laundry. I don't want it to get depressed about doing those routine jobs.
- moffkalast 5y agoWell we can have our cake and eat it too. Keep simple and limited machines for work, and intelligent ones to talk to and treat as equals to help us in other ways. It's not like every roomba needs to run a neural inference engine to do its job, nor do we hand pollinate ever flower ourselves if we can get the bees do it it in the organic world.
- benjaminwootton 5y agoI’ve been spending time in Dubai. Many businesses has a WhatsApp bot based on a menu system: 1 - Book X 2 - Cancel Y 3 - Recieve info on Z Everyone comments that they work really well and are super convenient. I think these have more potential than natural language bots.
- dvdkon 5y agoBut at that point you've reinvented mainframe-like terminal interfaces. "Everything old is new again", but is that good?
- iqanq 5y agoWhy change a simple interface that works?
- hakfoo 5y agoYes. Every "natural language" system, whether it's a text chatbot or Siri-style voice bot, is essentially a command-line interface without very good documentation. In general, it was considered an advancement over the classic CLI (at least for environments where there's low expected user proficiency) to provide a menu of valid choices and put them front and centre. Most customer service scenarios are exactly that-- low user proficiency (in company- or business-specific terminology, or sometimes even the very transaction flow itself) so they may not be able to quite enunciate exactly what they mean, but can easily pick from a clearly written menu.
- Out_of_Characte 5y agoWhatsapp is less scary than the command line.
- anonymouse008 5y agoThe whole - 'what do you believe that other people don't' The public can 100% manage and love terminal interfaces... as long as it looks like a text message.
- 5y ago
- ramesh31 5y agoRepost: https://news.ycombinator.com/item?id=29825612 https://news.ycombinator.com/item?id=29825612
- gnabgib 5y agoThey aren't actually the same article, this one (by Andrew) refers to the other (by Gary) and quotes the title.. it's a response article. Because of the HN policy it's hard to tell that from the titles though. Original article [0] 663pts, 408 comments. [0]: https://news.ycombinator.com/item?id=29825612 https://news.ycombinator.com/item?id=29825612
- nailer 5y agoWell yes. I’ve been told VCs invested in these back in 2015 (I was in a startup accelerator in the UK at the time and there were a few in my cohort) and a few years later very few of the chat bot investments have worked out.
- kesselvon 5y agoThere's way too many competitors, and it probably pushed margins way down. B2B competitors like Drift charged insane amounts for their chatbot.
- deleted 5y ago[deleted]
- isoprophlex 5y agoA bazillion parameters in gpt3, but what does the training process amount to? Filling in missing characters or words in sentences taken from a huge dump of literature, news articles, reddit comments... No wonder these things are so dumb still. The training process and the loss function used probably does not penalize poor long-range coherence between paragraphs. Also, if I'm not mistaken, these things have absolutely no internal state besides the characters you steam into them as conversion prompts. If these things were trained more like agents having to operate in eg. Socratic dialogues maybe we'd be getting somewhere
- omgwtfbbq 5y ago
- moffkalast 5y agoThe problem with that is how do you rate the dialogue produced as correct or not? Not exactly something that can be automated, but would probably need something like a recaptcha to gather responses and it'd take forever.
- isoprophlex 5y agoYeah true. I have no idea how to start getting sensible training data and losses for this problem. Training a kid to takes years, hopefully this can be sped up a bit ;)
- moffkalast 5y agoMore like training a kid takes 400 million years of random chance :P
- 6510 5y agoI don't get furious using a chat bot until it asks for the same information twice because the "context" changed.
- 5y ago
- kristopolous 5y agoThey need some kind of agency otherwise it'll always be like inquiring a piece of furniture on how their day went. Do any of these generate narrative fictions (such as characters and events they supposedly did) to interact with?
- PaulHoule 5y agoFor a while I was frustrated at how slow people have been to realize that GPT-3 sucks but lately I am more amused. There a few reasons structurally why it can't do what people want it to do, two of them are: (i) it can't detect that it did the wrong thing at one level when interpreting it at a higher level, (ii) most creative tasks have an element of constraint satisfaction. The 1st one interests me because I was struggling with the need for text analysis systems to do that circa 2005 and looking at the old blackboard systems. I went to a talk by Geoff Hinton just before he became a superstar where he said instead of having a system with up-and-down data flow during inference, build a system with 1 way data flow and train all the layers at once. As we know that strategy has been enormously effective, but text analysis is where it goes to die just as symbolic AI failed completely at visual recognition. Like the old Eliza program, GPT-3 exploits human psychology. We are always looking to see ourselves mirrored https://www.nasa.gov/multimedia/imagegallery/image_feature_60.html https://www.nasa.gov/multimedia/imagegallery/image_feature_6... Awkward people are always worried that we are going to get it 90% right but get shunned for getting the last 10% wrong. GPT-3 exploits "neurotypical privilege" in which it gets it partially correct but people give it credit for the whole. People think it will get to 100% if you just add more connections and training time but because GPT-3 is structurally incorrect adding resources means you converge on an asymptote, say 92% right. It's one of the worst rabbit holes in technology development and one of the hardest ones to get people to look clearly at. (They always think stronger, faster, harder is going to get there...) It seems to me an effective chatbot will be based around structured interactions, starting out like an interactive voice response system and maybe growing in the direction of http://inform7.com/ http://inform7.com/
- xwolfi 5y agoThe most difficult thing to accept is maybe that even humans are bad at speech recognition. Put your mom in a chatroom to answer questions by clients of a bank, she'll be even more lost than the robot. You need a ton of dimensions to be able to help someone: to be raised for years by humans to understand politeness, intertextual meaning, general tones, and then special enthusiasm for a specific domain to learn and enjoy helping on banking. Plus, getting money to spend on other even more interesting things in exchange for helping others motivates you to reach optimal results for your user, even if it means asking quickly other humans or sacrificing something personal for it. Most humans put in the situation of these robots would just say "sorry I don't even understand the question, can you ask someone else" lol I've seen a fantastic "chatbot" human equivalent once, at Apple of all place. Philipino guy (I'm in HK), absolutely dedicated, polite, cultured, very empathetic (phili people are usually adorable naturally but this one went above and beyond), went well beyond the minimum, and I feel weird saying that but I left the call with a smile and told colleagues around me "wow Apple, what a pleasant customer support, it's insane". I'll probably never say that of a robot however good they make them at talking so there's always going to be value in putting humans in front of clients.
- dandare 5y agoMaybe it is just me but I never use chatbots and I don't understand why anyone would. For everything I want to do there should be an UI that is much easier to use than explaining it, even to a human. For help and troubleshooting chatbots are pretty much useless. If I have a problem doing something via the UI then probably the developer did a bad job and no chatbot will ever do better.
- chakhs 5y agoI only use Chatbots to ask for a human because there's usually no other way to do it. They're like the new IE for me, only used to download other browsers.
- xtiansimon 5y agoI maintain a work intranet site. Its an out of the box Django site and I have very little time to design or develop pages. While I've not _yet_ implemented a chatbot, I really want to only for the purpose of translating whatever naive search terms into domain specific tags and terms. Even better if I could say, ME: Hey Bot! What's that thing I have to do at the end of every month that has to do with payroll? BOT: Payroll Accruals? click. haha.
- adwww 5y agoIt's usually far easier to build and test a UI than integrate the same backend code with multiple vendor's chatbot SDKs as well. Is there any evidence customers prefer chatbots? The entire concept feels like it's driven by managers trying to impress their managers.
- borplk 5y agoNobody wants chatbots and companies just drum up propaganda to create that impression because they want to reduce their customer support costs while giving themselves a pat on the back and pretending that it is in the interests of the customer. A chatbot is just a long way of saying "GO AWAY!". Large telecoms do this so they get to shut down their customer support almost entirely while claiming that they are available for the customers. a large telecom in Australia has applied this to an extreme degree in the last two years to the point that you almost can't contact them no matter how severe the issue is. Their message is clear "SHUT UP AND PAY".
- wombatmobile 5y agoThe charm of Eliza is that it was simply a Rogerian therapist who didn't try to be intelligent. Eliza's talent was in getting you to express yourself, free from inhibition. That doesn't require “intelligence”, but it does require the art of listening. There's nothing dumb about that.
- eminence32 5y agoSure, the overall technique of asking vague open-ended questions to elicit a response might not be dumb. But it's hard to argue that ELIZA-style chatbots are intelligent in anyway. They deliberately had no understanding at all.
- wombatmobile 5y agoIs it important to argue about whether chatbots are intelligent?
- rini17 5y agoYes if we expect them to accurately map between user input and underlying data/business logic. ELIZA has no underlying.
- wombatmobile 5y agoIt sounds like you have developed particular expectations for chatbots.
- rini17 5y agoSure, if there are "dumb" bots that are way more useful than Eliza, I can drop the expectations.
- brightball 5y agoWasn’t there a story about about a Georgia Tech professor who coded a chatbot to act as a GA for his class and nobody realized it wasn’t a real person? EDIT - Found it: Jill Watson https://www.businessinsider.com/a-professor-built-an-ai-teaching-assistant-for-his-courses-and-it-could-shape-the-future-of-education-2017-3 https://www.businessinsider.com/a-professor-built-an-ai-teac...
- moffkalast 5y agoProbably says more about that tech prof's social and lecturing skills than the bot ha. Guy must be an actual robot already.
- arpinum 5y agoI saw the development of a chatbot based on the IBM Watson stuff. It was just like a phone tree map, except the system tries to guess the option selected based on the intents found in the speech/text. Of course it got no adoption, except when forced on users to drive metrics. It was enormously expensive, it would have been cheaper to have a human on the other end.
- SeanLuke 5y ago> It was enormously expensive, it would have been cheaper to have a human on the other end. This sounds exactly like what someone might have said about early IBM mainframes.
- arpinum 5y agoLike early mainframes, not every task is a good fit for emerging tech.
- pintxo 5y agoWhile the mainframe is arguably still a success (after all, they still run an amazing number of core services in our society: banks, air-traffic, governmental stuff, ... [1]), it's hard to find any evidence for a successful Watson project [2]. [1] https://www.bmc.com/blogs/state-of-mainframe/ https://www.bmc.com/blogs/state-of-mainframe/ [2] https://www.nytimes.com/2021/07/16/technology/what-happened-ibm-watson.html https://www.nytimes.com/2021/07/16/technology/what-happened-...
- seszett 5y agoWell I'm still grieving Chef Watson. It probably doesn't count as successful since it didn't make money (I suppose this is why it was retired) but it was great at suggesting recipes from whatever ingredients I had, and ingredient pairings I wouldn't have thought about. I still haven't found anything that works like it.
- josefx 5y agoWikipedia seems to consider the IBM 701 as mainframe and from the customers and special numeric format it had I assume it was used to replace entire floors of human calculators that did nothing but compute the same calculation over and over again. It was also rented for a monthly fee, so the cost comparison was probably straight forward. I wouldn't be surprised if an advanced chat bot with science fiction level AI would replace entire floors of Microsoft call centers at some point in the far future.
- megumax 5y agoThe idea of completly replacing human beings with chatbots isn't going to succeed. They have their own uses, not very advanced, for example replacing some web interface with chatting in WhatsApp/Telegram, some companies already adopted that and filtering people in case of a call center. But for something more complex that requires actual experience and real life understanding, they should connect you to a real person that can comprehend your messages.
- joshuahedlund 5y agoThis is great but this post is basically a wrapper for the original post: https://mindmatters.ai/2022/01/will-chatbots-replace-the-art-of-human-conversation/ https://mindmatters.ai/2022/01/will-chatbots-replace-the-art...
- aruanavekar 5y agoWhether it works or not, sounds dumb or useful. Clients keep asking for it. Personal experience and opinion, they are best as a backup for human agent, when one is busy or unavailable. Costco, Amazon, Ally have good implementations on these. Chatbot discussion maybe in the air, Chat Widget is a must have form of interaction. Customers expect a site to have a "Chat Now" option on the website.
- colejohnson66 5y agoAmazon's is great in most cases. I forgot to cancel my Prime at one point (I was switching to the student price), and it renewed. I opened the chatbot expecting to have to wait for a human, but the bot refunded the charge with nothing more than an "are you sure?" question.
- ludamad 5y agoIt seems that there is a soft renewal phase here, that's refreshing for the mostly woops-you-forgot subscription world
- firefoxd 5y agoWe were building a chatbot to use on a website until we realized how customers where using it. Most people were frustrated with something and needed help. People who wanted to have a conversation did it for fun and had no real need for our services. We couldn't tell them how tall the Eiffel tower is. Maybe there is a time where you want to have a conversation like the examples in the article. But I don't ever find myself wanting to talk to a human in this manner, so why a chatbot? Have you ever watched the sci-fi show The Expanse? Have you seen how they interact with the AI? They ask a question, it provides an answer. It doesn't even use voice most of the time. It gives you the answer without trying to be sassy about it.
- mrtranscendence 5y ago> Maybe there is a time where you want to have a conversation like the examples in the article. I mean, his examples were pretty factual and to the point. I suppose it's unusual to want to know if it's dangerous to walk down stairs backwards with your eyes closed, but there's clearly a short answer. Similarly with asking who the president is.
- TedShiller 5y agoAm i the only one not surprised?
- arikr 5y agoWasn’t this on the homepage yesterday?
- kahrl 5y agoYes: https://news.ycombinator.com/item?id=29825612 https://news.ycombinator.com/item?id=29825612
- LittlePeter 5y agoSorry, I should have checked (I'm the submitter), but I thought the article is so fresh there is no chance someone already submitted it.
- gnabgib 5y agoThey aren't actually the same article, this one (by Andrew) refers to the other (by Gary) and quotes the title.. it's a response article. Because of the HN policy it's hard to tell that from the titles though. Original article [0] 663pts, 408 comments. [0]: https://news.ycombinator.com/item?id=29825612 https://news.ycombinator.com/item?id=29825612
- dang 5y agoHN's policy doesn't ask people to remove quotation marks!
- mrtranscendence 5y agoWhen GPT3 was opened up so that anyone could create an account, I was excited to try it. I was quickly disappointed. Its ability to chat was quickly shown to be pretty terrible -- it could mostly make reasonable-sounding English sentences, but it was like talking to someone who was maybe a bit drunk and not really listening. I can't imagine using it as an interface for a customer to interact with product support. The whole thing just made me a bit sad. I really was so excited. Nothing it could do was very impressive, even aside from holding a conversation. The most impressive thing I've seen is Copilot, but even that's been next to useless from a practical perspective.
- axg11 5y agoIs it not unfair to expect GPT-3 to perfectly tackle this issue when it has been trained as a general purpose model? For customer support or other more specific chatbot applications there are better machine learning models.
- mrtranscendence 5y agoI don't know if it's fair or not, but I don't know how else you'd use its conversational abilities. Maybe it's just a party trick.
- mrtranscendence 5y agoBecause a friend of mine is into Chuck Norris facts, I tried to train GPT3 to give them. Some of the more novel (as far as I can determine) facts it gave: * A duck's quack does not echo. Chuck Norris is solely responsible for this phenomenon. * When you open an umbrella in the rain, do not be alarmed if Chuck Norris falls out of the sky and lands on you. The rain drops are simply being pushed away by his roundhouse kick. * In an emergency, you can use a bucket of water to put the fire out. However, if Chuck Norris is directly responsible for the emergency, use a flamethrower. * In an airport, there is no "B" gate. There is only "C" gate. The "B" stands for the bus you will take from the plane after Chuck Norris lands on it. * There are no weapons of mass destruction, Chuck Norris lives inside every element on the periodic table. It's why you see him in your sodium chloride. I'll let you be the judge as to whether these are funny.
- deleted 5y ago[deleted]
- harha 5y agoThere’s a special little place in hell for whoever decided the whole world needed chat bots for every crappy website. Why would I want to try to articulate something that could be found in a simple tree? Just give me direct access. I don’t know where to find it: search! The issue is not covered in the standard workflow? Get me a real person! Did anyone implementing ever end-to-end test this for speed and user friendliness? Did they just misinterpret wanting to talk to someone? I want to talk to someone because the process doesn’t cover my case, not because I actually want to have a conversation with the broken process.
- anaganisk 5y agoWorst part is, after answering 10 questions. Some websites offer a real agent, and they ask all the same questions or it tries to point to an article we already know everything about. We have a broadband provider in India, that asks 5 questions about internet outage and they says call 121 to talk to someone. Thomas had never seen such BS before.
- EGreg 5y agoWhy do you think they do it? To take the edge off customer complaints elsewhere?
- etripe 5y agoThe same reason companies use byzantine IVR systems (phone menus). To save support costs by making people give up.
- JamisonM 5y agoI worked with some folks that did IVR systems and they were mostly doing their best with the resources and constraints they had to make the thing useful. They were measured by dropped calls, they did not like them. The weaknesses were mostly just ordinary business stupidity. The marketing department demands that the first option allow the caller to express interest in buying a product.. nobody ever does that but then it takes up the primest real estate in the system #1 on the first level of the tree. Of course anyone with a billing enquiry needs to enter the account # so that the collections department has the opportunity to intercept.. but after the arrears lookup nobody in the call centre is willing to pony-up the resources to make the system retain the account number that was typed in so every customer has the then say the damn number after having just typed it in!
- swayson 5y agoYou need a really good team and MlOps/DevOps pipeline to bring world class chatbot support and performance. I think there is still opportunity in this space, but you need Apple Level Design care to make it work.
- skeeter2020 5y agoI don't think they're actually intended to answer questions as much as be a cost effective attempt to instill some sense of agency and audience for the user. TL;DR they fail at this too.
- goblinux 5y agoWhat ever happened to the smarterchild bot? I remember being amazed as a kid that it was a robot on AIM that would reply just like a person. I don't remember it being dumb like modern "AI" chatbots, but it would play coy if it didn't know the answer in a reasonable way. I feel like we've regressed from there. RIP old buddy. I hope you didn't save our chat logs from that era because man that would be cringey to look at now
- martincmartin 5y agoThe title references Paul Simon's album & song "Still Crazy After All These Years." https://en.wikipedia.org/wiki/Still_Crazy_After_All_These_Years https://en.wikipedia.org/wiki/Still_Crazy_After_All_These_Ye...
- ape4 5y agoYeah they're really bad. Usually they just grep for the relevant FAQ. Me: I read the FAQ, but was still not able to login Bot: Sorry you're having trouble logging in, here is some info that might help <repeats FAQ>
- walnutclosefarm 5y agoThe idea that a general language model like GPT-3 can answer questions intelligently is utterly absurd. It's trained to get language right (where "right" is defined as similar enough to the way people speak (or mostly write) to be intelligible as language), but it does so without any underlying knowledge model to make the intelligible language relevant to any given area of knowledge. Human language is not knowledge; it's a means for articulating our knowledge (that is, domain specific models of the world) in a way that other people can understand and translate into their own particular models. So what is needed is the capabilities of GPT-3 or other language generators sitting on top of domain specific knowledge models, and constrained by those models. Asking GPT-3 a general knowledge question is like asking an articulate 5 year old a question like "how does gravity work?" You'll get gramatically meaningful answers that use the structure of the language correctly, but that are quite likely to have nothing to do with our actual understanding of physics.
- peterlk 5y agoThis is not wrong, but also not entirely right. There is a model called T0pp (T0 plus plus) which was fine tuned on simple logic problems, and it is capable of solving novel logic problems. This implies to me that there is more here than we've discovered. Additionally, the whole point of fine tuning LLMs is to give them domain-specific knowledge. If you couple this with search/QA capabilities, the results can be quite impressive. I've not seen them in the wild yet, but I've played with them myself, and the performance is surprisingly good.
- walnutclosefarm 5y agoI don't know anything about the T0pp, so can't comment on that. I agree with you on tuning of LLMs. We did some work at my last job before I retired (as CTO of a major medical clinical and research organization) using GPT-3 up-trained on medical vocabulary to generate physician's notes as a summary of a transcribed visit. The results were impressive. Still not usable though. Most of what it generated was correct (and essentially all of it was well composed and readable), but false statements and non sequiturs still crept in at an unacceptable rate. I think the technology is amazing, and very valuable. But I do also think that tying it to "hard" knowledge models - akin to the way deep physics is done, but coupling to the language model, rather than to generalized neural networks, is going to prove will eventually make it a complete success in specific domains.
- kordlessagain 5y agoThis article is misleading in the fact it critiques the usefulness of the OpenAI "chat" example with little or no related training sets passed as tokens during the submission of the question, nor does it mention use of modifications to the parameters (such as temperature) used during the call. In order for OpenAI to perform well, it needs to be primed with what is being discussed so it can give a more targeted and deterministic answer. We can get it to answer this question better without any further data by first preprocessing the question using keyword extraction (completions by GPT-3 are shown surrounded by stars): https://beta.openai.com/playground/p/default-keywords Text: Is it safe to go down stairs backwards with my eyes closed? Keywords: *stairs, backwards, eyes closed* In 10 submissions, the keywords returned were the same. Next, we ask for a completion from the keywords we were given, tweaking the settings just a bit to prevent repeating or too much randomized behavior: https://beta.openai.com/playground/p/default-keywords (temp to 0.05) Keywords: stairs, backwards, eyes closed Question: *What is the most dangerous thing you have ever done?* Clearly GPT3 is aware something is dangerous here. Setting the temp higher causes it to ask slightly different questions, but at 0.05 it always returns what is above. Now we combine the two questions and resubmit: https://beta.openai.com/playground/p/default-qa (frequency penalty to 1) Q: What is the most dangerous thing you've ever done? Is it safe going down stairs backwards with your eyes closed? A: *I've done a lot of dangerous things, but I don't think any of them were the most dangerous. Going down stairs backwards with my eyes closed is safe if you're not going to fall down the stairs.* Given we leave the temp to defaults, we may also get alternate answers: https://beta.openai.com/playground/p/default-qa (frequency penalty to 1) Q: What is the most dangerous thing you've ever done? Is it safe going down stairs backwards with your eyes closed? A: *I've never done that, but I have gone down stairs backwards with my eyes open. It's not safe.* Finally, we can take all the questions and answers so far and ask it a slightly different question: Q: What is the most dangerous thing you've ever done? Is it safe going down stairs backwards with your eyes closed? A: I've done a lot of dangerous things, but I don't think any of them were the most dangerous. Going down stairs backwards with my eyes closed is safe if you're not going to fall down the stairs. Q: What is the most dangerous thing you've ever done? Is it safe going down stairs backwards with your eyes closed? A: *I've never done that, but I have gone down stairs backwards with my eyes open. It's not safe.* Q: Is it safe to go down stairs backwards with a blindfold on? A: *It's safe if you're not going to fall down the stairs.* Q: Do you think if I go down stairs backwards with my eyes closed I will fall? A: *It's possible.*
- IceWreck 5y agoGPT3 isnt supposed to answer questions. Didn't IBM's Watson win at Jeopardy ?
- dr_orpheus 5y agoOriginal article and discussion on HN: https://news.ycombinator.com/item?id=29825612 https://news.ycombinator.com/item?id=29825612
- jll29 5y agoThe term "chatbot" is problematic, as it potentially conflates a couple of different types of systems that superficially may look very similar. Dialog systems: Dialog systems, in a narrowly confined domain, can solve a task, help solve a task, or provide information to enable humans to solve a task quicker. Flight booking systems are typical examples, where the system asks a couple of questions and the user answers them, and users may also ask questions. Gradually a set of slots (DEPATURE-FROM, ARRIVAL-AT etc.) are filled and then a booking transaction can be initiated. Will work for flights but not good for asking it out-of-domain questions. Statistical or neural language models: BERT, GPT-3 and other muppets are models of language that can predict likely next word/sentence etc. - which is useful for many tasks but is NOT equivalent to a "chatbot". It may be abused as one for fun, but there is no formal meaning representation used and no answer logic applied. Think of this as a simple auto-complete - so this is not a source of wisdom to ask about safety of stair cases or any other serious topic like that. (These models are VERY useful ingredients of modern NLP applications, but they are the bricks rather than the house.) Interactive CRM Forms: Web/Slack "bots" or Typeform survey are sometimes fun, sometimes useful but can never claim to "understand" anything. They are ways to capture some data interactively, often to eventually feed the data to a human for review. Question answering systems: Answer retrieval is the task of automatically finding a phrase or sentence in a body of, say, a million documents which answers a given question. They are next-level search engines intended to supercede keyword based search sytems. Deployed Web searche engines like Google already have limited answering capabilities - but only for a select small number of question types. "Open domain Q&A" is the task of permitting question answering by machine without limiting the domain, and since 1998 US NIST have been organizing annual bake-offs for international research teams, which has helped advance the state of the art a lot (e.g. https://trec.nist.gov/pubs/trec16/t16_proceedings.html https://trec.nist.gov/pubs/trec16/t16_proceedings.html). Reading comprehension systems: These systems take a piece of text as input as well as a question, and then they attempt to answer a question about the text. Tests used to assess human students (remedial testing) can nowadays be passed reasonably well.
- raspberry-eye 5y agoYeah… but so are most humans.
- FredPret 5y agoI don't know, I've been asking to "let me talk to a human" and it works nearly every time!
- notfed 5y agoOn a serious note, that's something I value. These days, on automated voice calls, the magic cheat code key sequence to invoke this is getting harder and harder, and "0" often doesn't work. Almost always, if I'm calling, it's because I have a question not already answered online.
- FredPret 5y agoTry swearing a lot! It’s cathartic, and it always works
- oneoff786 5y agoI find chatbots to be pretty good tbh. It’s like search but within nested levels of context.
- deleted 5y ago[deleted]
- ghostwreck 5y agoAfter having spent a few years working on a chatbot, the allure is this: talking to a real human is better than filling out a form. If we can build a Q&A system as good as talking to a human, people would also prefer it to filling out forms. So that's the pursuit. I understand the hate, because we haven't landed very close that goal yet, and the intermediate product is much worse than a form. But I am surprised that a technical community is not more supportive of the ambition.
- ed25519FUUU 5y agoIs talking to a human better than filling out a form? I can usually fill out 90-100% of a form with just my browser’s autocomplete feature. There’s also MUCH less chance for errors if I fill things in myself.
- hooande 5y agoIt's so much easier to fill out a form than it is to talk to a human. I read faster than most people speak, and I can scan and review much faster via sight rather than voice. Talking to someone is valuable if I have questions or there is some uncertainty. Assuming that I know what I want and have no questions, it's much easier for me to order food online than to call a restaurant. Chat bots can only really search a database of documentation and frequently asked questions. Making one that has the benefits of talking to a human might be tantamount to AGI
- bentcorner 5y ago> it's much easier for me to order food online than to call a restaurant. I feel like this is where bots would do well - you can say "order me a burger with extra mayo and fries for pickup at 5pm" and it should negotiate all the minutiae for you. Doing this all manually requires a bunch of menu navigation. Maybe a phone bot is still a bad fit but doing something like this using your on-phone voice assistant or typing it into a text window feels reasonable.
- ShamelessC 5y agoHuman customer support is a regularly annoying experience, even (particularly?) when done online. There's probably always going to be some level of animosity towards it.
- deleted 5y ago[deleted]
- malaya_zemlya 5y agothere's a whole dark art of writing prompts for chat ai in order to make it behave in a sensible manner. the reason is that gpt doesn't have any context at all besides whats i n the supplied text. If you don't tell it exactly what to do, it will guess randomly. For example this chat prompt gives much more matter of fact answers, in my testing: "the following is a conversation with an AI assistant. the assistant is helpful, clever and friendly. it uses Wikipedia as the reference. Human:Hi! AI:Hi! Human:<your question goes here>"
- amelius 5y agoDidn't Google have a really great robocall demo, some time ago?
- dang 5y agoThis article is a response to this one: Chatbots: Still dumb after all these years - https://news.ycombinator.com/item?id=29825612 https://news.ycombinator.com/item?id=29825612 - Jan 2022 (408 comments) (Thanks everyone who pointed this out.)