21 ms·
Natural language is an unnatural interface
- deleted 3y ago[deleted]
- blululu 3y agoStrong disagree. These are straw man arguments about the limitations of models. Senior executives, politicians and people in high leadership positions can see their wildest dreams realized using little more than natural language and discussions with the world’s leading experts. Satya Nadella or Joe Biden don’t enjoy and special abstractions and yet they are able to get everything they want using natural language.
- esafak 3y agoThe ability to express oneself through natural language is tangentially related to its suitability as the foundation of a user interface. A natural language should let you express your thoughts effectively. A user interface should let you accomplish the task you want effectively, within the limited capabilities of the program you are using. And it should expose those limits, not make you discover them through trial and error, as opaque, conversational interfaces do.
- hLineVsLineH 3y agoOr maybe they use natural language because they have no choice. If Nadella or Biden could just push some buttons to make things happen, they probably would. I know I would.
- deleted 3y ago[deleted]
- m3kw9 3y agoIt can only be good when the interface already knows a massive amount of context. For example the LLM in a future use case knows all your email and you just ask it to draw a time line of this thread and etc. If you are asking it to do stuff from scratch like most of us do on ChatGPT, it’s quite a pain
- biql 3y agoExactly. It doesn't need extra buttons in the UI, it needs a giant context window to contain all the important info about your life or a business and the tools for gathering and maintaining that context. Then merely saying that you are seek becomes enough to autocomplete the rest of the actions including sending an email.
- esafak 3y agoI don't want it to know everything about me unless it runs locally.
- deleted 3y ago[deleted]
- Kostchei 3y agoMost of the arguments I hear from my mates-who-code about adding deterministic old school things like buttons and "ui" on top of LLMs is "sometimes it's faster and more specific"... Ok, sure, but go to voice to text, and give the llm in a voice version, to a grandma. They can just use it. and if you, the programmer need specifics "I want to do root cause analysis of these 20 incidents, break them up into mobile and web app, then draw common threads in each type of RCS".. then write/say a better prompt that exactly details what you want. Might take 2 or 3 goes, but you will get there. Adding metadata tabs like some vector-db franken-sharepoint seems like a step back. At least from a Rich Sutton world view. Let the LLM work it out- and if it fails, improve it for a many uses cases, not just one. IMHO (tldr- strong disagree)
- Terr_ 3y ago> unnatural interface A simple example might be the problem of "pick a color". Even the best natural-language interface is going to suck about as much as if you're trying to ask another human to do it for you, even if that assistant is capable of displaying 1-5 color swatches in their replies. Instead of just seeing the entire palette and choosing, you need to say "I want a gold color", "lighter than that" and "darker" and "less like urine" etc. > People fundamentally don’t know how to prompt. There are no better examples than Stable Diffusion prompts. You know, this reminds me of the Good Old Days of internet search engines, where a little expertise in choosing terms/operators was very powerful, before the advanced-case was cannibalized to help the average-case.
- iamflimflam1 3y agoYet this is how we communicate with designers and end up with good results.
- Terr_ 3y agoI think that confuses doing versus delegating. Delegation is easy to do via a text-box, because you're just kicking the interactive complexity-can down the road to someone else, often in a way which can be problematic even with actual humans. For example, a project-manager or executive could verbally delegate "make a new registration page for the site" and "needs more rounded corners", either to an AI or to an employee or offshore contractor. However that's not the same as trying to program exclusively by typing (or dictating) prose to a text-box. ("Page down more. Go to method pee-reg underscore apply. Show me its caller methods. Go to caller method two. Type the following into line 7 position 43...") There might be some parallels we can draw with the last few decades of "programming business logic will be replaced by drawing diagrams" predictions.
- totetsu 3y agoI have been using Dall-e, and testing the dall-e chatgpt plugin. Even tough both are supposedly natural language interfaces, I find I approach the dall-e prompt more like writing a formula than real language. Using the gpt plugin is like delegating to a designer to write the prompt for you. personally I don't like the results of that compared to what I would make myself.
- azhenley 3y agoVery relevant to my post from January, “Natural language is the lazy user interface”. https://austinhenley.com/blog/naturallanguageui.html https://austinhenley.com/blog/naturallanguageui.html
- afavour 3y agoReminds me of the boom in voice assistants, when we were told it was the interface of the future. I’d ask my Google Home what the temperature was going to be today. It would tell me. I’d ask what the temperature was yesterday. It would tell me it didn’t understand the question. ChatGPT etc obviously aren’t quite that bad but the core problem remains discoverability. It isn’t at all clear what these interfaces can and cannot do, all you can do is keep trying different things until you get a result.
- pprotas 3y agoI ask my HomePod if it’s going to rain tonight and it tells me that tomorrow is sunny…
- dustrider 3y agoAdd to this that voice assistants fundamentally got worse and worse as their makers tried to monetise, increase usage etc. it took cool tech and made it annoying. I can see llms going much the same way.
- kaba0 3y agoFunny anecdote, but I have literally have started programming by doing “smart” assistants like the former, and I don’t think they were much worse. What I did was to lemmatize the words (my native tongue is agglutinative so it was somewhat harder than with English), and simply look at a fixed set of “commands” like “play it”, and pass the rest of the words as parameters when needed (I searched youtube for a video in this case). Unfortunately I have seem to lost these “beautiful php” codes, even though I would be very curious how bad that code was :D
- TeMPOraL 3y agoThe best voice assistant I ever had was the one I DIY-ed around 2007, using Microsoft Speech API and a cheap piezoelectric mike I soldered to a long cable, hung off the wardrobe, and plugged into PC. The code itself was a mashup of MS SAPI demo of "controlled language" interface and some tutorials for how to control WinAMP with WM_USER messages in WinAPI. I designed a little tree of commands, maybe 3 level deep, wrote the magic XML for it, some trivial C++ logic for driving the voice recognizer and reacting to identified commands (including one that held the recognizer two levels deep in the command tree, so I could issue multiple commands from a subtree without having to repeat two extra words for each). The result of this couple afternoons of working on this (instead of learning for my maturity exams, as I was supposed to), was a system that, I kid you not, was more reliable and delivered more value to me than any of the current voice assistants. For one, its recognition was flawless. The typical interaction would look like: $ Computer! > <appropriate beep from Star Trek: TNG, because of course doing this was 90% of my reason for building the program> $ Music, Playlist Alpha > <appropriate confirmation beep, WinAMP begins to play> I had commands for usual play/pause/resume, next/previous, four playlist (alpha through delta), and volume control at different granularity ("mute", "one quarter", "two quarters", "three quarters", "full", plus "louder" and "quieter" for IIRC +/- 5% or +/- 10% jumps). Plus some stubs for non-music thing that IIRC I never eventually implemented. Here's the thing: it worked flawlessly. It heard me across the room. It heard me through music so loud that it was uncomfortable to talk in. It never self-triggered (except that one person who managed to make a swear word be read as the wake word, a single case out of many who tried). It worked fast - I could complete the whole command chain in less time than Google Assistant takes to start listening after "OK Google". The secret? Constrained grammar and training. In order to use speech recognition in Windows back then, you had to turn it on and let it analyze a sample of your voice (offline! those were the days!), based on a recording of you reading some calibration text it gave you. This process was additive - you could repeat it to improve recognition accuracy. But a little known fact was that you could also supply your own text - and that was the other half that made the magic happen. I created myself a training text, consisting of individual command words and their sequences, and trained the Windows speech recognition on it multiple times, under varying conditions. Specifically, I run: {three locations in the room} x ({no background} + ({classical music, pop music, whatever was on FM radio} x {quiet playback, normal playback, very loud playback})) training sessions. That's 30 sessions of repeating the same text. Each one took maybe a minute or less, so I was done with it in about an hour. And after that training, no matter where I was standing in the room and what I was doing, the voice control system worked with near-zero false positives and near-zero false negatives. I say "near" because I had maybe two or three cases of each, over months of continued use. And yes, I could play music so loud you couldn't talk in the room, and I could scream out commands, outshouting the music, and it would work. Try that with Google Assistant. To recap: I had a system I hacked together in couple evenings, whose software was a relatively small tweak to a default example project (but done with love!) and hardware was hand-soldered from cheapest, locally-sourced parts, that did everything I wanted from a voice assistant, did it flawlessly, much faster than any of the voice assistants on the market today, completely off-line, in 2007, on a mid-range PC, without noticeably taxing its resources. This is why I occasionally rant that voice assistants are bloated and done backwards - all because they're designed to suit vendor needs first, user needs second. -------- But hey, I know a way Google, Apple, Samsung (!) et al. could fix the shitty performance of their voice assistants and dictation software. They need to fine-tune a LLM on a dataset made of target words/sentences, and transcripts of them being misheard in great many ways. Then they need to feed the output of their voice-to-text pipeline through that LLM, so it can correct the text wholesale. That, or maybe, you know, do whatever Microsoft was doing in 2007 that made dictation work well and offline.
- Barrin92 3y agoI think on the input side natural language is a pretty reasonable interface. For me the output side is much more problematic. Natural language makes search output so uniform not only can't you tell whether something's real or not, if it is you can't tell whether it came from the Encyclopedia Britannica or the Youtube comment section. Taking the ability to discern sources yourself away turns me off from using these systems anywhere where that is relevant. Also of course not outputting structured data greatly diminishes having AI systems interface with other traditional automated systems. Natural language is not a great format for processing or analysis of any kind.
- brunorsini 3y ago"There is nothing, I repeat nothing, more intimidating than an empty text box staring you in the face." Talk about a hyperbolic opening line. Is it really that intimidating to have an empty text box on Whatsapp or your favorite SMS app? No, as you expect to have an appropriate response coming from the other side, pretty much regardless of what your input is. As a frequent user of ChatGPT, I've come to expect the same in there. And it works great, without me having to study any "prompt engineering". In fact, as it gets updated, I get frustrated less and less often — unlike my experience using Bard, which can be better for a few tasks but often returns opaque errors that do feel frustrating. The solution here is clearly for the model to improve, and one doesn't even need a leap of faith — just look at what OpenAI is already delivering! Talking to a competent LLM is nothing like talking to bash or dos. I also get frustrated when I sometimes have to ask for the same thing in a slightly different way... but that's still almost always faster than searching for the right button or submenu in most creation-oriented software. Whoever is waiting for Word or Google Docs to add a "write this in business-formal email tone" dropdown menu to the UI clearly hasn't grokked the true shift we're about to go through in computing. Incidentally, I am often using ChatGPT to help me do more advanced / rarely used tasks in software from Avid Pro Tools to Adobe Premiere. And I can't remember a single time when doing this was slower or more frustrating than reaching out to either Google or the software's own "help" section. Of course we'll have more input options. It makes tons of sense for things like image or video generation. I bet the models will also soon be outputting more and more "interactive elements" that will aid in refining results. But I have a feeling the opening text box (or, better yet, the open ears of a friendly audio assistant) is here to stay.
- 9dev 3y ago> or, better yet, the open ears of a friendly audio assistant It’s interesting you mention this. I’ve been wondering this for a while now - there have been made leaps recently in LLMs, speech synthesis and speech recognition. There are sophisticated language models, computer voices that are hard to distinguish from real humans, and software that can reliably understand even the worst recording of someone speaking. Yet still, those three components have not yet been integrated in a next generation Alexa yet. But why? It doesn’t even sound particularly complicated (on the scale of all the prior art necessary).
- specproc 3y agoSome odd Python "class" syntax in there. ## Python class called Car def Car(make, model, year): """This function creates a new car object""" car = {"make": make, "model": model, "year": year} return car
- deleted 3y ago[deleted]
- _nalply 3y agoI started reading and I found the text not engaging, perhaps it's me being not in the right mood, but while deciding to stop reading, Substack showed a popup. What the hell. Enough enshittification. I dropped the thing and started to complain here. (shrugs)
- mercurialsolo 3y agoNatural language is the ultimate programming language. We use it to program us humans all the time
- qsort 3y agoConsidering we move away from natural language whenever we have the chance to do so (many early formal languages were meant to target the humans who would in turn program computers!), I don't think this is really the case except for extremely trivial cases. Natural languages are terrible programming languages for many of the same reasons programming languages are terrible natural languages.
- checkyoursudo 3y ago> Natural language is the ultimate programming language. We use it to program us humans all the time Unreliably, with mixed results. Just yesterday, I said something loudly from my basement so that my family upstairs could hear, and the response was for them to ask why I was angrily yelling nonsense at them from the dungeon. YMMV.
- palata 3y agoNot at all. Programming languages are meant to be unambiguous. Natural languages are meant to be ambiguous. They just solve different problems.
- mercurialsolo 3y agoHuman programming is all about thriving in ambiguity
- nottorp 3y agoI suppose that's why legalese and even engineering have strictly defined terms to eliminate any ambiguity.
- seydor 3y agoThe problem is that most people don't know how to express what they want (and often they dont know what they want). But it's a bit myopic to stick to language interfaces, the LLMs are already integrated with images, and over time they will be integrated with entire GUIs.
- TalktoCrystal 3y agoMany companies laid off UX designer for the new AI trend
- inciampati 3y agoI'd like to take issue with the characterization of copilot as not needing textual prompting. To get it to work, you need to use comments. The more the better. You can put huge amounts of context and information in them. This involves writing and description. It's exactly the same as what the author is arguing against.
- monkeydust 3y agoWe have been using natural language interface where I work for 5+ years (pre LLM) and honestly if applied correctly it can be very sticky and effective. For example...Your application may have multiple capabilities serving multiple user types. Rather than smack an empty text box in the front page you can try embedding the box into a specific capability and develop parsers focused on the particular domain of that capability. This limits the scope and thus chances for not delivering on the users intention.
- gaazoh 3y agoWhile I agree with the title, I find it very lackluster that it entirely focuses on some specific AI interfaces. First of all, ChatGTP implies by its name that it's designed to chat. It's a sophisticated chatbot, but a chatbox nonetheless, you are not going to have a chat by filling a form. Then the examples given for StableDiffusion are not natural language. In this case, there can probably be a better interface than a single textbox, but the issue is the textbox, not that it expects natural language (it doesn't). Other types of interface for AI stuff do exist. Copilot is also an LLM, but it takes surrounding code as input, not a natural language prompt. Plenty (most) of models take whatever format is convenient for the task and output an adapted format (image/video/text classification, feature detection...). On the other hand, there are some interfaces that force natural language processing where it is one of the most unnatural and ineffective option, and no AI whatsoever is involved. Anyone who ever tried to book a train ticket in France in the past couple of years know what I mean[0]. Having to spell out an itinerary and date information is very, very confusing and error prone. [0] https://www.sncf-connect.com/ https://www.sncf-connect.com/
- mo_42 3y agoI agree with the general intention of the post. > Natural language is an unnatural interface However, I think the title is misleading. Maybe I am too much of a technical person and take interface too broadly. But natural language is probably the most natural interface as humans learn this really early and use it every day. Also, the word "natural" may not be suitable for this discussion. Maybe the author has this in mind already but I think we should rather talk about intuitive, efficient, motivating, … interfaces. In general, I guess HCI is not about natural but rather useful.
- deleted 3y ago[deleted]
- intrasight 3y agoThis article isn't about natural language or AI, IMHO. Typing into a keyboard is not natural. Adding buttons makes it even less so. We don't have AI avatars today. But we will within my lifetime. I'll be able to converse with most any historical figure in VR. We will see each others facial and hand expressions. Let's not go backwards and add buttons. Put your efforts into going forwards.
- ericol 3y agoAs opposed to what? I don't know about you, but for me on a personal level ChatGPT is a SERIOUS paradigm shift. I work on OSX, I use an app (MacGPT) that is always a keyboard shortcut away, and more often than not the responses are several times better than what Google will give me. I know there are a lot of areas with high friction, but those will go away eventually and compared to what I was doing before, the friction is much much less. Not to mention that it is a tool _I didn't have before_ and for me in some cases - specially those tasks that are unavoidable, you have to generally start from zero, and are a burden - my productivity has increased 3X I'll take an unnatural interface any day if that's the end result.
- parpfish 3y agofor an app like MacGPT, how much do you end up paying each month in API fees? I'd love to try it out, but I know that I'd constantly be second-guessing myself and saying "nah, i guess I don't really need to look that up" because I'd that pay-per-token API fee would always be sitting in the back of my mind.
- ericol 3y agoTBH I have no idea [1] but I have quite a low limit of expenses using the API (10 bucks) and I'm yet to reach it. Besides that, you can as well log in via web (That I haven't done, but it's possible) As a side note, even thought the friction might be too high I couple it with the same programmer's MacWhisper, so that when I need a large text I just speak to the computer. [1] Geez. I just checked, and my cumulative expenses for June are a staggering 0.27 USD
- parpfish 3y agoOof. By the end of the month that may climb all the way up to 0.30USD. /s
- niam 3y agoI don't particularly agree with the title but agree that there's a present awkwardness in the way that users are expected to derive correct insights from LLMs. In the same way that tokenization exists to overcome a would-be shortcoming of present computing resources and architectures, but may eventually become unnecessary: a more streamlined interface would be helpful in tiding us over this hump of awkwardness even if it too eventually becomes unnecessary.
- throwuwu 3y agoThe GUIfication of LLMs. Yuck. I’d better join the local model camp before bad ideas like this catch on and fully nerf this tech.
- paisible 3y agoI'm really confused by "Anecdotally, most people use LLMs for ~4 basic natural language tasks" and "Most LLM applications use some combination of these four". I'm not sure about the `ELI5` use-case, and feels like this is only true for a very limited type of use-cases people currently use LLMs for. For conversational FAQ-type use-cases like the ones described by OP perhaps a few basic prompts suffice (although anything requiring the agent to have "agency" in its replies would necessarily require prompt engineering) - but what about all the other ways that people can use LLMs to analyze, transform and generate data?
- FloatArtifact 3y agoYes, this is what also is frustrating about saying a few words for over the phone voice interactions to describe problems.