5 ms·
I agree. I think apps that would initially benefit from LLM-powered conversational interfaces are those that have the following traits: - constrained context
by mbil 3y ago
I agree. I think apps that would initially benefit from LLM-powered conversational interfaces are those that have the following traits:
- constrained context
- part of a hands-free workflow
A couple use-cases I have been pondering are driving assistant and cooking assistant.
People are already used to using their phone or car's nav system to give them directions to an unfamiliar place. But even with such a system it's useful to have a human navigator in the car with you to answer various questions:
- What's my next turn again?
- How long till we get there?
- Are there any rest stops near here?
- What was that restaurant we just passed?
- Is there another route with less traffic?
These questions are all answerable with context that can be provided by the mapping app:
- List of upcoming directions
- Overall route planning
- Surrounding place data
- Traffic data and alternate route information
It's possible to pull over to the side of the road, take off your distance glasses, put on your reading glasses, and zoom/pan the map to try to answer these questions yourself. But if the map application can just expose its API to the language interface layer, then a user can get the answers without taking their eyes off the road.
The information is contextual and constrained based on a current task. In some cases it might be more desirable to whip out your phone and interact with the map to look up the answers on a screen, but often it won't be worth stopping the car, and so the conversational interface is better.
Cooking assistant is a similar case: you are busy stirring something and checking on the oven -- you don't want to wipe the flour off your hands to pick up your phone and ask how many teaspoons of sugar you need. Again: contextual and constrained info based on a current task, and your hands and eyes -- the instruments of traditional UIs -- are otherwise occupied.
Today, our software interfaces generally have one of two kinds of entity on the other end: humans, or other software. In the near future there will be another type of entity: language models. We need to start thinking of how our APIs will change when they're interacting with an LLM -- e.g. they'll need to be discoverable and self-describing; error states will need to be standardized or explicit with instructions on how to correct; they'll need to be fast enough to fit in a conversational interface; etc. It's arguable that such traits are part of good API design today, but in the future they may be required for the API to function in a landscape of virtual agents.
- RandomLensman 3y agoIn the cooking example, you either need the AI to have full awareness of the step you are at or you need to describe the step you are at, which could be cumbersome ("I did ..., how much sugar do I need now"). I venture, having the recipe projected in front of you would be much faster.
- travoc 3y agoand a piece of paper wins again.
- mbil 3y agoI imagined the AI would be reading the steps aloud to you, and so would be aware of your progress. I don’t think an AI assistant precludes the recipe being projected tho, just as in the driving example it wouldn’t replace an on screen map.
- troupo 3y agoHaving it both in front of my eyes, and being able to get answers to questions like "I've added the eggs, now what?" or "what does folding a dough mean?" at the same time would be very valuable.