3 ms·
The problem is in remembering the voice commands. I could never do it. Word the command slightly “wrong” and it won’t work at all (at least not in my 2014 VW).
by rlpb 9mo ago
The problem is in remembering the voice commands. I could never do it. Word the command slightly “wrong” and it won’t work at all (at least not in my 2014 VW).
I’m optimistic that the latest progress in AI will fix this when the technology matures in cars. I reckon this is still a decade away though.
- ssl-3 9mo agoHonestly, other than that one single command ("Climate control defrost and floor") I never really use voice for anything else while actively driving. The temperature knob usually does what I want when driving, and I'll be stopped again soon enough if I want to fiddle with something else. And that one voice command is easy-enough to remember, and the resulting manually-selected mode is easy-enough to cancel with the Auto button (which is the entire middle of the temperature knob -- simple enough). AI is too easy to get wrong. For example: At home when my hands are full and I'm headed to/from the basement, I might bark out the command "Alexa! Basement lights!" This command sometimes results turning the lights on or off. But sometimes, it results in entering a conversation about the basement lights, when all anyone really wants from such simple diction is for the lights to toggle state -- like interacting with a regular light switch just toggles state. I simply want computers to follow instructions. I am very particularly disinterested in ever having conversation -- a negotiation -- with a computer in my car. But I can see plenty of merit to adding some context-aware tolerance for ambiguity to the accepted commands. Different people sometimes (quite rightly) use different words to describe the end result they want. That doesn't take an LLM to accomplish, I don't think. After all, a car has a limited number of functions. It should be mostly a matter of broadening the voice recognition dictionary and expanding the fixed logic to deal with that breadth. I reckon that this should have happened 5 years ago. :)
- rlpb 9mo ago> That doesn't take an LLM to accomplish, I don't think. After all, a car has a limited number of functions. It should be mostly a matter of broadening the voice recognition dictionary and expanding the fixed logic to deal with that breadth. I think the most effective way to get this accurate and effective is to give an LLM the user’s voice prompt and current context and ask to convert the user’s request into an API call. The user wouldn’t be chatting with the LLM directly. The point is that it doesn’t require a static dictionary to already have your exact phrasing and will just work with plain English.
- ssl-3 9mo agoThat requires either an online connection to a datacenter somewhere or something that (at present time) is a fairly substantial on-board computing system, and those are things that I think are worth trying to avoid for tasks like adjusting the HVAC in a Honda. Maybe some day. Right now, when we do have substantial on-board computing systems, they're trying to drive the car -- not change the temperature. Adding in an additional computational timeslot for LLM voice commands seems both foolhardy and expensive. Meanwhile: Broader dictionaries and static flows with greater breadth for voice recognition? We can do that right now. (We can even use LLMs to help generate the static flows, along with people to evaluate and test them. Once implemented, they become cheap to run. This is in-keeping with a fairly common theme here on HN: Don't use the bot to process the data. Instead, use the bot to write the data processor.)