4 ms·
> For instance, in customer service, an RL agent might discover that sometimes asking a clarifying question early in the conversation, even when seemingly obvio
by moojacob 2y ago
> For instance, in customer service, an RL agent might discover that sometimes asking a clarifying question early in the conversation, even when seemingly obvious, leads to much better resolution rates. This isn’t something we would typically program into a wrapper, but the agent found this pattern through extensive trial and error. The key is having enough computational power to run these experiments and learn from them.
I am working on a gpt wrapper in customer support. I’ve focused on letting the LMs do what they do best, which is writing responses using context. The human is responsible for managing the context instead. That part is a much harder problem than RL folks expect it to be. How does your AI agent know all the nuance of a business? How does it know you switched your policy on returns? You’d have to have a human sign off on all replies to customer inquiries. But then, why not make an actual UI at that point instead of an “agent” chatbox.
Games are simple, we know all the rules. Like chess. Deepmind can train on 50 million games. But we don’t know all the rules in customer support. Are you going to let an AI agent train itself on 50 million customer interactions and be happy with it sucking for the first 20 million?
- ip26 2y agoThe bitter lesson would suggest eventually the LM agent will train itself, brute force, on something and extract the context itself. Perhaps it will scrape all your policy documents and figure out which ones are most recently dated.