13 ms·
The problem here is where these actions come from. Generic LLM cannot generate correct actions in many (if not most) real life cases. So, it will have to learn,
by two_in_one 3y ago
The problem here is where these actions come from. Generic LLM cannot generate correct actions in many (if not most) real life cases. So, it will have to learn, and LLMs aren't good at learning. For example: "I'm tired, play my favorite". The action depends on _who_ is saying and on what's going on right now. There may be someone sleeping, or watching TV. I'm afraid that acceptable solution is much more complicated.
- pmx 3y agoDO we have to expect _that_ level of understanding from the agent, though? If my wife said that to me, I may have a good chance of queuing up the song she has in mind, but anyone else? No chance. I don't expect tools like this to be able to understand cryptic requests and always come to the right answer. I'm happy if I can request a song or an action, or anything else in the same way i might ask another human who doesn't know me intimately.
- swexbe 3y agoIf not, how is this more useful than something like Siri?
- famouswaffles 3y agoNatural language understanding. Siri doesn't get context at all. You can twist unstructed data or requests however you like and the LLM will deal with it just fine. "Play my favorite" is just a knowledge problem. If GPT fails there, it's because it doesn't know your favorite, not because it can't parse the request or understand what you need it to do. You have to speak certain ways to Siri to get it to do things. Unless specifically hard-coded, Siri will never receive "damn I'm finding it hard to read" as input as decide to turn on the lights. GPT will.
- troupo 3y ago"Siri doesn't get context at all." and yet immediately "GPT fails there, it's because it doesn't know your favorite" "Knowing your favorite" is the context. > Unless specifically hard-coded, Siri will never receive "damn I'm finding it hard to read" as input as decide to turn on the lights. GPT will. Of course it won't. You have to very specifically fine tune it to understand what light conditions are, where you are in the house, and what it is you need to turn on.
- circuit10 3y agoThey mean context in the sentence or that or that be inferred from “common sense” and without any specific knowledge, I think
- troupo 3y agoI hate Siri as much as anyone, but Chat GPT has no context in the "common sense" either. The sibling comment literally says "I had to provide a long-ish sentence as a context/programming instructions before it could do anything". https://news.ycombinator.com/item?id=37464563 https://news.ycombinator.com/item?id=37464563
- circuit10 3y agoYou expect it to know things like that the user wants it to act as a home assistant without being told? That’s not common sense, that’s mind reading
- troupo 3y agoNo, I expect people to stop pretending that LLMs somehow know context unlike the stupid Siri. There's considerably more in that prompt besides just "you need to act like a home assistant"
- omniglottal 3y agoThere was insufficient context. Imagine I tell you "turn on that light, where I'm pointing". You'd do no better. No one here is under the conviction magical prescience is involved. This tooling provides the mechanism for an initial API call to be tied to the event described, in natural language, as "look where I'm pointing". The first response (to ask for clarification) is precisely what a human agent would do to get context to clarify the coarse-grained request. The second guess, assuming you disabled the (explicit) allowance for clarifying questions, is also a magnificent recognition of implicit, common-sense context. Seems it's even more effective than you at following the true context for this tools appropriate placement.
- jakderrida 3y agoHuman: "damn I'm finding it hard to read" GPT w/ memory: "Because you still have dyslexia."
- ftkftk 3y agoThanks for the chuckle.
- eternityforest 3y agoWhy would we want this at all if it doesn't know you that well? Current voice assistants without AI can already handle songs and actions like that. Seems like it's largely solved.
- pplonski86 3y agoI think this can be easily fixed, if LLM can do notes on what's going on. If it will have additional context before the prompt: ``` You are home assistant. Here is information what's going on in the house: It's 4PM. Bob likes Chopin Fantaisie-Impromptu. Alice likes Mozart Rondo in D. Bob is in the house. Alice will be back from office at 5PM. You get a prompt: I'm tired, play my favorite ``` For the above input any LLM will play Chopin.
- troupo 3y agoWhere is that input coming from?
- pplonski86 3y agoIt's just example, I've manually created it. But I think LLM can do such memory notes for itself and include as context.
- koolba 3y agoMemory notes by an LLM for its own consumptions reminds me of the Polaroids in the movie Momento.
- wizzwizz4 3y agoAt least those were deliberate.
- pplonski86 3y agoNice analogy - yes, something like this. What is more, LLM notes can be hierarchical, to have some kind of generalization.
- troupo 3y agoSo, an LLM that has no context, and must have context provided to it via notes, prompts etc. will somehow create these notes for itself?
- selalipop 3y ago> So, it will have to learn, and LLMs aren't good at learning LLMs are bad at human-like learning, but their zero-shot performance + semantic search more than make up for it. If you give an LLM access to your Spotify account via an API, it has access to your playlists and access to details about each song like `BPM`, `vocality`, even `energy` : https://developer.spotify.com/documentation/web-api/reference/get-list-users-playlists https://developer.spotify.com/documentation/web-api/referenc... https://developer.spotify.com/documentation/web-api/reference/get-audio-features https://developer.spotify.com/documentation/web-api/referenc... An LLM with no prior explanation of either endpoint, can figure out that it should look at your favorites playlists, and find which songs in your favorite list are most suitable for a tired person. - But it can go even further and identify its own sorting criteria for different situations with chain of thought: Bedroom at night: https://chat.openai.com/share/6b1787ef-cd84-4834-b582-5024f8f1c6a8 https://chat.openai.com/share/6b1787ef-cd84-4834-b582-5024f8... Kitchen at 5pm: https://chat.openai.com/share/7ddaa047-0855-48c1-bcea-3080833a670e https://chat.openai.com/share/7ddaa047-0855-48c1-bcea-308083... Rather than blindly selecting the most relaxing songs it understands nuance like: > Room State: "lights on" and "garage door open" can imply either returning home from work or engaging in some evening activity. The environment is probably not yet set for relaxation completely. And genuinely comes up with an intelligently adapted strategy based on the situation - And say it gets your favorite wrong, and you correct it: an LLM with no specialized training can classify your follow up as a correction vs an unrelated command. It can even use chain-of-thought to posit why it may have been wrong. You can then store all messages it classified as corrections and fetch those using semantic similarity. That addresses both the customization and determinism issues: You don't need to rely on the zero-shot performance getting it right every time, the model can use the same chain of thought to translate past corrections into future guidance without further training. For example, if your last correction was from classical music to hard metal when you got back from work, it's able to understand that you prefer higher energy songs, but still able to understand that doesn't mean every time in perpetuity it should play hard metal Kitchen w/ memory: https://chat.openai.com/share/43635427-55d5-4394-b282-46acae5a100d https://chat.openai.com/share/43635427-55d5-4394-b282-46acae... Bedroom w/ memory: https://chat.openai.com/share/8c146dd5-2233-4aba-8f6a-b97b7a778272 https://chat.openai.com/share/8c146dd5-2233-4aba-8f6a-b97b7a... - I experimented heavily with things like this when GPT came out; part of me wants to go back to it since I've seen shockingly few projects do what I assumed everyone would do. LLMs + well thought out memory access can do some incredible things as general assistants right now, but that seemed so obvious I moved on from the idea almost immediately. In retrospect, there's an interesting irony at play: LLMs make simple products very attractive. But if you embed them in more thoroughly engineered solutions, you can do some incredible things that are far above what they otherwise seem capable of. Yet a large number of the people most experienced in creating thoroughly engineered solutions view LLMs very cynically because of the simple (and shallow) solutions that are being churned out. Eventually LLMs may just advanced far enough that they bridge the gap in implementation, but I think there's a lot of opportunity left on the table because of that catch-22
- regularfry 3y agoI'm genuinely not seeing a problem there that the Planner part of the paper couldn't cover. "Who said that" and "what's going on right now" are just API calls. Besides which, if one person says "play my favourite" while another person is watching TV, that's not the LLM's job to unpack. The point is that the ability to call APIs gives them the ability to learn so that the actions that are eventually taken are correct in context. It's like a more generic version of https://code-as-policies.github.io/ https://code-as-policies.github.io/.
- jsemrau 3y agoThe whole notion of "memory" in LLM research solves this problem.
- liampulles 3y agoI have investigated use of agents for real support agent type work and the rate of failure made it unacceptable for my use case. This is even after giving it very explicit and finely tuned context. I suspect that if engineering of LLM solutions utilizes unseen testing data more, it's going to become apparent that it really does not have sufficiently reliable "cognitive" ability to do any practical agent type work.
- pplonski86 3y agoHave you used any benchmark to test agents? I'm currently looking for REST API usage benchmark for LLMs.
- powerapple 3y agohopefully it can be solved with the target API, the target API knows who is calling this API, the service has user information. Or this will be translated into "Play the most played playlist", and the action will be enough. I agree with you in general though, more useful AI is, more data it will need to see. I strongly believe companies like Microsoft, Google or Apple will bring the best experience because they own operating systems. It is going to be very hard for a third party to build a general AI assistant.