6 ms·
> So, it will have to learn, and LLMs aren't good at learning LLMs are bad at human-like learning, but their zero-shot performance + semantic search more than
by selalipop 3y ago
> So, it will have to learn, and LLMs aren't good at learning
LLMs are bad at human-like learning, but their zero-shot performance + semantic search more than make up for it.
If you give an LLM access to your Spotify account via an API, it has access to your playlists and access to details about each song like `BPM`, `vocality`, even `energy` :
https://developer.spotify.com/documentation/web-api/reference/get-list-users-playlists https://developer.spotify.com/documentation/web-api/referenc...
https://developer.spotify.com/documentation/web-api/reference/get-audio-features https://developer.spotify.com/documentation/web-api/referenc...
An LLM with no prior explanation of either endpoint, can figure out that it should look at your favorites playlists, and find which songs in your favorite list are most suitable for a tired person.
-
But it can go even further and identify its own sorting criteria for different situations with chain of thought:
Bedroom at night: https://chat.openai.com/share/6b1787ef-cd84-4834-b582-5024f8f1c6a8 https://chat.openai.com/share/6b1787ef-cd84-4834-b582-5024f8...
Kitchen at 5pm: https://chat.openai.com/share/7ddaa047-0855-48c1-bcea-3080833a670e https://chat.openai.com/share/7ddaa047-0855-48c1-bcea-308083...
Rather than blindly selecting the most relaxing songs it understands nuance like:
> Room State: "lights on" and "garage door open" can imply either returning home from work or engaging in some evening activity. The environment is probably not yet set for relaxation completely.
And genuinely comes up with an intelligently adapted strategy based on the situation
-
And say it gets your favorite wrong, and you correct it: an LLM with no specialized training can classify your follow up as a correction vs an unrelated command. It can even use chain-of-thought to posit why it may have been wrong.
You can then store all messages it classified as corrections and fetch those using semantic similarity.
That addresses both the customization and determinism issues: You don't need to rely on the zero-shot performance getting it right every time, the model can use the same chain of thought to translate past corrections into future guidance without further training.
For example, if your last correction was from classical music to hard metal when you got back from work, it's able to understand that you prefer higher energy songs, but still able to understand that doesn't mean every time in perpetuity it should play hard metal
Kitchen w/ memory: https://chat.openai.com/share/43635427-55d5-4394-b282-46acae5a100d https://chat.openai.com/share/43635427-55d5-4394-b282-46acae...
Bedroom w/ memory: https://chat.openai.com/share/8c146dd5-2233-4aba-8f6a-b97b7a778272 https://chat.openai.com/share/8c146dd5-2233-4aba-8f6a-b97b7a...
-
I experimented heavily with things like this when GPT came out; part of me wants to go back to it since I've seen shockingly few projects do what I assumed everyone would do.
LLMs + well thought out memory access can do some incredible things as general assistants right now, but that seemed so obvious I moved on from the idea almost immediately.
In retrospect, there's an interesting irony at play: LLMs make simple products very attractive. But if you embed them in more thoroughly engineered solutions, you can do some incredible things that are far above what they otherwise seem capable of.
Yet a large number of the people most experienced in creating thoroughly engineered solutions view LLMs very cynically because of the simple (and shallow) solutions that are being churned out.
Eventually LLMs may just advanced far enough that they bridge the gap in implementation, but I think there's a lot of opportunity left on the table because of that catch-22
- troupo 3y ago> Yet a large number of the people most experienced in creating thoroughly engineered solutions view LLMs very cynically because of the simple (and shallow) solutions that are being churned out. Maybe, just maybe, because even simple solutions are invariably an incomplete brittle complicated unpredictable mess that you can't use to build anything complex with? As eloquently demonstrated by your "simple" solutions
- selalipop 3y agoYour reply is not indicative of someone capable of a good faith conversation on the topic, but I'll bite. I think you don't understand what the hard and easy problems are that underly the solutions I'm talking about. For example: you repeatedly reply to people talking about the length of the prompts, but end users don't need to write prompts. It's trivial to append instructions around what a user says. On the other hand, you keep replying to people with "how is that not just Siri" when people describe the LLM demonstrating zero-shot classification for example, but you don't seem to understand how difficult of a problem that has been for ML. Those contrived chat logs you see are demonstrating multiple discrete classifications that would have each cost untold hundreds of thousands of dollars in development of recommender systems to replicate just a few years ago. — Most people couldn't even dream of building a Spotify song recommender from first principles that could capture nuance like that chat demonstrated with an army of engineers. The fact is today, right now, that's something someone could hack into a real usable personal recommender in a weekend. At the end of the day LLMs don't make all problems easier, and they make some problems harder: but the problems they make easier are extremely hard problems. I think if you're not familiar with how hard some of the things they're doing are, then the things they're doing poorly glare out much brighter. If you spend half that weekend is spent fighting the LLM to output JSON the right way, it sure sounds like LLMs are just dumb hype machines... but it doesn't reflect the sheer impossibility of the value they're providing within that same system.
- troupo 3y ago> Your reply is not indicative of someone capable of a good faith conversation on the topic, but I'll bite. You think so because replies to me have willfully ifgnored and misunderstood the point of my replies. And have willfully ignored the context (which, as I already said, is funny and ironic). The whole discussion started with - "LLMs can't generate actions in real life situations" - "We can't expect LLMs to do that" - "How is it more useful than Siri" - and here's the most important one: "Siri doesn't have context ... GPT fails there, it's because it doesn't <know context>" So, Siri is bad, because it doesn't have context. But somehow even though GPTs are the same, they are good because... someone somewhere can come up with an imprecise unpredictable prompt for a rather specific situation tat may or may not work for some people... and that's why they are better than Siri and have context. "Where is this context/input coming from?" - "end users don't need to write prompts. It's trivial to append instructions around what a user says." This is literally magical thinking. "Someone somehwere will maybe somehow create a proper prompt that maybe will definitely work, and users won't have to do anything". This... is literally Siri. It even asks for clarifications when it can't understand something. You keep harping on about "zero-shot classification". And completely ignored what I wrote: I ran your amazing zero-shot classification, and it immediately failed. It raised the brightness in the garage. I guess someone (not the end user) should write another model to correct the first one. And when that one inevitably, and immediately, fails, someone (not the end user) should trivially write corrections for that. It's all turtles all the way down, isn't it? (On a second try it did say that the user is likely in the kitchen or in the bathroom, and increased brightness in the bathroom). Thing is: I don't subscribe to this magical thinking. I see innumerable failure modes and "edge cases" (which are not edge cases, but actual every day scenarios) where none of this works. This is also the reason why we haven't seen any complex product (apart from specialised fine-tuned ones) built with LLMs: they fail very much like Siri does in even the simplest scenarios. No one knows how to provide an actual proper context of a person's life so that it works reliably more than half of the time (and when it seemingly works, a simple MRU would probably work better). The "trivial annotations" for user actions are anything but. (There's also a separate discussion here: https://news.ycombinator.com/item?id=37464277 https://news.ycombinator.com/item?id=37464277). > Most people couldn't even dream of building a Spotify song recommender from first principles that could capture nuance like that chat demonstrated with an army of engineers. The fact is today, right now, that's something someone could hack into a real usable personal recommender in a weekend. As an engineer who works at Spotify (not in recommendations, but I know the details at least superficially), thank you for a hearty laugh this sentence brought me. As I said, magical thinking.