3 ms·
Your reply is not indicative of someone capable of a good faith conversation on the topic, but I'll bite. I think you don't understand what the hard and easy p
by selalipop 3y ago
Your reply is not indicative of someone capable of a good faith conversation on the topic, but I'll bite.
I think you don't understand what the hard and easy problems are that underly the solutions I'm talking about.
For example: you repeatedly reply to people talking about the length of the prompts, but end users don't need to write prompts. It's trivial to append instructions around what a user says.
On the other hand, you keep replying to people with "how is that not just Siri" when people describe the LLM demonstrating zero-shot classification for example, but you don't seem to understand how difficult of a problem that has been for ML. Those contrived chat logs you see are demonstrating multiple discrete classifications that would have each cost untold hundreds of thousands of dollars in development of recommender systems to replicate just a few years ago.
—
Most people couldn't even dream of building a Spotify song recommender from first principles that could capture nuance like that chat demonstrated with an army of engineers. The fact is today, right now, that's something someone could hack into a real usable personal recommender in a weekend.
At the end of the day LLMs don't make all problems easier, and they make some problems harder: but the problems they make easier are extremely hard problems. I think if you're not familiar with how hard some of the things they're doing are, then the things they're doing poorly glare out much brighter.
If you spend half that weekend is spent fighting the LLM to output JSON the right way, it sure sounds like LLMs are just dumb hype machines... but it doesn't reflect the sheer impossibility of the value they're providing within that same system.
- troupo 3y ago> Your reply is not indicative of someone capable of a good faith conversation on the topic, but I'll bite. You think so because replies to me have willfully ifgnored and misunderstood the point of my replies. And have willfully ignored the context (which, as I already said, is funny and ironic). The whole discussion started with - "LLMs can't generate actions in real life situations" - "We can't expect LLMs to do that" - "How is it more useful than Siri" - and here's the most important one: "Siri doesn't have context ... GPT fails there, it's because it doesn't <know context>" So, Siri is bad, because it doesn't have context. But somehow even though GPTs are the same, they are good because... someone somewhere can come up with an imprecise unpredictable prompt for a rather specific situation tat may or may not work for some people... and that's why they are better than Siri and have context. "Where is this context/input coming from?" - "end users don't need to write prompts. It's trivial to append instructions around what a user says." This is literally magical thinking. "Someone somehwere will maybe somehow create a proper prompt that maybe will definitely work, and users won't have to do anything". This... is literally Siri. It even asks for clarifications when it can't understand something. You keep harping on about "zero-shot classification". And completely ignored what I wrote: I ran your amazing zero-shot classification, and it immediately failed. It raised the brightness in the garage. I guess someone (not the end user) should write another model to correct the first one. And when that one inevitably, and immediately, fails, someone (not the end user) should trivially write corrections for that. It's all turtles all the way down, isn't it? (On a second try it did say that the user is likely in the kitchen or in the bathroom, and increased brightness in the bathroom). Thing is: I don't subscribe to this magical thinking. I see innumerable failure modes and "edge cases" (which are not edge cases, but actual every day scenarios) where none of this works. This is also the reason why we haven't seen any complex product (apart from specialised fine-tuned ones) built with LLMs: they fail very much like Siri does in even the simplest scenarios. No one knows how to provide an actual proper context of a person's life so that it works reliably more than half of the time (and when it seemingly works, a simple MRU would probably work better). The "trivial annotations" for user actions are anything but. (There's also a separate discussion here: https://news.ycombinator.com/item?id=37464277 https://news.ycombinator.com/item?id=37464277). > Most people couldn't even dream of building a Spotify song recommender from first principles that could capture nuance like that chat demonstrated with an army of engineers. The fact is today, right now, that's something someone could hack into a real usable personal recommender in a weekend. As an engineer who works at Spotify (not in recommendations, but I know the details at least superficially), thank you for a hearty laugh this sentence brought me. As I said, magical thinking.
- selalipop 3y agoYou had a chance to prove my assumption wrong by writing this same exact comment without all the snark. At the end of the day if you're just unmoved by the implications that an ML model went from a bag of tokens to a structured, explained chain of thought, and a final response on an unknown task with rewards defined in natural english (!) and intentional ambiguity most humans wouldn't even try to confront... there's not much conversation to be had. I think the rest of us (including your colleagues) will continue to build on these models, and like most advancements there'll be a vocal crowd insisting the car isn't useful because it can't be fed with grass. > not in recommendations You didn't have to say that after complaining ChatGPT's web interface didn't give both us the same reply (most people in ML understand how temperature relates to LLM output) _ By the way, if making your own personal music recommender seems like "magical thinking", maybe you're a little lost on which parts of Spotify's recommender systems are complex due to scale: if Spotify only needed to make song selection work for one person at a time, they'd have a lot more leeway in architecture. In fact, when the problem was flipped and they were scaling humans attaching their recommendations to songs, they built on OpenAI: https://newsroom.spotify.com/2023-02-22/spotify-debuts-a-new-ai-dj-right-in-your-pocket/ https://newsroom.spotify.com/2023-02-22/spotify-debuts-a-new... Not unexpected when you're a founding organizer of the "NLP 4 Music and Audio Forum".