5 ms·
I don't buy it. LLMs cannot do anything reliably, no matter how constrained the domain. Their outputs are of acceptable quality when back to a person who will u
by aofjfdgnionio 3y ago
I don't buy it. LLMs cannot do anything reliably, no matter how constrained the domain. Their outputs are of acceptable quality when back to a person who will use their human brain to paper over the cracks. People can recognize when the output is garbage, figure out minor ambiguities, and subconsciously correct minor factual or logical errors. But I would never feed LLM results directly into another computer program This rules out most traditional NLP tasks.
- benreesman 3y agoI’m sympathetic to the instinct to push back on the absurd boosterism (these things are an existential threat to humanity this year), it’s fucking annoying. But they can do plenty of useful stuff reliably. It’s not “be generally intelligent”, which they are just nothing even remotely close to, but know you don’t dig the LLM hype from that comment? Yeah, they get that every time.
- aofjfdgnionio 3y agoI have tried to use LLMs, namely GPT4 and Llama-2, for sentiment analysis. They did quite poorly. I asked them to identify which sentiments from a list are found in a given text, with output formatted as a comma-separated list. In response I usually got just that. But sometimes I got a prose explanation of the sentiment, a list containing different keywords than requested, a list formatted differently than I wanted, or nothing useful at all. Sentiment analysis is easy. It wasn't the end goal, just a "hello world" example to verify my tools were set up correctly. I ran into unsolvable problems in the tutorial. I have no use for tools which do amazing things sometimes but which cannot be reasoned about and cannot be prevented from producing garbage. Maybe other people will find uses for them, though. I'll keep an open mind and check back in five years.
- FeepingCreature 3y agoWould you feed human output directly into a computer program? I'm just saying, we invented backspace for a reason. LLMs have no backspace. It's insane they work as well as they do.
- benreesman 3y agoSo your other reply got flagged which I thought was a little harsh (I mean you were pushing it but who am I to talk). If you’re not convinced about sentiment analysis on e.g. LLaMA 2, I think you’re wrong, but maybe I’m wrong. If you’re up for it, let’s run an expedient, I’ve got a GPU or two in my living room. This thread seems like a pretty great test set actually. Maybe we both learn something richer than some benchmark stat?
- YeGoblynQueenne 3y agoI too thought flagging the comment was a bit too harsh so I tried vouching for it, but it didn't get resurrected. I note that the comment is [dead] not [flagged] [dead], so maybe its state has to do with something else than the content of the comment? Just [dead] is, I think, shadowban. I checked the poster's comments, but since it's a new account there's very few of them and I can't determine the reason for the [dead] from them.