6 ms·
Despite the hype about LLMs, many of the answers are pretty terrible. The 12-bar blues progressions seem mostly clueless. The question is will any of these ever
by coldcode 3y ago
Despite the hype about LLMs, many of the answers are pretty terrible. The 12-bar blues progressions seem mostly clueless. The question is will any of these ever get significantly better with time, or are they mostly going to stagnate?
- smokel 3y agoWhat alternative technology do you think is better? In other words, what is your frame of reference for labeling this "pretty terrible"?
- salil999 3y agoHumans. After all, LLMs are designed to reason equal to or better than humans.
- maweaver 3y agoBy "Humans", I assume you mean something like "adult humans, well-educated in the relevant fields". Otherwise, most of these responses look like they would easily beat most humans.
- DylanDmitri 3y agoI think most high-school educated adults, with the ability to make a couple web searches, would do fine on all these questions. It would take the humans minutes instead of seconds because they don't have the internet memorized. Me, Kubernetes Haikus, time taken 84 seconds: ---------- Kubernetes rules With its smooth orchestration You can reach web scale ---------- Kubernetes sucks Lost in endless YAML hell Why is it broken?
- seabass-labrax 3y agoI think you're spot on here. Yes, if one's trying to compare human and GPT intelligence, then you have to define what counts as memorisation and what counts as reasoning. But what most people outside of academia are trying to do is work out whether a GPT can effectively replace a human in some time-consuming task, and to be able to do so without access to the internet is rarely an important factor.
- sirk390 3y agoHumans are pretty bad at these questions. Even with the simplest questions like "Sally (a girl) has 3 brothers. Each brother has 2 sisters. How many sisters does Sally have?" I think that a lot of people will give an incorrect answer. And for questions like "Argue for and against the use of kubernetes in the style of a haiku", 99.99% will not be able to do it.
- earthboundkid 3y agoThe thing with humans is they will say “I don’t remember how many syllables a haiku has” and “what the hell is kubernetes?” No LLM can reliably produce a haiku because their lexing process deprives them of reliable information about syllable counts. They should all say “I’m sorry, I can’t count syllables, but I’ll try my best anyway.” But the current models don’t do that because they were trained on texts by humans, who can do haiku, and not properly taught their own limits by reinforcement learning. It’s Dunning Kruger gone berserk.
- pixl97 3y agoEh, it's not D&K gone berserk, it's what happens when you attempt to compress reality down to a single dimension (text). If you're doing a haiku, you will likely subvocalize it to ensure you're saying it correctly. It will be interesting when we get multimodal AI that can speak and listen to itself to detect things like this.
- earthboundkid 3y agoThe problem isn’t just that everything is text. It’s that everything is a Fourier transform of text in such a way that it’s not actually possible for an LLM to learn to count syllables.
- pixl97 3y agoAgain, that is just using text only. Imagine you have a lot more computing resources in a multimodal LLM. It sees your request of count the syllables and realizes it can't do them from text alone (hell I can't and have to vocalize it). It then sends your request to a audio module and 'says' the sentence, then another listening module that understand syllables 'hears' the sentence. This is how it works in most humans, now if you do this every day you'll likely make some kind of mental shortcut to reduce the effort needed, but at the end of the day there is no unsolvable problem on the AI side.
- rvz 3y ago> After all, LLMs are designed to reason equal to or better than humans. No. I doubt you would fully trust a LLM to replace high risk jobs such as lawyers, doctors or pilots such that when something goes wrong as it is used unattended, there is no-one held to account for it to transparently explain its own mistakes and errors. It is just nonsense to suggest that such systems are capable of ‘reasoning’ when it pretends to do so and repeats itself without understanding their own errors. Thus, LLMs and other black-box AIs cannot be trusted for those high risk situations over a consensus of human professionals.
- NoraCodes 3y agoGiven that people are already firing real human workers to replace them with worse but cheaper LLMs, I'd argue that we're not talking about a competing technology, but that the competition is simply not firing your workforce. And, as an obligate customer of many large companies, you should be in favor of that as well. Most companies already automate, poorly, a great deal of customer service work; let us hope they do not force us to interact with these deeply useless things as well.
- YetAnotherNick 3y agoHow many humans in your office do you think could solve the questions with better success ratio than GPT-4? I would say less than 20%. If the primary complaint is the blues that GPT-4 wrote is not that great, I think it is definitely worth the hype, given that a year before people argued that AI can never pass turing test.
- gtowey 3y agoThat's a false dichotomy. Language models will always confidently give you answers, right or wrong. Most humans will know if they know the answer or not, they can do research to find correct information, and they can go find someone else with more expertise when they are lacking. And this is my biggest issue with the AI mania right now -- the models don't actually understand the difference between correct or incorrect. They don't actually have a conceptual model of the world in which we live, just a model of word patterns. They're auto complete on steroids which will happily spit out endless amounts of garbage. Once we let these monsters lose with full trust in their output, we're going to start seeing some really catastrophic results. Imagine your insurance company replaces thier claims adjuster with this, or chain stores put them in charge of hiring and firing. We're driving a speeding train right towards a cliff and so many of us are chanting "go faster!"
- famouswaffles 3y ago>Most humans will know if they know the answer or not, No they won't. >they can go find someone else with more expertise when they are lacking. They can but they often don't. >the models don't actually understand the difference between correct or incorrect. They certainly do https://imgur.com/a/3gYel9r https://imgur.com/a/3gYel9r
- dwaltrip 3y agoLooking at recent history, things have progressed very quickly in the past 5 years. I expect additional advances at some point in the future.
- bottlepalm 3y agoIt's like watching a baby learn how to talk..
- yard2010 3y ago...and saying it would never replace you in your job because he talks like a baby
- bottlepalm 3y agoBabies are so small and weak, no threat to anyone whatsoever.
- ilaksh 3y agoIt's like most new technologies. In the beginning there are only a few instances that really stand out, and many with issues. I remember back in like 2011 or 2012 I wanted to use an SSD for a project in order to spend less time dealing with disk seeks. My internet research suggested that there were a number of potential problems with most brands, but that the Intel Extreme was reliable. So I specified that it must be only that SSD model. And it was very fast and completely reliable. Pretty expensive also, but not much compared to the total cost of the project. Then months later a "hardware expert" was brought on and they insisted that the SSD be replaced by a mechanical disk because supposedly SSDs were entirely unreliable. I tried to explain about the particular model being an exception. They didn't buy it. If you just lump all of these together as LLMs, you might come to the conclusion that LLMs are useless for code generation. But you will notice if you look hard that OpenAIs models are mostly nailing the questions. That's why right now I only use OpenAI for code generation. But I suspect that Falcon 180B may be something to consider. Except for the operational cost. I think OpenAI's LLMs are not the same as most LLMs. I think they have a better model architecture and much, much more reinforcement tuning than any open source model. But I expect other LLMs to catch up eventually.
- guerrilla 3y ago> It's like most new technologies. In the beginning there are only a few instances that really stand out, and many with issues. Except this isn't new. This is after throwing massive amounts of resources at it multiple decades after arrival.
- gjm11 3y agoWhat are you taking "it" to be here? The transformer architecture on which (I think) all recent LLMs are based dates from 2017. That's only "multiple decades after" if you count x0.6 as "multiple". Neural networks are a lot older than that, of course, but to me "these things are made out of neural networks, and neural networks have been around for ages" feels like "these things are made out of steel, and steel has been around for ages".
- Sohcahtoa82 3y ago
- caturopath 3y agoThe majority of these LLMs are not cutting edge, and many of them were designed for specific purposes other than answering prompts like these. I won't defend the level of hype coming from many corners, but it isn't fair to look at these responses to get the ceiling on what LLMs can do -- for that you want to look at only the best (GPT4, which is represented, and Bard, which isn't, essentially). Claude 2 (also represented) is in the next tier. None of the other models are at their level, yet. You'd also want to look at models that are well-suited to what you're doing -- some of these are geared to specific purposes. Folks are pursuing the possibility that the best model would fully-internally access various skills, but it isn't known whether that is going to be the best approach yet. If it isn't, selecting among 90 (or 9 or 900) specialized models is going to be a very feasible engineering task. > The 12-bar blues progressions seem mostly clueless. I mean, it's pretty amazing that they many look coherent compared to the last 60 years of work at making a computer talk to you. That being said, I played GPT4's chords and they didn't sound terrible. I don't know if they were super bluesy, but they weren't _not_ bluesy. If the goal was to build a music composition assistant tool, we can certainly do a lot better than any of these general models can do today. > The question is will any of these ever get significantly better with time, or are they mostly going to stagnate? No one knows yet. Some people think that GPT4 and Bard have reached the limits of what our datasets can get us, some people think we'll keep going on the current basic paradigm to AGI superintelligence. The nature of doing something beyond the limits of human knowledge, creating new things, is that no one can tell you for sure the result. If they do stagnate, there are less sexy ways to make models perform well for the tasks we want them for. Even if the models fundamentally stagnate, we aren't stuck with the quality of answers we can get today.
- VerminOctopus1 3y agoI coincidentally tried to get ChatGPT 4 to give me some chord progressions today. I was wanting some easy inspiration and figured that’d be a good place to start. I was wrong, it produced total nonsense. The chord names did not match up with the key or the degrees.