5 ms·
My only experience so far is with chatgpt, and I am not an expert in AI, or even just in LLMs. With those disclaimers, my interactions inform my opinion that t
by fipar 3y ago
My only experience so far is with chatgpt, and I am not an expert in AI, or even just in LLMs.
With those disclaimers, my interactions inform my opinion that the LLM behind chatgpt has no internal world model. It shows no understanding of basic facts and makes very silly mistakes very easily. I have my bias, like anyone else, but in the case of AI in particular, I should say that I don't think there's anything especially magical or sacred about conscience and the human brain (or the brain of other animals, for that matter), and I'm sure it must be possible to arrive at different forms of intelligence starting with the "inanimate" building blocks of hardware and software, but what I've seen so far in LLMs doesn't make me support your statement that they've crossed a certain line. In fact, I'm very much let down by the experience and I'm afraid once the hype goes down, we may be in for another AI winter.
As I said at the start of the comment, I'm well aware I'm not a subject-matter expert. I'm open to being wrong. I just wanted to point out that not all of us unimpressed are in denial. I'd love to be in awe at an AI breakthrough, and I kind of feel I will be in awe sometime during my lifetime, but LLMs are not (yet?) that for me.
Edit: s/are not in denial/are in denial/
- ilaksh 3y agoHave you used GPT-4?
- fipar 3y agoNo, I've only used GPT-3 so far. Should I be excited about GPT-4? :) Examples of things I'd file as silly mistakes: - mentioning non-existing settings when asking it about a particular software. - responding to roughly "give me a way to list all rds clusters that are running a version that will reach EOL within the next 12 months" (which I know is almost impossible to do without scraping, as the EOLs are not returned by the API, but it was worth trying) with a query that compares the version with < 12.0. It did acknowledge the error when I asked "won't that just compare the version number" and, interestingly, it then did add to the response that the API does not provide EOL for versions, but the first response was IMHO bad in a very basic way. - providing wrong steps in an answer to "give me steps to do X", wrong in a way that would cause the process to go very bad, and then when I point out the problematic step, it responds with a seemingly random alternate step that is also wrong. I guess this part is the one that is less bad, because it loosely reminds me of people who are bad at reasoning and just give back random answers to a logic or math question. I must say it is surprisingly good (though it will still make up facts at times) when answering questions about elisp though. Paraphrasing Hofstadter, I'd say if I'm asked about the intelligence of chatgpt using gpt-3 I'd say it's slightly better than a thermostat or a mosquito, but not much more. Furthermore, my (admittedly not very valuable to others) gut feeling is this is not the path to get to what I'd consider an intelligent being, mostly because in my interactions, it makes lots of bad mistakes, and none of the good ones (here I'm using good and bad as qualifiers for mistakes based on my experiences as a parent, and as a mentor to junior peers, where a good mistake is something that lets you know they're on the right path even if they haven't arrived at the right answer/place yet). Try asking it about analogies, for example ("what's a good analogy to explain X to someone new to it"), the results I've gotten are underwhelming, when not just plain wrong. Edit: s/fall/file/ Formatting.
- School-Cotton 3y ago> No, I've only used GPT-3 so far. Should I be excited about GPT-4? The difference is extremely dramatic. Any experience with 3 or 3.5 is completely irrelevant for 4.
- kaffekaka 3y agoThere is a clear difference, but "extremely dramatic" is obviously hyperbole. (I use 3.5 and 4, both in the chatgpt interface and via API.)
- nopinsight 3y agoGPT-4 should do significantly better, but still worse than expert humans, on the tasks you mentioned. They would also benefit from good prompting, such as those in - https://lilianweng.github.io/posts/2023-03-15-prompt-engineering https://lilianweng.github.io/posts/2023-03-15-prompt-enginee... - https://help.openai.com/en/articles/6654000-best-practices-for-prompt-engineering-with-openai-api https://help.openai.com/en/articles/6654000-best-practices-f... Not a large percentage of humans would be able to do the tasks listed above without significant experience or training either. Average humans are also not great at reasoning.
- fipar 3y agoThanks, I'll wait for it to be available (I'm a casual user and I'm not the one setting this up or paying for it, if someone is paying for it, so I'm not in control of which version I use). I'm in full agreement about humans and reasoning, btw. I just don't think chatgpt (with the version I've used) is anywhere on the same league as the worst (way below-average) humans either. I do think it's quite useful as a writing aid. You do need to review what it gives you, but you'd have to do that even if you were relying on a human writing aid.
- baq 3y agoTry Bing chat, it has a GPT-4 mode, which you have to look for but it’s there and free.
- nopinsight 3y agoCurrent LLMs likely have multiple world models with varying qualities. Good prompts are required to activate suitable models for each task. 1. Did you use GPT-3.5 (free version) or GPT-4 (paid one) to form the judgment? Their performances on harder tasks differ significantly as shown in https://openai.com/research/gpt-4 https://openai.com/research/gpt-4. 2. Have you tried adding "Please think step by step" to some harder requests? This simple phrase gets most current LLMs to perform significantly better. It's a bit like asking students to show their work, which forces them to think more clearly. Current LLMs, without additional mechanisms, tend to be perform like a drunk or sleepy human, i.e. using mostly intuition or System-1 (as defined in "Thinking Fast and Slow"). The prompt such as "think step by step" asks it to think more in the System-2 style. (There are other techniques which get them to perform even better still.) I think of current LLMs as a very-well-read, but often sleepy intern who needs strict instructions, feedback, and sometimes extra training if you want them to perform well.
- fipar 3y agoAs stated in another response, I'm using 3, so I'm probably missing some good stuff. I have tried adding things like "please think step by step" (quite literally that question actually), and also "please make sure you check the facts before answering so you don't include non-existing arguments" (when asking about a cli tool that takes arguments), but I didn't notice a significant improvement. I like what you say in your last line, though I think a big challenge is that a very-well-read intern got to be very-well-read for at least two reasons: 1) they like to read and can do that a lot (LLMs can do this, not saying they like, but saying training them is analogous to someone reading a lot), 2) they went through a curated list of reading material. My experiences make me think part 2 is a weak part of chatgpt. I still think LLMs are not the way towards an interesting AI. I don't know why we're insisting so much on natural language. I mean, I can understand this for simpler tasks like a support chatbot, but I wish there was (maybe there is?) good research on building an AI that is not based on human language, since it's one of the worst mediums to communicate with rigor.
- gpderetta 3y agoI don't think baseline ChatGPT si capable to "please check the facts before answering so you don't include non-existing arguments".
- pyth0 3y agoPeople are already asking you if you have tried <insert version here> but this pretty much echos my experience and I have used both free and paid versions of ChatGPT (3.5 and 4 respectively) as well as the GPT4 API directly. I had a similar experience between all three of them, although the GPT4 based versions were certainly better in terms of output quality. In my opinion it seems pretty clear that this is a fundamental limit of the architecture and all the RLHF and increase in model sizes wont change that.
- famouswaffles 3y agoIt's so fascinating going through previous GPT-X threads. You can almost everybody making the same kind of extrapolation mistakes you're making. People just don't seem all that interested in revising their model when it's obviously wrong.
- webdood90 3y agoI really don't understand replies like this. This is just the beginning. It is a tiny, tiny view into this new technology. Nobody with an ounce of intelligence looked at the first iPhone and thought "ah well it's just another phone". It's the potential of this new technology that is so impressive. How can you not be impressed by this? I am literally incapable of empathizing with your perspective. All I can think is that the world is going to pass you by.
- ric2b 3y agoI felt similarly to you with cryptocurrency but then not much improved in the 10 years since. Maybe LLM's will be different.
- corethree 3y agoTrue but in those 10 Years AI has been moving at a rocket pace. Draw the trend line.
- webdood90 3y agoI personally didn't have that "ah ha" moment with crypto/blockchain, but I understand your point. to me AI is different - it's current iteration is useful, today. just this year alone it has improved at an insane pace, so I'm excited to see where we are in a few years.
- fipar 3y agoI'm not sure I follow how you get from me not being impressed by the LLMs I've interacted with (you may be incapable of empathizing with that perspective, but that doesn't make me less unimpressed at them) to some hypothetical people thinking the first iPhone was just another phone, and therefore not having an ounce of intelligence. I'm not worried about the world passing me by, I think LLMs and similar technology will increase, not decrease, the need for people who can troubleshoot and fix live production software. I mean I do have my passive income if I'm forced into early retirement, but I'm not fearing that to happen soon. Mind you, I did say I'd love to be in awe at an AI breakthrough; I'm not a luddite. I just don't see anything in LLMs to impress me, and on the other side, I do see a lot of hype behind it.