4 ms·
A problem I have with this type of experiment is that it's almost pre-messaging the bot what to say. Simply filtering down to very specific variables leads the
by bdg 4y ago
A problem I have with this type of experiment is that it's almost pre-messaging the bot what to say. Simply filtering down to very specific variables leads the outcomes in a very strong way. I suspect this exact approach will not work well for finding future problems.
I would be more interested if the variables were not pre-selected with the knowledge of what was important and could determine things that haven't happened, not because of the prediction but because this means it would understand how to pick the variables that are relevant (like the human did in this experiment).
- sillysaurusx 4y agoCould you give an example of what other variables could’ve been included?
- bdg 4y agoElsewhere in this discussion thread there's a few examples from im3w1l who saw this problem too: > Like it could have mentioned jobs numbers but it didn't. It could have mentioned covid hospitalization statistics but it didn't etc
- deleted 4y ago[deleted]
- nopinsight 4y agoPerhaps someone could try other variations of the experiment as suggested above and by many here.
- czx4f4bd 4y agoAnecdotally, the times I've felt most impressed by ChatGPT usually ended up being exactly like this. Even as an LLM skeptic, I've had chat sessions with ChatGPT that felt unbelievable at the time, where it almost seemed to be generating incredibly cogent ideas and arguments all on its own. Then, after a few days, I went and reread the transcript, only to realize I'd been hinting and nudging it along much more than I consciously thought at the time. I only remembered the answers that were the most impressive and forgot about the ones that missed the mark. ChatGPT is really good at picking out the key ideas in the prompt and responding to them. It's really easy to inadvertently nudge it to give a certain type of response, which of course makes it feel all the more astounding when it gives you exactly the type of answer you already unconsciously expect.
- alpaca128 4y agoThe AI equivalent of cold reading - just like with "psychics" where the audience wants to believe in the claimed abilities.
- czx4f4bd 4y agoYou must be psychic. I was literally thinking of cold reading as I wrote my comment.
- ASalazarMX 4y agoIt helps to think of it as augmentation rather than a separate entity. That way you can lead it consciously.
- dmreedy 4y agoIndeed; I've been trying to describe these models as being consummate bullshitters. Or consummate improvisers, if you want to be kinder. Since its core task model is "autocomplete", it's less "trying to answer a question" than it is "trying to predict the most likely next thing to say after a question". Doing its best to sound like the most likely thing that comes after the thing prompting it, which in this case, often looks like an 'answer'. And when you have access to as much context as these models do, that can get as specific and nuanced as "giving you the answer you were expecting". It has got me wondering if this might be an interesting fundamental model of consciousness. A bullshit generator modulated by other systems that may have more "logical" qualities.
- bdg 4y agoI discuss this topic briefly in a post I recently made where I call it "decorative knowledge" and try to pinpoint exactly what GPT is doing and fails to do. https://buildingbetterteams.de/profiles/brian-graham/navigating-ai-businesses https://buildingbetterteams.de/profiles/brian-graham/navigat...
- tpoacher 4y agoSo ChatGPT is Clever Hans
- LanceH 4y agoWe should be running it on 1000 banks right now, and post the result next year.
- lamontcg 4y agoYeah my immediate response to reading this was "objection! leading the witness!" You know it failed due to interest rate risk and you prompt it to score about interest rate risk. It turns out that the bank that failed due to interest rate risk scored highly on it. If you already knew to ask that question back in 2021, you would have known the answer as well, but neither you or the LLM are predicting the future accurately.