3 ms·
How do you get the chatbot to tell the truth more? I’m a scientist and this is my biggest showstopper with AI tools. I just can’t know if I can trust what it’s
by icapybara 3y ago
How do you get the chatbot to tell the truth more? I’m a scientist and this is my biggest showstopper with AI tools. I just can’t know if I can trust what it’s telling me.
- unblough 3y agoYou are unable to. “LLMs can’t self-correct in reasoning tasks, DeepMind study finds“ https://news.ycombinator.com/item?id=37823543 https://news.ycombinator.com/item?id=37823543 Anyone who says otherwise is either ignorant of the underlying function of llms or trying to sell you something.
- mistermann 3y agoOr, they realize that absence cannot be scientifically proven, and that scientists use language loosely, confusing those who take their loose language literally. The study didn't "find" (discover) what they claim, rather, they didn't find validation that it "can" (the implementation of which varies per observer, sub-perceptually). If you had to code something like this at work for a different domain, I bet you'd have no problem realizing that a nullable boolean is required to accurately model the problem space.
- unblough 3y agoI recognize your clarification of “discovery” and conclusion from that research, but I do think there is a strong argument that in terms of the stochastic usage of a nonlinear system the “undefined” state of your nullable boolean is itself a falsey state.
- mistermann 3y agoYou can argue whatever you like, but if the unknown IS actually known, why can't scientists tell us their secrets? How many people would have to be in on the scheme? And this isn't just a one off, this is a systemic, institutional shortcoming, I encounter several instances of it every day just in my regular social media feeds.
- famouswaffles 3y ago>You are unable to. This is just wrong lol. GPT-4 logits calibration pre RLHF - https://imgur.com/a/3gYel9r https://imgur.com/a/3gYel9r Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback - https://arxiv.org/abs/2305.14975 https://arxiv.org/abs/2305.14975 Teaching Models to Express Their Uncertainty in Words - https://arxiv.org/abs/2205.14334 https://arxiv.org/abs/2205.14334 Language Models (Mostly) Know What They Know - https://arxiv.org/abs/2207.05221 https://arxiv.org/abs/2207.05221
- unblough 3y ago> This is just wrong lol. The needless condescension of your “lol” feels a bit premature. How can you have self correction without superintelligence?
- jiggawatts 3y agoDon't try to use AIs as omniscient oracles. Use them to replace grad students / interns. In other words, AIs are useful when you can confidently evaluate their output for correctness. Note that this also applies in novel scenarios, such as new scientific discovery. It's often much more difficult to come up with an answer than it is to verify it. As a random example, modern AIs are very good at image recognition tasks, which would make them ideally suited for finding rare phenomena in all-sky surveys. The AI doesn't need to write the research paper! It can be valuable if it can just find candidates for detailed evaluation by human researchers. Similarly, I've found AIs are quite good at spotting errors and inconsistencies. Don't make it write the paper... just ask it to proofread it for you. Etc...
- oersted 3y agoNever rely on the AI's own "memory". GPT-4 can remember a lot of deep technical things correctly from its training, but it can also write non-factual statements that look perfectly reasonable. We only use LLMs to retrieve and extract information from reliable sources (mostly papers and patents). If you provide extensive background information in the prompt, it will stick to that and avoid making up facts, and it has the added benefit of including verifiable references for each statement in the answer. Again, we treat the AI mostly as an "automatic reader", it's very good at that, it's not hard to keep it factually grounded and that alone can save hours of research. But when asking it to do deep reasoning or original writing it can be a bit more tricky to keep it controlled.
- jstarfish 3y agoTreat it like you're dealing with some vendor's salesperson. You have to bring your own source of truth, like the interfaces that let you "talk to your PDFs". Prep it with data you know to be true, then start your discussion. If you show up unprepared and let it tell you what's real, it will bullshit you in its favor.