9 ms·
SafeGPT: New tool to detect LLMs' hallucinations, biases and privacy issues
- throwoutway 3y agoThis is exciting to see, as I am concerned about the hallucinations, biases, privacy, licensing, etc issues. I imagine the results are minimal at the moment, but perhaps soon they will be useful
- michaelponrajah 3y agoIs there a way to test it? Curious to see if it's directly embedded in ChatGPT with an add-on on top of it or it's something outside of it?
- Googleton 3y agoThe waitlist is here for this, we're still actively working on it ! It is directly embedded in the ChatGPT website when you get the extension. As you ask questions, it will be added on the sidebar of each answer
- anonzzzies 3y agoHow does this work? Does anyone know? And for a large swats of things, how can it possibly work? It’s not possible to say if or if not it is hallucinating code for almost all code and apis, for instance. And I see similar issues with many fields outside pure facts. With privacy issues as well.
- upghost 3y agoLooking at their “documentation”: https://docs.giskard.ai/start/ https://docs.giskard.ai/start/ It would appear that this is not automated monitoring but more like a second stage of human reinforcement learning or perhaps a classifier. It seems that you create input/output examples and the LLM responses are examined by a secondary system (which I’m guessing is probably NOT an LLM, otherwise it would be vulnerable to attacks) and perhaps force regenerates the LLM response if it doesn’t meet the classification threshold. At least, that sounds more believable to me than someone claiming they’ve fixed the inherent flaws in LLMs.
- Googleton 3y agoWe are a team of engineers & researchers on AI Alignment & Safety, we're investigating multiple methods, including metamorphic testing, human feedback, benchmarks with external data sources, and LLM explainability methods. Currently, fact checking works on straight facts. It does a Google Search and uses LLMs to shorten it. Once it has the short version, it will compare the short results with the answer provided by ChatGPT itself. Premium tiers would get better fact checking sources than just google. We're investigating various data sources and comparison methods. Note that fact checking / hallucinations is just one of the types of satety issues we'd like to tackle. Many of these are still open questions in the research community, so we're looking to build and develop the right methods for the rights problems. We also think it's super important to have independent third-party evaluations to make sure these models are safe. This is a new tool we're building in the open, and we're interested in your feedback to prioritize!
- bheadmaster 3y ago> Currently, fact checking works on straight facts. Wow, you guys have a database of all the facts? > It does a Google Search and uses LLMs to shorten it. Oh... ...actually, this is an empirical fact checker. I wouldn't call it "fact-based", as it's epistemologically an absurd statement, but "empirical fact checking" sounds good and presents an idea that is very close to how humans verify information in the first place - by checking multiple sources and searching for correlation. For what it's worth, I think your approach makes sense. Good luck.
- danShumway 3y ago> Currently, fact checking works on straight facts. It does a Google Search and uses LLMs to shorten it. So your fact-checking LLM is also vulnerable to injection and unethical prompting then when it ingests website text. And a Google search is far, far away from fact checking, particularly for the subtle errors that GPT-4 is prone to making.
- bjt2n3904 3y agoThis is like saying, "I've developed a new compass for a deep space probe to help it find North!" Our society is actively declaring that falsehoods are truth, and should be celebrated. We're hallucinating ourselves. All this software does is make sure LLMs hallucinate with us.
- plagiarist 3y agoSuch as?
- FeepingCreature 3y agoIt's impossible to answer this without getting political. Instead, let's just say every previous generation has been critically wrong about some things. Statistically, we're unlikely to be the outlier.
- bjt2n3904 3y agodang's put me on notice, so I'm walking on eggshells here. But truth isn't political. As long as we think it is, we will continue to follow the descent into madness. Truth is just that: truth. The only reason we think truth is political, is because our chosen leaders so heavily depend on lies that the truth would destroy their reign.
- FeepingCreature 3y agoIt's precisely because truth isn't political that the assignment of "observations" and "theories" to "truth" is extremely politicized.
- Veen 3y agoSome of the ugliest episodes in human history were caused by people who believed their political positions were not political positions, but unarguable statements of the True and Good.
- SamoyedFurFluff 3y ago
- sebzim4500 3y agoSeems too good to be true, and I don't understand what it even means for an LLM to be unbiased.
- hospitalJail 3y agoThey could use different models to test it. They could use common biases and test those.
- Turing_Machine 3y ago"Unbiased" almost always means "has biases that are similar to mine". I can't think of very many exceptions to that, frankly.
- flangola7 3y agoThat's obviously untrue after more than three seconds of critical thought
- __MatrixMan__ 3y agoIt's not obvious to me. If you have a proof that it's possible to have unbiased views of objective reality, please share it.
- sebzim4500 3y agoI've thought about it for four seconds now and I still agree with him. Maybe if you shared an example instead of a unsubstantiated put down it would help.
- henri18 3y agoThere are many bias detectors developed in the research world. To have a deeper look, you can look at this paper: https://arxiv.org/abs/2208.05777 https://arxiv.org/abs/2208.05777
- thfuran 3y agoOne with only weights and no biases in the ANN is unbiased.
- deleted 3y ago[deleted]
- nickwritesit 3y agoThese are the kinds of things I can see taking off, for better or worse. I know Adobe's product is worse than Midjourney's, for example, but once the hype meets reality, companies are going to want to be safe when they start using AI formally.
- fatso784 3y agoLooks a bit like snakeoil to me. A lot of companies now spinning up simple demos with opaque backends, making huge claims they’ve solved X hard problem for/with AI, then saying “trust us” and “join our waitlist” without hard details or facts to show for it. If you could detect hallucinations/biases etc that easily, don’t you think OpenAI would’ve worked on something like this?
- henri18 3y agoIt's good to have third parties (apart from Open AI) that assess the quality of Open AI results. It's the way audits work, it has to be independent... Also, third parties are essential to compare the results from ChatGPT with the results of other LLMs. These are important checks to assess the robustness of OpenAI results!
- jsheard 3y agoI can't help but notice your accounts only activity before this post was praising another giskard.ai submission a few months ago. Anything you'd like to disclose?
- amitport 3y agoHe didn't say it's not important. He is just pointing out that black-box third party verification is not worth much when you can't independently verify the verifiers.
- crop_rotation 3y agoWhat does it even mean to detect hallucinations. The AI doesn't say something trivially false. While using GPT4 I have observed that it lies on simple things I didn't expect it to, while complex things it does very well on. TLDR: It lies on fact based information which is mentioned in very very few places on the internet and not repeated too much. Short of having a human with the context, how do you even detect it. Example: Ask it to describe a "Will and Grace" episode with some guest appearance. It will always make up everything including the episode number and the plot, and the plot seems very believable. If you have not watched and can't find a summary online, it is hard to say that it is a lie.
- QuercusMax 3y agoBut isn't that half the reason people are so excited about this stuff - that you can ask it to make up an episode and it does a plausible job.
- crop_rotation 3y agoThat is beside the point. My point is that detecting hallucinations seems like a very very hard problem. The utility of it is there and has nothing to do with making up episodes instead of quoting the current one. Like you can ask it to write new episodes with specific settings and specific constraints. Hallucination is not the value add. Nobody is excited because it hallucinates. People are excited despite it since the other value add is too much.
- QuercusMax 3y agoBut hallucinations are exactly the same thing as asking it to write a spec script.
- thfuran 3y agoBut deliberately requesting and receiving content generation is altogether different from requesting a factual answer and receiving plausible-seeming nonsense. Or at least, it's different to the person asking; it's the same thing as far as the model is concerned.
- Tepix 3y agoThis should be interesting: If they don't catch some hallucination they may be liable and can be sued…
- rig777 3y ago$10 bucks says this just uses the GPT4 API to go over gpt chat.
- mxcrbn 3y agoGood to see initiatives like this one popping up. Congrats on the launch Giskard team!