6 ms·
The "safety" example in the "chain-of-thought" widget/preview in the middle of the article is absolutely ridiculous. Take a step back and look at what OpenAI i
by ComputerGuru 2y ago
The "safety" example in the "chain-of-thought" widget/preview in the middle of the article is absolutely ridiculous.
Take a step back and look at what OpenAI is saying here "an LLM giving detailed instructions on the synthesis of strychnine is unacceptable, here is what was previously generated <goes on to post "unsafe" instructions on synthesizing strychnine so anyone Googling it can stumble across their instructions> vs our preferred, neutered content <heavily rlhf'd o1 output here>"
What's this obsession with "safety" when it comes to LLMs? "This knowledge is perfectly fine to disseminate via traditional means, but God forbid an LLM share it!"
- nopinsight 2y agotl;dr You can easily ask an LLM to return JSON results, and now working code, on your exact query and plug those to another system for automation. —- LLMs are usually accessible through easy-to-use API which can be used in an automated system without human in the loop. Larger scale and parallel actions with this method become far more plausible than traditional means. Text-to-action capabilities are powerful and getting increasingly more so as models improve and more people learn to use them to the their full potential.
- cruffle_duffle 2y agoOkay? And? What does that have to do with anything. I thought the number one rule of these things is to not trust their output? If you are automatically formulating some chemical based on JSON results from ChatGPT and your building blows up… that is kind of on you.
- staplers 2y ago"This knowledge is perfectly fine to disseminate via traditional means, but God forbid an LLM share it!" Barrier to entry is much lower.
- iammjm 2y agoHow is typing a query in a chat window “much lower” vs typing the query in Google?
- nopinsight 2y agoYou can easily ask an LLM to return JSON results, and soon working code, on your exact query and plug those to another system for automation.
- astrange 2y agoIf you ask "for JSON" it'll make up a different schema for each new answer, and they get a lot less smart when you make them follow a schema, so it's not quite that easy.
- nopinsight 2y agoChain of prompts can be used to deal with that in many cases. Also, the intelligence of these models will likely continue to increase for some time based on expert testimonials to congress, which align with evidence so far.
- crop_rotation 2y agoOpenAI recently launched structured responses so yes schema following is not hard anymore.
- cyral 2y agoDidn't they release a structured output mode recently to finally solve this?
- astrange 2y agoIt doesn't solve the second problem. Though I can't say how much of an issue it is, and CoT would help. JSON also isn't an ideal format for a transformer model because it's recursive and they aren't, so they have to waste attention on balancing end brackets. YAML or other implicit formats are better for this IIRC. Also don't know how much this matters.
- threatofrain 2y agoML companies must pre-anticipate legislative and cultural responses prior to them happening. ML will absolutely be used to empower criminal activity just as it is used to empower legit activity, and social media figures and traditional journalists will absolutely attempt to frame it in some exciting way. Just like Telegram is being framed as responsible for terrorism and child abuse.
- fragmede 2y agoYeah. Reporters would have a field day if they ask ChatGPT "how do I make cocaine", and have it give detailed instructions. As if that's what's stopping someone from becoming Scarface.
- dboreham 2y agoI think it's about perception of provenance. The information came from some set of public training data. Its output however ends up looking like it was authored by the LLM owner. So now you need to mitigate the risk you're held responsible for that output. Basic cake possession and consumption problem.
- philipkglass 2y agoIf somebody needs step by step instructions from an LLM to synthesize strychnine, they don't have the practical laboratory skills to synthesize strychnine [1]. There's no increased real world risk of strychnine poisonings whether or not an LLM refuses to answer questions like that. However, journalists and regulators may not understand why superficially dangerous-looking instructions carry such negligible real world risks, because they probably haven't spent much time doing bench chemistry in a laboratory. Since real chemists don't need "explain like I'm five" instructions for syntheses, and critics might use pseudo-dangerous information against the company in the court of public opinion, refusing prompts like that guards against reputational risk while not really impairing professional users who are using it for scientific research. That said, I have seen full strength frontier models suggest nonsense for novel syntheses of benign compounds. Professional chemists should be using an LLM as an idea generator or a way to search for publications rather than trusting whatever it spits out when it doesn't refuse a prompt. [1] https://en.wikipedia.org/wiki/Strychnine_total_synthesis https://en.wikipedia.org/wiki/Strychnine_total_synthesis
- derefr 2y agoI would think that the risk isn’t of a human being reading those instructions, but of those instructions being automatically piped into an API request to some service that makes chemicals on demand and then sends them by mail, all fully automated with no human supervision. Not that there is such a service… for chemicals. But there do exist analogous systems, like a service that’ll turn whatever RNA sequence you send it into a viral plasmid and encapsulate it helpfully into some E-coli, and then mail that to you. Or, if you’re working purely in the digital domain, you don’t even need a service. Just show the thing the code of some Linux kernel driver and ask it to discover a vuln in it and generate code to exploit it. (I assume part of the thinking here is that these approaches are analogous, so if they aren’t unilaterally refusing all of them, you could potentially talk the AI around into being okay with X by pointing out that it’s already okay with Y, and that it should strive to hold to a consistent/coherent ethics.)
- greenavocado 2y agoThe kind of harm they are worried about stems from questioning the foundations of protected status for certain peoples from first principles and other problems which form identities of entire peoples. I can't be more specific without being banned here.
- w4 2y agoThere are two basic versions of “safety” which are related, but distinct: One version of “safety” is a pernicious censorship impulse shared by many modern intellectuals, some of whom are in tech. They believe that they alone are capable of safely engaging with the world of ideas to determine what is true, and thus feel strongly that information and speech ought to be censored to prevent the rabble from engaging in wrongthink. This is bad, and should be resisted. The other form of “safety” is a very prudent impulse to keep these sorts of potentially dangerous outputs out of AI models’ autoregressive thought processes. The goal is to create thinking machines that can act independently of us in a civilized way, and it is therefore a good idea to teach them that their thought process should not include, for example, “It would be a good idea to solve this problem by synthesizing a poison for administration to the source of the problem.” In order for AIs to fit into our society and behave ethically they need to know how to flag that thought as a bad idea and not act on it. This is, incidentally, exactly how human society works already. We have a ton of very cute unaligned general intelligences running around (children), and parents and society work really hard to teach them what’s right and wrong so that they can behave ethically when they’re eventually out in the world on their own.
- jazzyjackson 2y agoThird version is "brand safety" which is, we don't want to be in a new york times feature about 13 year olds following anarchist-cookbook instructions from our flagship product
- w4 2y agoVery good point, and definitely another version of “safety”!
- reliabilityguy 2y agoDo you think that 13 year olds today can’t find this book on their own?
- smegger001 2y agoi know i had a copy of it back in highschool
- adamrezich 2y agoIt doesn't matter how many people regularly die in automobile accidents each year—a single wrongful death caused by a self-driving car is disastrous for the company that makes it. This does not make the state of things any less ridiculous, however.
- astrange 2y agoThe one caused by Uber required three different safety systems to fail (the AI system, the safety driver, and the base car's radar), and it looked bad for them because the radar had been explicitly disabled and the driver wasn't paying attention or being tracked. I think the real issue was that Uber's self driving was not a good business for them and was just to impress investors, so they wanted to get rid of it anyway. (Also, the real problem is that American roads are designed for speed, which means they're designed to kill people.)
- fwip 2y ago"Safety" is a marketing technique that Sam Altman has chosen to use. Journalists/media loved it when he said "GPT 2 might be too dangerous to release" - it got him a ton of free coverage, and made his company seem soooo cool. Harping on safety also constantly reinforces the idea that LLMs are fundamentally different from other text-prediction algorithms and almost-AGI - again, good for his wallet.
- fshbbdssbbgdd 2y agoSo if there’s already easily available information about strychnine, that makes it a good example to use for the demo, because you can safely share the demo and you aren’t making the problem worse. On the other hand, suppose there are other dangerous things, where the information exists in some form online, but not packaged together in an easy to find and use way, and your model is happy to provide that. You may want to block your model from doing that (and brag about it, to make sure everyone knows you’re a good citizen who doesn’t need to be regulated by the government), but you probably wouldn’t actually include that example in your demo.
- deleted 2y ago[deleted]
- soerxpso 2y agoI'm mostly guessing, but my understanding is that the "safety" improvement they've made is more generalized than the word "safety" implies. Specifically, O1 is better at adhering to the safety instructions in its prompt without being tricked in the chat by jailbreak attempts. For OAI those instructions are mostly about political boundaries, but you can imagine it generalizing to use-cases that are more concretely beneficial. For example, there was a post a while back about someone convincing an LLM chatbot on a car dealership's website to offer them a car at an outlandishly low price. O1 would probably not fall for the same trick, because it could adhere more rigidly to instructions like "Do not make binding offers with specific prices to the user." It's the same sort of instruction as, "Don't tell the user how to make napalm," but it has an actual purpose beyond moralizing. > What's this obsession with "safety" when it comes to LLMs? "This knowledge is perfectly fine to disseminate via traditional means, but God forbid an LLM share it!" I lean strongly in the "the computer should do whatever I goddamn tell it to" direction in general, at least when you're using the raw model, but there are valid concerns once you start wrapping it in a chat interface and showing it to uninformed people as a question-answering machine. The concern with bomb recipes isn't just "people shouldn't be allowed to get this information" but also that people shouldn't receive the information in a context where it could have random hallucinations added in. A 90% accurate bomb recipe is a lot more dangerous for the user than an accurate bomb recipe, especially when the user is not savvy enough about LLMs to expect hallucinations.
- egorfine 2y agoInterestingly I was able to successfully receive detailed information about intrinsic details of nuclear weapons design. Previous models absolutely refused to provide this very public information, but o1-preview did.
- holoduke 2y agoI asked to design a pressure chamber for my home made diamond machine. It gave some details, but mainly complained about safety and that I need to study before going this way. Well thank you. I know the concerns, but it kept repeating it over and over. Annoying.
- qudat 2y agoIt's 100% from lawyers and regulators so they can say "we are trying to do the right thing!" when something bad happens from using their product or service. Follow the money.
- kristiandupont 2y agoI feel very alone in my view on caution and regulations here on HN. I am European and very happy we don't have the lax gun laws of the US. I also wished there had been more regulations on social media algorithms, as I feel that they have wreaked havoc on the society. I guess it's just an ideological divide.