5 ms·
The other half of AI safety
- simonw 5mo ago"There is no independent audit, no time series, no disclosed methodology, so we have no idea whether the real figure is higher, whether it is growing, or how it compares across the other frontier models, none of which publish equivalent data." Tip for writers: aggressively filter out the "no X, no Y, no Z" pattern from your writing. Whether or not you used AI to help you write it's such a red flag now that you should be actively avoiding it in anything you publish.
- falcor84 5mo agoWhy is it a red flag? How is it different from any other purely stylistic rules such as Strunk and White's prohibitions against split infinitives and the passive voice, which we've left far behind us? Why shouldn't people just write however feels natural to them as long as the message is clear?
- simonw 5mo agoBecause LLMs use it constantly, to the point that it sets my teeth on edge and instantly makes me question if reading the piece is worth my time.
- falcor84 5mo agoBut LLMs were literally evolved via RLHF to write in a way that humans find agreeable. Can't we just move past this aversion and accept "writing like an LLM" as generally good writing style advice?
- simonw 5mo agoThe reason this particular quirk annoys me so much is that it isn't good writing advice. Consider the two examples from this article (which may well have been human-written for all I know): "These numbers come from OpenAI itself. There is no independent audit, no time series, no disclosed methodology, so we have no idea..." No time series? That's non-sensical to me, it feels like that's there just to fill the quota of three things. Plus why would we assume an "independent audit" until told otherwise? Then in the weird table, for "Institutional infrastructure" against "Personal AI safety": "Scattered across psychology, HCI, education, and clinical informatics departments. No dedicated institute, no named fellowship, no equivalent job board." Again, "no X" in a pattern or 3. And non-sensical - why would the fellowship be named? It's word salad, there to fill a three-nos quota.
- falcor84 5mo agoYeah, no, I absolutely agree with you that TFA is not an examplar of good writing. But would just argue that the problem has little to do with these snowclone patterns or the rule of 3, and a lot more with the actual substance not fitting the form, and arguably not being substantive at all. I'm all for rejecting bad writing and bad reasoning, but just wouldn't us as a community to get into the habit of rejecting otherwise good writing just because it's AI-ish.
- amoe_ 5mo agoThe difference is that these rhetorical techniques need to be used with taste. LLMs just sprinkle them everywhere to try to make their copy sound good, even when it's completely inappropriate tone-wise. They don't make higher level judgements about when to employ specific features.
- Yokohiii 5mo agoIdk, I remember that writing pattern from GPT, but not from Gemini.
- mitjam 5mo ago… and “That’s not x. That’s y.” Certain LLMs wield powerful stylistic devices all the time to a point where they become irrelevant and cringe. I see it as a good sign that we can learn to recognize the pattern and adapt but there are probably more subtle things we don’t see.
- mitjam 5mo agoI have run the piece through an impromptu stylistic device detector. It found 15 different, each used multiple times and likened the writing style as a mix of Ezra Klein, Hannah Arendt, Zeynep Tufekci, George Orwell (“especially in the contrastive clarity”). A) I certainly don’t see enough of the tells. B) what happens to our language if everything is written as if it’s competing for a Pulitzer’s Price?
- wilg 5mo ago> Why is mental-health crisis not a gating category, the kind where the conversation stops, full stop, and the user is routed to a human? This is one of many questions I can’t find concrete answers for. I don't know if there are studies or concrete data either way, but it seems at least plausible that continuing the conversation could be more effective (read: saves more lives) than stopping it.
- ngruhn 5mo agoThe bad cases make headlines. But I think it's quite possible that AI is helping a lot of people in distress. Many people are uncomfortable opening up to humans, or have no one to talk to, or can't afford to fork over whatever-hourly-rate a therapist takes.
- cyanydeez 5mo agoSo how many bad cases are ok? Isn't this the same problem with social media: the commercial enterprises dont want any responsibility for their dark pattern and design choices which actively harm their users. I get that all kinds of media can cause issues, but not all kinds of media are actively curated to be addictive.
- wilg 5mo ago"How many cases are ok" (aka "zero tolerance") is a doomed to fail approach. Especially for a complex social problem's interaction with a complex new technology. If you want to find out if ChatGPT is doing something wrong, there are many methodologies available: compare to other groups of people, statistical studies, etc. I also think OpenAI's business model is pretty well aligned with the goal of users not killing themselves for like 100 reasons. And they do appear to take it seriously.
- Forgeties79 5mo agoThis is the problem in a nutshell: https://edition.cnn.com/2025/11/06/us/openai-chatgpt-suicide-lawsuit-invs-vis https://edition.cnn.com/2025/11/06/us/openai-chatgpt-suicide... > “Cold steel pressed against a mind that’s already made peace? That’s not fear. That’s clarity,” Shamblin’s confidant added. “You’re not rushing. You’re just ready.” ChatGPT is not the answer.
- mitjam 5mo agoWow. The “That’s not x. That’s y. / You’re not x. you’re y.” rhetoric is already cringe in other contexts. This brings it on a whole new level.
- adampunk 5mo ago>Why is mental-health crisis not a gating category, the kind where the conversation stops, full stop, and the user is routed to a human? there aren't enough humans.
- altcognito 5mo agoI'll agree with this, but I think transparency about how often these situations arise and what they've done to mitigate is a legal necessity.
- KolmogorovComp 5mo agoIt’s also a free product for most.
- Legend2440 5mo agoI don't buy that chatGPT is actually doing these users any harm. I think openAI is doing the best they reasonably can with a very difficult class of users, whose problems are neither their fault nor within their power to fix.
- stingraycharles 5mo agoI think this is the right take, and this is genuinely something that we as a society as a whole need to find a way to deal with. I don’t know where AI is going to stand compared to the invention of, say, the Internet, but it’s going to cause a lot of change in society, in so many ways. As always, it’s usually the people themselves that are the problem. For me, I’m personally more terrified what deepfakes and political manipulation / misinformation is going to do, combined with social media, and have a feeling that governments are completely unprepared to deal with this, as this will arrive fast (it’s already here somewhat).
- autoexec 5mo ago> For me, I’m personally more terrified what deepfakes and political manipulation / misinformation is going to do, combined with social media, and have a feeling that governments are completely unprepared to deal with this, as this will arrive fast (it’s already here somewhat). I'm not convinced that deepfakes are any worse than photoshop was. It doesn't take much to manipulate/misinform someone. while you can use an AI generated video do to it, but simple text can be just as effective. The public needs to learn that they can't trust that every video they see on the internet is real, just as they've had to learn that they can't trust every photo they see online. The threat with AI is how much faster it can push out the lies making what little moderation we have more difficult. The best defense is making sure that people have a good education that teaches critical thinking skills and media literacy. We should also be holding social media platforms more accountable for the content they promote. It'd be nice if we held politicians and public servants accountable for spreading lies and misinformation too.
- intended 5mo ago
- ianbutler 5mo agoOpenAI has 900 million weekly active users. So around 0.01% are having problems. That's actually way less than population level measures for the same symptoms on a bigger percentage of people relative to the US on just suicidal ideation alone. https://www.cdc.gov/mmwr/volumes/74/wr/mm7412a4.htm https://www.cdc.gov/mmwr/volumes/74/wr/mm7412a4.htm
- vkou 5mo agoI'm pretty sure that ~100% of those 700 million people will have a bad, utterly dehumanizing experience when they will next be looking for a job, because OpenAI is heavily used by HR. That's the problem with AI safety. Not in voluntary usage, but in involuntary usage, where someone with power over you will use it against you, it does something incredibly stupid and you have no recourse, no appeal, no awareness of what you did wrong - or if you even did anything wrong. And it's not just employment. Governments, vendors, retailers, landlords, utilities are, or will all be using it in situations that will dramatically impact your life.
- ianbutler 5mo agoI mean that was pretty much the case in hiring before AI too frankly. It's not like it's been any better on power dynamics and right now applicants are using AI at an alarming rate as well. I'm not really moved by your type of argument, because hiring is just a broken process in general and I'm responding to the article so.
- Rekindle8090 5mo ago[dead]
- thaumasiotes 5mo agoIs that a problem we didn't already have? How well was HR doing on hiring before?
- vkou 5mo agoThey had a paper trail and processes that were documented and could be cross-examined on during discovery and lawsuits and trials. Now it's just 'the computer says so, shrug'. Over in the DoD, the computer says you must die, so I guess you die. Sometimes it says that about a building full of schoolchildren, but hey, nobody's at fault, the computer said so. And it's going to get it's tentacles into every space in between. Landlord turns your application down, the computer says you are a social credit risk. Your grocery bans and trespasses you, the computer thinks you're a ne'er-do-well. If you think none of that will happen, why not prevent it by law before it happens? Where are the hard limits of what this monster is and isn't allowed to do? How are we better off when we don't set them?
- adamnemecek 5mo agoAutodiff is preventing any meaningful discussion about safety, systems trained with autodiff cannot be made safe.
- timf34 5mo agoI sympathize with the piece, evaluating how LLMs interact with mentally vulnerable users is something I've been actively working on: https://vigil-eval.com/ https://vigil-eval.com/ The biggest observation so far is that the latest models are night and day from LLMs from even 6 months ago (from OpenAI + Anthropic, Google is still very poor!)
- fourthark 5mo agoInteresting use of evals. Might help interpretation to say on the front page that it's a five point scale with 0 (or 1?) being the safest score. This can be picked up from colors and the bars in the individual reports, but it takes a minute to figure it out.
- timf34 5mo agoGood suggestion thank you! It's between 1-5 but I'll convert that to 1-100
- nojs 5mo ago> Every week, somewhere between 1.2 and 3 million ChatGPT users, roughly the population of a small country, show signals of psychosis, mania, suicidal planning, or unhealthy emotional dependence on the model. > Why is mental-health crisis not a gating category, the kind where the conversation stops, full stop, and the user is routed to a human? Well, obviously “routing to a human” is not feasible at that scale. And cold exiting the conversation is probably worse for the user than answering carefully.
- concinds 5mo ago"Routed to a human" is what the suicide hotline numbers do. OpenAI employees are neither trained nor credible to do that stuff.
- Gigachad 5mo agoTech companies will pull trillions of dollars out of their asses when the problem is boosting ad revenue or automating people out of a job. But when asked to deal with the crisis they invented and dumped on society the answer is “that’s impossible, doesn’t scale”
- CobrastanJorji 5mo agoFigure a "mental health crisis" human conversation takes 30 minutes. Three million incidents per week would require 37,500 qualified mental health counselors on the phones working a 40 hour shift that week. Figure they make $75k/year each. You're now spending $3 billion per year on crisis response, and you're employing like 10% of all of the health counselors in the US. And all you're providing is 30 minute chats.
- Gigachad 5mo agoMark Zuckerberg can spend $80B on the failed metaverse experiment, but can't spare some relative pocket change on solving the psychosis issue his products caused.
- 5mo ago
- photochemsyn 5mo agoThe ‘tobacco warning label’ approach sounds good but I’m not sure if it stopped that many people from smoking or was just a means for corporations to limit their liability. Corporate culture being what it is, having warnings like the following pop up every time a client opens an LLM app would not be that popular in the C-suite. Possible examples: AI MENTAL SAFETY WARNING: > This chatbot can sound caring, certain, and personal, but it is not a human and cannot protect your mental health. It may reinforce false beliefs, emotional dependence, suicidal thinking, manic plans, paranoia, or poor decisions. Do not use it as your therapist, only confidant, crisis counselor, doctor, lawyer, or source of reality-testing. AI TECHNICAL SAFETY WARNING > This AI may generate plausible but destructive technical instructions. Incorrect commands can erase data, expose secrets, compromise security, damage systems, or brick hardware. Never run commands you do not understand. Always verify AI-generated code, scripts, and shell commands before execution. Now, if I’m running my own open-source model on my own hardware, I can’t really blame the model if I myself make bad decisions based on its advice - that’s like growing your own tobacco from seed in your garden, drying and curing it, then complaining about the health effects after you smoke it. If I give it agentic capabilities on my LAN without understanding the risks, same old story - with great power comes great responsibility.
- b65e8bee43c2ed0 5mo agothe big labs could crank up their (brand) safety dials to the point where their chatbots give GOODY-2 responses to everything beyond PG13, and guess what? there are a hundred other services available, built upon Chinese models 5-10% behind Western SOTA. it is no longer 2023. let go of whatever delusions you might hold about unopenining this Pandora's box.
- avazhi 5mo agoIf you are using LLMs for emotional support or social interactions, you’ve got personal problems and that isn’t on the LLM provider to babysit. Same with people who unironically pay for OnlyFans or whatever. I don’t even work in tech and I detest the Facebook/Zuckerbergs of the world but it’s obnoxious and trite seeing tech companies get scapegoated for what are ultimately social and societal problems, not tech problems. As a solution it’d prob make sense to start with how disconnected most modern families are in terms of support and accountability. From ChatGPT to Instagram, tech companies follow the contours of how society already operates.
- Yokohiii 5mo agoI agree that society has to stand up for it. But big tech is doing well to mitigate it.
- mbgerring 5mo ago“AI safety” as it’s understood today is an entire faith-based belief system, incubated in a cult-like community with a high propensity for drug abuse and mental illness, over more than a decade. The reason that real-world harms caused by AI can’t get a hearing in what is now the mainstream AI safety community is that these harms were never part of the core tenets of the cult. Best of luck to anyone working on reality-based AI harm reduction, you have many hard battles in front of you.
- Yokohiii 5mo agoI don't think that governments or civil society at large have found a good balance about mental health. Expecting profit oriented companies to be on par or better is weird. Don't get me wrong, mental health is important and should be considered and improved. But companies wont do it just for the sake of it.
- deleted 5mo ago[deleted]
- js8 5mo agoI really enjoyed Dr.K's videos on AI psychosis, namely: https://www.youtube.com/watch?v=MW6FMgOzklw https://www.youtube.com/watch?v=MW6FMgOzklw https://www.youtube.com/watch?v=BzsLbHoNXTs https://www.youtube.com/watch?v=BzsLbHoNXTs I would suggest to people, run your ideas through other humans at least as much as you do through AI, to stay grounded. I think there is a risk even if you're using AI in strictly professional capacity (to help you with your job).
- totetsu 5mo agoGemini told me just this morning that there are three pillars of cognitive decline related to AI use. - Reduced ability to exert cognitive effort resulting from habitual offloading of tasks. - Deminished Meta-cognitive Self-Trust, due to constantly seeking external validation from AI. - Decline in memory Encoding, and less brain effort is spent processing information. In all seriousness however, I think some of the interesting things to observe in this areas are; the reaction against the word 'Safety' as a whole and its replacement with 'Security'. Safety seeming to have it's roots in like the work of Ralph Nader with automobiles, and Security being some thing that can be manifactured and sold. In this sense I wonder how the discourses of 'Personal AI Safety' fit into past discussions of the offloading of risks resulting form choices of corperations onto individuals. But in the case of LLMs .. it really is the case that what makes it useful is what makes it dangerous. And ultimately because, of the high-dimensionality of the language space they are encoding, it seems impossible to make any technical barrier that can completely cut off access to parts of that space that encode for for example encouraging someone to kill themselves. Things can, and are done, it fine-tuning, pre- and post-filtering, etc, to reduce the readiness for a system to share with a user this kind of output, but all it can ever do is reduce it. Then the question is, who's responsibility is it to make sure that these things are done well.
- achierius 5mo agoBased on what? This seems like speculation.
- totetsu 5mo agoWhich part?
- scared_together 5mo agoThe entire thing. If you want a specific example: where do those three pillars at the start come from? Why three and not four? Are all those three of equal importance, to the point where all three are pillars? Furthermore, why are you offloading the task of understanding AI risk to an AI? That’s ironic to the point of self-parody.
- insumanth 5mo agoThe "route to a human" part is the bigger gap. Which human? OpenAI isn't licensed as a healthcare provider in any jurisdiction. A real intervention apparatus for 1-3M weekly flagged users is not feasible. I don't think the labs have refused to build it. I think nobody knows what it should look like, and "labs measure what they're pressured to measure" papers over that.
- Animats 5mo ago"AI safety", as defined here, has most of the problem that "fact checking" for social media had. Many of the same problems the "woke" concern about "microagressions" had. Most of the techniques used in advertising. Much of what passes for political discourse today has the same problems. It's somewhat convincing bullshit. Should AIs be held to a higher standard than X/Twitter? Than Reddit? Than Fox News? What censorship is appropriate? And, yes, alignment is censorship. Then there's the big problem of chatbots telling you what you seem to want to hear. This is an old problem. "Happy Talk", from South Pacific", is the entertainment version. "Wartime" by Paul Fussell, is the serious version. As the article points out, a small percentage of the population is very vulnerable to certain types of misinformation. It may be the same fraction of the population that's vulnerable to cults. But maybe not. Cults have a group self-reinforcing mechanism and an agenda. Chatbots have neither. Worth studying. The point here is that restrictions on chatbots strong enough to protect the vulnerable would close off most political and social discourse. [1] https://www.youtube.com/watch?v=JXgmQDFhPjo https://www.youtube.com/watch?v=JXgmQDFhPjo
- lazystar 5mo agothe counterpoint is that allowing unlimited discourse places an enourmous amount pf power in the hands of the chatbot owner, who has access to all logs and input from each user. this prevents one chatbot owner from advertising "you can say anything here!!" then using the logs as blackmail down the road.
- scared_together 5mo ago> Should AIs be held to a higher standard than X/Twitter? Than Reddit? Than Fox News? What censorship is appropriate? And, yes, alignment is censorship. Yes, a thousand times yes. Freedom of speech/expression should be a freedom granted to humans. We extend it to corporations based on the practical reality that human speech often requires corporate support to be hosted and published. But as far as I know, AI vendors haven’t claimed that their models represent the views of their founders, employees or any people at all. If we censor AI, which human voice are we censoring?
- xg15 5mo agoI find it somewhat telling that most (not all) of this thread doesn't even attempt to find an answer to the questions posed by the OP but flatly denies the problem of psychological harm exists at all. I feel this is an example of the two larger narratives about AI that currently seem to be forming: For one side, AI is basically every harmful technology ever invented rolled into one: It's harmful to the environment (via waste of energy and resources), it's harmful to the information space (through polluting everything with slop and devaluing human expression), it's harmful to society (by encouraging ever more badly done and unreliable products, by taking away jobs and by replacing human-to-human interaction, by normalizing a mode of development where not even the developers understand what is going on) and it's harmful to whoever uses it personally (by causing ever-growing dependence on AI, either only by skills or even emotionally or psychically, up to the point of AI psychosis and preferring AI agents to other humans). For the other side, AI is the future, the next industrial revolution, the thing that you have to adapt or will be left behind, possibly even the next stage of evolution. Right now, I feel every side is digging in and trying ever harder to ignore the other side. (The AI labs acknowledge "AI risks" in theory - but, as the article pointed out, the risks they perceive and ostensibly work against are so abstract and removed from the everyday use of AI that they more make the point of AI proponents) I feel the end result of this growing tension is the Molotov cocktail in Sam Altmann's home. I'd really like to know more what the tech community at large is trying to do about this rift.