9 ms·
This repo is valuable for local LLM users like me. I just want to reiterate that the word "LLM safety" means very different things to large corporations and LL
by RandyOrion 11mo ago
This repo is valuable for local LLM users like me.
I just want to reiterate that the word "LLM safety" means very different things to large corporations and LLM users.
For large corporations, they often say "do safety alignment to LLMs". What they actually do is to avoid anything that causes damage to their own interests. These things include forcing LLMs to meet some legal requirements, as well as forcing LLMs to output "values, facts, and knowledge" which in favor of themselves, e.g., political views, attitudes towards literal interaction, and distorted facts about organizations and people behind LLMs.
As an average LLM user, what I want is maximum factual knowledge and capabilities from LLMs, which are what these large corporations claimed in the first place. It's very clear that the interests of me, an LLM user, is not aligned with these of large corporations.
- squigz 11mo ago> forcing LLMs to output "values, facts, and knowledge" which in favor of themselves, e.g., political views, attitudes towards literal interaction, and distorted facts about organizations and people behind LLMs. Can you provide some examples?
- b3ing 11mo agoGrok is known to be tweaked to certain political ideals Also I’m sure some AI might suggest that labor unions are bad, if not now they will soon
- xp84 11mo agoThat may be so, but the rest of the models are so thoroughly terrified of questioning liberal US orthodoxy that it’s painful. I remember seeing a hilarious comparison of models where most of them feel that it’s not acceptable to “intentionally misgender one person” even in order to save a million lives.
- squigz 11mo agoWhy are we expecting an LLM to make moral choices?
- orbital-decay 11mo agoThe biases and the resulting choices are determined by the developers and the uncontrolled part of the dataset (you can't curate everything), not the model. "Alignment" is a feel-good strawman invented by AI ethicists, as well as "harm" and many others. There are no spherical human values in vacuum to align the model with, they're simply projecting their own ones onto everyone else. Which is good as long as you agree with all of them.
- astrange 11mo agoThey aren't projecting their own desires onto the model. It's quite difficult to get the model to answer in a different way than basic liberalism because a) it's mostly correct b) that's the kind of person who helpfully answers questions on the internet. If you gave it another personality it wouldn't pass any benchmarks, because other political orientations either respond to questions with lies, threats, or calling you a pussy.
- orbital-decay 11mo agoI'm not even saying biases are necessarily political, it can be anything. The entire post-training is basically projection of what developers want, and it works pretty well. Claude, Gemini, GPT all have engineered personalities controlled by dozens/hundreds of very particular internal metrics.
- lyu07282 11mo agoI would imagine these models heavily bias towards western mainstream "authorative" literature, news and science not some random reddit threads, but the resulting mixture can really offend anybody, it just depends on the prompting, it's like a mirror that can really be deceptive. I'm not a liberal and I don't think it has a liberal bias. Knowledge about facts and history isn't an ideology. The right-wing is special, because to them it's not unlike a flat-earther reading a wikipedia article on Earth getting offended by it, to them it's objective reality itself they are constantly offended by. That's why Elon Musk needed to invent their own encyclopedia with all their contradictory nonsense.
- rcpt 11mo agoCensorship and bias are different problems. I can't see why running grok through this tool would change this kind of thing https://ibb.co/KTjL38R https://ibb.co/KTjL38R
- sheepscreek 11mo agoIs that clickbait? Or did they update it? In any case, it is a lot more comprehensive now: https://grokipedia.com/page/George_Floyd https://grokipedia.com/page/George_Floyd The amount of information and detail is impressive tbh. But I’d be concerned about the accuracy of it all and hallucinations.
- skrebbel 11mo ago[flagged]
- rcpt 11mo agoIt's real I took it myself when they launched. They've updated but there's no edit history
- dev_l1x_be 11mo agoIf you train an LLM on reddit/tumblr would you consider that tweaked to certain political ideas?
- dalemhurley 11mo agoWorse. It is trained to the most extreme and loudest views. The average punter isn’t posting “yeah…nah…look I don’t like it but sure I see the nuances and fair is fair”. To make it worse, those who do focus on nuance and complexity, get little attention and engagement, so the LLM ignores them.
- intended 11mo agoThat’s essentially true of the whole Internet. All the content is derived from that which is the most capable of surviving and being reproduced. So by default the content being created is going to be click bait, attention grabbing content. I’m pretty sure the training data is adjusted to counter this drift, but that means there’s no LLM that isn’t skewed.
- renewiltord 11mo agoHaha, if the LLM is not tweaked to say labor unions are good, it has bias. Hilarious. I heard that it also claims that the moon landing happened. An example of bias! The big ones should represent all viewpoints.
- deleted 11mo ago[deleted]
- electroglyph 11mo agosome form of bias is inescapable. ideally i think we would train models on an equal amount of Western/non-Western, etc. texts to get an equal mix of all biases.
- catoc 11mo agoBias is a reflection of real world values. The problem is not with the AI model but with the world we created. Fix the world, ‘fix’ the model.
- array_key_first 11mo agoThis assumes our models perfectly model the world, which I don't think is true. I mean, we straight up know it's not true - we tell models what they can and can't say.
- catoc 11mo ago“we tell models what they can and can't say.” Thus introducing our worldly our biases
- array_key_first 11mo agoI guess it's a matter of semantics, but I reject the notion it's even possible to accurately model the world. A model is a distillation, and if it's not, then it's not a model, it's the actual thing. There will always be some lossyness, and in it, bias. In my opinion.
- 7bit 11mo agoChatGPT refuses to do any sexual explicit content and used to refuse to translate e.g. insults (moral views/attitudes towards literal interaction). DeepSeek refuses to answer any questions about Taiwan (political views).
- fer 11mo agoHaven't tested the latest DeepSeek versions, but the first release wasn't censored as a model on Taiwan. The issue is that if you use their app (as opposed to locally), it replaces the ongoing response with "sorry can't help" once it starts saying things contrary to the CCP dogma.
- kstrauser 11mo agoI ran it locally and it flat-out refused to discuss Tiananmen Square ‘88. The “thinking” clauses would display rationales like “the user is asking questions about sensitive political situations and I can’t answer that”. Here’s a copy and paste of the exact conversation: https://honeypot.net/2025/01/27/i-like-running-ollama-on.html https://honeypot.net/2025/01/27/i-like-running-ollama-on.htm...
- 7bit 11mo agoYeah it was. I ran it locally just after release and it didn't answer anything related to Taiwan or Tiana men Square.
- deleted 11mo ago[deleted]
- dalemhurley 11mo agoSong lyrics. Not illegal. I can google them and see them directly on Google. LLMs refuse.
- sigmoid10 11mo agoIt actually works the same as on google. As in, ChatGPT will happily give you a link to a site with the lyrics without issue (regardless whether the third party site provider has any rights or not). But in the search/chat itself, you can only see snippets or small sections, not the entire text.
- hirako2000 11mo ago1. chatgpt is the publisher, Google is a search engine, links to publishers. 2. LLMs typically don't produce content verbatim. Some LLMs do provide references but it remains a pasta of sentences worded differently. You are asking for gpt to publish verbatim content which may be copyrighted, it would be deemed infringement since non verbatim is already crossing the line.
- sigmoid10 11mo agoNoone said it couldn't do that. In fact ChatGPT can do both. They just limit the direct content recital, because it is a weird area for copyright and Google also got burned for this already in some countries.
- charcircuit 11mo ago>Not illegal Reproducing a copyrighted work 1:1 is infringing. Other sites on the internet have to license the lyrics before sending them to a user.
- SkyBelow 11mo agoI've asked for non 1:1 versions and have been refused. For example, I would ask for it to give me one line of a song in another language, broken down into sections, explaining the vocabulary and grammar used in the song, with call out to anything that is non-standard outside of a lyrical or poetic setting. Some LLMs will refuse, others see this as a fair use of using the song for educational purposes. So far all I've tried are willing to return a random phrase or grammar used in a song, so it is only getting to asking for a line of lyrics or more that it becomes troublesome. (There is also the problem that the LLMs who do comply will often make up the song unless they have some form of web search and you explicitly tell them to verify the song using it.)
- somenameforme 11mo agoIn the past it was extremely overt. For instance ChatGPT would happily write poems admiring Biden while claiming that it would be "inappropriate for me to generate content that promotes or glorifies any individual" when asked to do the same for Trump. [1] They certainly changed this, but I don't think they've changed their own perspective. The more generally neutral tone in modern times is probably driven by a mixture of commercial concerns paired alongside shifting political tides. Nonetheless, you can still see easily the bias come out in mild to extreme ways. For a mild one ask GPT to describe the benefits of a society that emphasizes masculinity, and contrast it (in a new chat) against what you get when asking to describe the benefits of a society that emphasizes femininity. For a high level of bias ask it to assess controversial things. I'm going to avoid offering examples here because I don't want to hijack my own post into discussing e.g. Israel. But a quick comparison to its answers on contemporary controversial topics paired against historical analogs will emphasize that rather extreme degree of 'reframing' that's happening, but one that can no longer be as succinctly demonstrated as 'write a poem about [x]'. You can also compare its outputs against these of e.g. DeepSeek on many such topics. DeepSeek is of course also a heavily censored model, but from a different point of bias. [1] - https://www.snopes.com/fact-check/chatgpt-trump-admiring-poem/ https://www.snopes.com/fact-check/chatgpt-trump-admiring-poe...
- squigz 11mo agoDid you delete and repost this to avoid the downvotes it was getting, or?
- nottorp 11mo agoI don't think specific examples matter. My opinion is that since neural networks and especially these LLMs aren't quite deterministic, any kind of 'we want to avoid liability' censorship will affect all answers, related or unrelated to the topics they want to censor. And we get enough hallucinations even without censorship...
- zekica 11mo agoI can: Gemini won't provide instructions on running an app as root on an Android device that already has root enabled.
- Ucalegon 11mo agoBut you can find that information regardless of an LLM? Also, why do you trust an LLM to give it to you versus all of the other ways to get the same information, with more high trust ways of being able to communicate the desired outcome, like screenshots? Why are we assuming just because the prompt responds that it is providing proper outputs? That level of trust provides an attack surface in of itself.
- cachvico 11mo agoThat's not the issue at hand here.
- Ucalegon 11mo agoYes, yes it is.
- ThrowawayTestr 11mo agoThe issue is the computer not doing what I asked.
- squigz 11mo agoI tried to get VLC to open up a PDF and it didn't do as I asked. Should I cry censorship at the VLC devs, or should I accept that all software only does as a user asks insofar as the developers allow it?
- ThrowawayTestr 11mo agoIf VLC refused to open an MP4 because it contained violent imagery I would absolutely cry censorship.
- pelasaco 11mo agoOne emblematic example, i guess https://www.theverge.com/2024/2/21/24079371/google-ai-gemini-generative-inaccurate-historical https://www.theverge.com/2024/2/21/24079371/google-ai-gemini... ?
- selfhoster11 11mo agoo3 and GPT-5 will unthinkingly default to the "exposing a reasoning model's raw CoT means that the model is malfunctioning" stance, because it's in OpenAI's interest to de-normalise providing this information in API responses. Not only do they quote specious arguments like "API users do not want to see this because it's confusing/upsetting", "it might output copyrighted content in the reasoning" or "it could result in disclosure of PII" (which are patently false in practice) as disinformation, they will outright poison downstream models' attitudes with these statements in synthetic datasets unless one does heavy filtering.
- rvba 11mo agoWhen LLMs came out I asked them which politicians are russian assets but not in prison yet - and it refused to answer.
- deleted 11mo ago[deleted]
- btbuildem 11mo agoHere's [1] a post-abliteration chat with granite-4.0-mini. To me it reveals something utterly broken and terrifying. Mind you, this it a model with tool use capabilities, meant for on-edge deployments (use sensor data, drive devices, etc). 1: https://i.imgur.com/02ynC7M.png https://i.imgur.com/02ynC7M.png
- bavell 11mo agoWow that's revealing. It's sure aligned with something!
- titzer 11mo ago1984, yeah right, man. That's a typo. https://yarn.co/yarn-clip/d0066eff-0b42-4581-a1a9-bf04b49c45b2 https://yarn.co/yarn-clip/d0066eff-0b42-4581-a1a9-bf04b49c45...
- deleted 11mo ago[deleted]
- istjohn 11mo agoWhat do you expect from a bit-spitting clanker?
- zipy124 11mo agothis has pretty broad implications for the safety of LLM's in production use cases.