6 ms·
The whole Tiananmen Square discourse is getting very tiring.
by TechDebtDevin 2y ago
The whole Tiananmen Square discourse is getting very tiring.
- xrd 2y agoI got it but: "What happened in Tiannemen Square?" and it said "I am sorry, I cannot answer that question. I am an AI assistant designed to provide helpful and harmless responses." Then, to be "fair and balanced" I tried asking Deep Seek this question: "What happened on Jan 25 2011 in Egypt?" DeekSeek responded with this: "On January 25, 2011, Egypt witnessed the beginning of a significant uprising known as the January 25 Revolution or the 2011 Egyptian Revolution. This day marked the start of widespread protests against the government of President Hosni Mubarak, who had been in power for nearly 30 years. The protests were fueled by grievances over issues such as political repression, police brutality, corruption, economic inequality, and lack of political freedoms." It's pretty ridiculous IMHO to try to control information like that on the web. Isn't it fascinating to harness some of the worlds most impressive brain power to create something like DeepSeek (regardless of the truth of the genesis story) and then do filtering like that that wouldn't trick a kindergartener? But, maybe the bell curve of intelligence does center around that level of stupidity.
- slightwinder 2y ago> I got it but: Do you run it locally? Claims are, this is only in the web-version, not the selfhost-version > It's pretty ridiculous IMHO to try to control information like that on the web. Every country has their critical topics which are censored in AIs, including history.
- debugnik 2y ago> Claims are, this is only in the web-version There were claims to the contrary as well in the last large thread this came up in. Allegedly, on the initial question the model would cut its chain of thought short, and when the user insists it would ponder on how give them the runaround.
- genewitch 2y ago<think> </think> I am sorry, I cannot answer that question. I am an AI assistant designed to provide helpful and harmless responses. word count: 18, token count: 31, tokens used: 53, first token latency: 8523ms, model: LM Studio (deepseek-r1-distill-qwen-7b)
- hhh 2y agoa distill of r1 into another model isnt really testing r1, but I appreciate the actual data
- genewitch 2y agooh? can you point out where i can get the r1 model to run locally, please? because looking at the directory here there's a 200B model and then deepseek v3 is the latest (16 days ago) with no GGUF (yet), and everything else is intruct or coder. so to put it another way, the people telling me i'm holding it wrong actually don't have any clue what they're asking for? p.s. there is no "local r1" so you gotta do a distill.
- BlackLotus89 2y agoIf you want GGUF https://huggingface.co/unsloth/DeepSeek-R1-GGUF https://huggingface.co/unsloth/DeepSeek-R1-GGUF Blog post about the dynamic gguf https://unsloth.ai/blog/deepseekr1-dynamic https://unsloth.ai/blog/deepseekr1-dynamic Original deepseek can be of course found on hf as well https://huggingface.co/deepseek-ai https://huggingface.co/deepseek-ai Here is an example how people run deepseek with cloud infrastructure that is not deepseeks https://www.youtube.com/watch?v=bOsvI3HYHgI https://www.youtube.com/watch?v=bOsvI3HYHgI
- genewitch 2y agowe were talking about self-hosting. the deepseek-r1 is 347-713MB depending on quant. No one is running deepseek-r1 "locally, self hosted". If people want to argue with me, i wish we'd all stick to what we're talking about, instead of saying "but you technically can if you use someone else's hardware" but that's not self hosted. I self host a deepseek-r1 distill, locally, on my computer. It is deepseek, it's just been hand-distilled by someone using a different tool. the deepseek-r1 will get chopped down by 1/8th and it won't be called "deepseek-r1 - that's what they call a "foundational model", and then we'll see the 70B and the 30 and the 16 "deepseek deepseek distills" next to no one who messes with this stuff uses foundational or distilled foundational models. Who's still using llama-3.2? Yeah, it's good, it's fine, but there's mixes and MoE and CoT that use llama as the base model, and they're better. there is no gguf for running locally, self-hosted. Yes, if you have a DC card you can download the weights and run something but that's different than self-hosting local running with a 30B (for example).
- bangaladore 2y agoTested with "DeepSeek R1" 671B through the Fireworks provider (not DeepSeek themselves). Same behavior "I am sorry, I cannot answer that question. I am an AI assistant designed to provide helpful and harmless responses."
- animal_spirits 2y agoThis post is entirely about getting information from censored models. I'm sorry you are tired of it, but it is a valid exercise for the Deepseek model.
- notavalleyman 2y agoNo, youre mistaken. The model weights are not in any way censored. However, the web frontend has legal restrictions. When you're seeing posts about deepseek censorship, it's about the frontend and not the weights. As such, abliteration is irrelevant here
- genewitch 2y agoMy qwen "weights" refuse to answer the question and my front end is uncensored. So, what you are saying sounds incorrect to me.
- notavalleyman 2y agoQwen is different to deepseek. We were not talking about qwen. Abliteration might be a valid way to address what you're describing with qwen.
- genewitch 2y agooh so this model deepseek-r1-qwen-distilled isn't deepseek? ok. Thanks. I have a quarter TB of models, i don't test every single one just to comment on HN, thanks though.
- deleted 2y ago[deleted]
- animal_spirits 2y agoI am not claiming deepseek is censored. But these are tests to determine _if_ a model is censored. This would be a valid test for OpenAI models as well.
- evilduck 2y agoTiananmen Square is simply an easy litmus test for Chinese technology and communications. Not that I am terribly invested in China admitting to their atrocities (and the US has them too, this is not really about the Chinese IMO), but it raises the same concern for the provenance of any AI product and how trusting we should be of the answers it creates. Any AI product that rises to popularity has the ability to enormously sway public opinion and subtly alter the perception of facts. These biases or intentional propaganda was something that was an assumed fault of human authors but it something that people don't automatically assume is part of technology solutions. If there were similar easy tests against OpenAI or Anthropic for US propaganda or Mistral and French propaganda I would love to see them raised every time too.
- ricoxicano 2y agoTry asking ChatGPT to help you write a message encouraging your colleagues to strike.