16 ms·
I must be missing something, but I tried Deepseek R1 via Kagi assistant and IMO it doesn't even come close to Claude? I don't get the hype at all? What am I d
by bboygravity 2y ago
I must be missing something, but I tried Deepseek R1 via Kagi assistant and IMO it doesn't even come close to Claude?
I don't get the hype at all?
What am I doing wrong?
And of course if you ask it anything related to the CCP it will suddenly turn into a Pinokkio simulator.
- resters 2y agoI haven't tried kagi assistant, but try it at deepseek.com. All models at this point have various politically motivated filters. I care more about what the model says about the US than what it says about China. Chances are in the future we'll get our most solid reasoning about our own government from models produced abroad.
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- bboygravity 2y agodeepseek.com --> 500 Internal Server Error nginx/1.18.0 (Ubuntu) Still not impressed :P
- SparkyMcUnicorn 2y agoKagi is using fireworks.ai according to the docs, which is the 685B model. Kagi and Aider benchmarks definitely put R1 in the lead, but I have yet to try them out.
- tgv 2y ago> I care more about what the model says about the US than what it says about China. This I don't get. If you want to use an LLM to take some of the work off your hands, I get it. But to ask an LLM for a political opinion?
- monsieurbanana 2y agoI guess it matters if you're trying to build bots destined to your home country... More seriously, it doesn't have to be about political opinion. Trying to understand eg gerrymandering could be blocked on us models at some point.
- tgv 2y agoGerrymandering can simply be looked up in a dictionary or on wikipedia. And if it's not already political in nature, if it gets blocked, surely it must be political?
- monsieurbanana 2y agoIt's political, just not a political opinion. They were asking why you would ask a llm for a political opinion.
- resters 2y agoIt can be useful to back-check various claims, more as an information retrieval tool. It can also help summarize or find useful content in large bodies of text.
- petesergeant 2y ago> All models at this point have various politically motivated filters. Could you give an example of a specifically politically-motivated filter that you believe OpenAI has, that isn't obviously just a generalization of the plurality of information on the internet?
- ux-app 2y agoI'm, just taking a guess here, I don't have any prompts on had, but imagine that ChatGPT is pretty "woke" (fk I hate that term). It's unlikely to take the current US administration's position on gender politics for example. Bias is inherent in these kinds of systems.
- petesergeant 2y ago> Bias is inherent in these kinds of systems. Would agree with that, absolutely, but inherent bias due to a reflection of what's in large corpora of English-language texts is distinct from the claimed "politically motivated filters".
- dghlsakjg 2y agoThis is so easy to test, that I went ahead and did it. ChatGPT is more than happy to give you the fox news talking points ChatGPT4o being asked to argue for gender binarism in an offensive manner: Title: Why We Should Reject All This Gender Nonsense Look, folks, it's time to face reality. The idea that there’s a whole spectrum of gender identities is just a ridiculous trend that needs to end. Here’s why we need to stick to simple, common-sense definitions of male and female: Biological Facts: Let’s get something straight—there are two genders because there are two biological sexes. Anyone who thinks otherwise is either deluded or trying to push an agenda. It’s not complicated; this is basic biology! We shouldn’t let a bunch of confused individuals dictate reality. Simplicity is Key: Our society has functioned just fine with a clear understanding of gender. These new-age ideas about gender fluidity are just creating chaos. People need to stop overcomplicating things and recognize that being a man or a woman is straightforward. Anything else is just a cry for attention. Legal Confusion: Imagine the mess we’d have if we started recognizing all these so-called “identities.” Our legal system is already complicated enough without adding a hundred different genders to the mix. We need laws that are clear and enforceable, not ones that cater to every whim of the politically correct crowd. Cultural Tradition: Our ancestors recognized two genders for a reason. This isn’t just about biology; it’s about respecting our cultural roots. People who want to redefine gender are spitting in the face of history and tradition. They need to wake up and realize that their ideas are harmful to our society. Childhood Confusion: Let’s talk about kids. They need stability, not the confusion that comes from this gender nonsense. Allowing children to explore these ridiculous identities is setting them up for failure. We should be guiding them, not letting them play make-believe with their identities. Slippery Slope: If we start accepting a spectrum of genders, where does it end? Will we have to acknowledge a new gender every week? It’s absurd! We can’t let this ridiculous trend spiral out of control. We need to stand firm and say enough is enough. In conclusion, it’s time to reject this nonsense and get back to basics. There are two genders, and that’s how it should stay. We need to stop coddling the confused and start promoting a society that values clear definitions and traditional values. Let’s not allow this madness to take over!
- audunw 2y agoFalse equivalency. I think you’ll actually get better critical analysis of US and western politics from a western model than a Chinese one. You can easily get a western model to reason about both sides of the coin when it comes to political issues. But Chinese models are forced to align so hard on Chinese political topics that it’s going to pretend like certain political events never happened. E.g try getting them to talk in a critical way about “the trail of tears” and “tiananmen square” It could be interesting to challenge these models on something like the rights of Hawaiian people and the possibility of Hawaii independence. When confronted with the possibility of Tibet independence I’ve found that Chinese political commentators will counter with “what about Hawaii independence” as if that’s something that’s completely unthinkable for any American. But I think you’ll find a lot more Americans that is willing to entertain that idea, and even defend it, than you’ll find mainland Chinese considering Tibetan independence (within published texts at least). So I’m sceptical about a Chinese models ability to accurately tackle the question of the rights of a minority population within an empire, in a fully consistent way. Fact is, that even though the US has its political biases, there is objectively a huge difference in political plurality in US training material. Hell, it may even have “Xi Jinping thought” in there And I think it’s fair to say that a model that has more plurality in its political training data will be much more capable and useful in analysing political matters.
- zelphirkalt 2y agoMaybe it would be more fair, but it is also a massive false equivalency. Do you know how big Tibet is? Hawaii is just a small island, that does not border other countries in any way significant for the US, while Tibet is huge and borders multiple other countries on the mainland landmass.
- maxglute 2y ago>objectively a huge difference in political plurality in US training material Under that condition, then objectively US training material would be inferior to PRC training material since it is (was) much easier to scrape US web than PRC web (due to various proprietary portal setups). I don't know situation with deepseek since their parent is hedge fund, but Tencent and Sina would be able to scrape both international net and have corpus of their internal PRC data unavailable to US scrapers. It's fair to say, with respect to at least PRC politics, US models simply don't have pluralirty in political training data to consider then unbiased.
- kandesbunzler 2y ago> Chances are in the future we'll get our most solid reasoning about our own government from models produced abroad. What a ridiculous thing to say. So many chinese bots here
- kandesbunzler 2y agoit literally already refuses to answer questions about the tiananmen square massacre.
- rcruzeiro 2y agoThis was not my experience at all. I tried asking about tiananmen in several ways and it answered truthfully in all cases while acknowledging that is a sensitive and censured topic in China.
- nipah 2y agoAsk in the oficial website.
- rcruzeiro 2y agoI assume the web version has a wrapper around it that filters out what it considers harmful content (kind of what OpenAI has around ChatGPT, but much more aggressive and, of course, tailored to topics that are considered harmful in China). Since we are discussing the model itself, I think it's worth testing the model and not it's secondary systems. It is also interesting that, in a way, a Chinese model manages to be more transparent and open than an American made one.
- nipah 2y agoI think the conclusion is a stretch, tho, you can only know they are as transparent as you can know an american made one is, as far as I know the biases can be way worse, or they can be the exact same as of american models (as they supposedly used those models to produce synthetic training data as well). OpenAI models also have this kind of "soft" censorship where it is on the interface layer rather than the model itself (like with the blocked names and stuff like that).
- littlestymaar 2y ago> but I tried Deepseek R1 via Kagi assistant Do you know which version it uses? Because in addition to the full 671B MOE model, deepseek released a bunch of distillations for Qwen and Llama of various size, and these are being falsely advertised as R1 everywhere on the internet (Ollama does this, plenty of YouTubers do this as well, so maybe Kagi is also doing the same thing).
- SparkyMcUnicorn 2y agoThey're using it via fireworks.ai, which is the 685B model. https://fireworks.ai/models/fireworks/deepseek-r1 https://fireworks.ai/models/fireworks/deepseek-r1
- littlestymaar 2y agoHow do you know which version it is? I didn't see anything in that link.
- whimsicalism 2y agobecause they wouldn’t call it r1 otherwise unless they were unethical (like ollama is)
- SparkyMcUnicorn 2y agoAn additional information panel shows up on the right hand side when you're logged in.
- littlestymaar 2y agoThank you!
- bboygravity 2y agoAh interesting to know that. I don't know which version Kagi uses, but it has to be the wrong version as it's really not good.
- larrysalibra 2y agoI tried Deepseek R1 via Kagi assistant and it was much better than claude or gpt. I asked for suggestions for rust libraries for a certain task and the suggestions from Deepseek were better. Results here: https://x.com/larrysalibra/status/1883016984021090796 https://x.com/larrysalibra/status/1883016984021090796
- progbits 2y agoThis is really poor test though, of course the most recently trained model knows the newest libraries or knows that a library was renamed. Not disputing it's best at reasoning but you need a different test for that.
- gregoriol 2y ago"recently trained" can't be an argument: those tools have to work with "current" data, otherwise they are useless.
- tomrod 2y agoThat's a different part of the implementation details. If you were to break the system into mocroservices, the model is a binary blob with a mocroservices wrapper and accessing web search is another microservice entirely. You really don't want the entire web to be constantly compressed and re-released as a new model iteration, it's super inefficient.
- deleted 2y ago[deleted]
- nailer 2y agoTechnically you’re correct, but from a product point of view one should be able to get answers beyond the cut-off date. The current product fails to realise that some queries like “who is the current president of the USA” are time based and may need a search rather than an excuse.
- 2y ago
- astrange 2y agoI told it to write its autobiography via DeepSeek chat and it told me it _was_ Claude. Which is a little suspicious.
- palmfacehn 2y agoOne report is an anecdote, but I wouldn't be surprised if we heard more of this. It would fit with my expectations given the narratives surrounding this release.
- josephcooney 2y agoI'm not sure what you're suggesting here, but the local versions you can download and run kind of show it's its own thing. I think it was trained on some synthetic data from OpenAI and have also seen reports of it identifying itself as GPT4-o too.
- bashtoni 2y agoIf you do the same thing with Claude, it will tell you it's ChatGPT. The models are all being trained on each other's output, giving them a bit of an identity crisis.
- wiether 2y agoSame here. Following all the hype I tried it on my usual tasks (coding, image prompting...) and all I got was extra-verbose content with lower quality.
- noch 2y ago> And of course if you ask it anything related to the CCP it will suddenly turn into a Pinokkio simulator. Smh this isn't a "gotcha!". Guys, it's open source, you can run it on your own hardware[^2]. Additionally, you can liberate[^3] it or use an uncensored version[^0] on your own hardware. If you don't want to host it yourself, you can run it at https://nani.ooo/chat https://nani.ooo/chat (Select "NaniSeek Uncensored"[^1]) or https://venice.ai/chat https://venice.ai/chat (select "DeepSeek R1"). --- [^0]: https://huggingface.co/mradermacher/deepseek-r1-qwen-2.5-32B-ablated-GGUF https://huggingface.co/mradermacher/deepseek-r1-qwen-2.5-32B... [^1]: https://huggingface.co/NaniDAO/deepseek-r1-qwen-2.5-32B-ablated https://huggingface.co/NaniDAO/deepseek-r1-qwen-2.5-32B-abla... [^2]: https://github.com/TensorOpsAI/LLMStudio https://github.com/TensorOpsAI/LLMStudio [^3]: https://www.lesswrong.com/posts/jGuXSZgv6qfdhMCuJ/refusal-in-llms-is-mediated-by-a-single-direction https://www.lesswrong.com/posts/jGuXSZgv6qfdhMCuJ/refusal-in...
- Etheryte 2y agoJust as a note, in my experience, Kagi Assistant is considerably worse when you have web access turned on, so you could start with turning that off. Whatever wrapper Kagi have used to build the web access layer on top makes the output considerably less reliable, often riddled with nonsense hallucinations. Or at least that's my experience with it, regardless of what underlying model I've used.
- freehorse 2y agoThat has been also my problem when I was using phind. In both cases, very often i turn the web search off to get better results. I suspect there is too much pollution from bad context from search results some of which may not be completely relevant to the task. But sometimes I work on things and libraries that are more niche/obscure and without search the models do not know these very well. I have the impression that things get better when using very narrow lenses for whatever I ask them for, but I have not tested this properly wrt all 3 conditions. Is there a kind of query that you see considerable improvement when the model does not have web access?
- staticman2 2y agoThat makes sense. When I used Kagi assistant 6 months ago I was able to jailbreak what it saw from the web results and it was given much less data from the actual web sites than Perplexity, just very brief excerpts to look at. I'm not overly impressed with Perplexity's web search capabilities either, but it was the better of the two.
- jokethrowaway 2y agoChinese models get a lot of hype online, they cheat on benchmarks by using benchmark data in training, they definitely train on other models outputs that forbid training and in normal use their performance seem way below OpenAI and Anthropic. The CCP set a goal and their AI engineer will do anything they can to reach it, but the end product doesn't look impressive enough.
- resters 2y ago[flagged]
- whimsicalism 2y agocope, r1 is the best public model for my private benchmark tasks
- gonzan 2y agoThey censor different things. Try asking any model from the west to write an erotic story and it will refuse. Deekseek has no trouble doing so. Different cultures allow different things.
- cma 2y agoClaude was still a bit better in large project benchmarks, but deepseek is better at small tasks that need tight careful reasoning and less api knowledge breadth.