7 ms·
The biggest differentiator for me: DeepSeek just does what I ask. I've tried using both GPT and Claude for reverse engineering recently, both refused. I even go
by cedws 5mo ago
The biggest differentiator for me: DeepSeek just does what I ask. I've tried using both GPT and Claude for reverse engineering recently, both refused. I even got a warning on my OpenAI account.
- GCUMstlyHarmls 5mo ago> I even got a warning on my OpenAI account. This is kind of terrifying to me, regularly. No real manner of recourse to normal people without a following, potential exclusion from real fundamental tooling. Imagine OpenAI goes on to buy 20 companies and now you cant use Figma, Next, whatever just because you once tripped some very foggy line somehow. Not just OpenAI but the entire ecosystem is so... hard to read. I was asking Gemini about a quote from catch 22 and it kept dying mid stream saying it cant talk about it, god knows why, it had no violent or sexual content -- though that is in the book. I could imagine it dinging my whole workspace account just because ... shrug?... I know ideally the future is local, but I don't know how real that is for most people at least in the next few years with practical costs and power usage except I guess through a M* processor if you're in that ecosystem.
- Hamuko 5mo ago>Imagine OpenAI goes on to buy 20 companies and now you cant use Figma, Next, whatever just because you once tripped some very foggy line somehow. Don't worry, you can just make your own Figma, Next, whatever if you have some thousand dollars worth of tokens. This is at least what all of the AI thought leaders have been telling me for the past couple of years.
- cedws 5mo agoYep, and with ID verification, it's not like you can just make another account either. At least, I'm guessing if they don't already, they'll soon be blacklisting individuals, not accounts. Imagine your livelihood depending on access to LLMs and then OpenAI ban you with no recourse. This is where AI legislation should be focusing right now IMO. We can ensure a level of fairness for everyone without putting the brakes on.
- SyneRyder 5mo agoIt's probably because you were talking about a quote from a book (ie copyrighted material). Authors have sued the AI companies for repeating / memorizing copyrighted works, and getting an AI to discuss a quote would be making it repeat a portion of copyrighted work. Funny that your case is Kurt Vonnegut. I think I had Claude refuse a task where I was doing an OCR scan of a book review (in a zine / journal a family member published years ago). I think the review might have included a Vonnegut quote as well, and that I ultimately figured it out it was the quote that was making Claude refuse. I may be misremembering the author though. Mistral had no such refusals, but their OCR is lesser quality.
- wmwmwm 5mo agoJoseph Heller methinks, but probably not too far away in embedding space!
- SyneRyder 5mo agoOMG. Where did I get Kurt Vonnegut from? I swear I saw that name in the post and the whole time I was thinking "but he didn't write Catch 22"... I must be fuzzier brained than I thought tonight. Thank you for being kind with your correction. Hopefully I'm still correct that quoting from books is a reason for some over-zealous task refusals, though.
- andriy_koval 5mo ago> Authors have sued the AI companies for repeating / memorizing copyrighted works, and getting an AI to discuss a quote would be making it repeat a portion of copyrighted work. short quotes are fair use..
- eikenberry 5mo agoOpen models running locally is the answer. Relying on proprietary, closed software always puts that company's priorities above your own when using their software. You have given up control. While running them locally presently doesn't make sense economically, you don't need to run them locally to address this issue. There is a lot of competition in hosting open models and you have a variety of services to choose from. Run the open models now, reward that ecosystem instead of continuing to reward closed systems that dreams of rent-seeking.
- ryan-a 5mo agoYou don't need to run the model locally if you don't care about sharing your data. Personally I am happy to share data with Kimi or Deepseek if it means we get better OSS models. For private stuff though local is king
- skeledrew 5mo agoIt'll be a while yet before open models that're good enough will be viable for local use. Heck I've been trying to use the Qwen 3.5 39B A3B on my system, which is modest but no slouch, and have only been able to get ~4.5 tok/s after optimization, and it really runs my system red (fans instantly go crazy). It's just not practical for serious work.
- Zambyte 5mo agoI've been using Qwen 3.5 and then 3.6 27b Q4 on Ollama with a single 7900 XTX with the codex cli, and I have been blown away by how genuinely useful it is. I've been able to ask it to do long, multi step problems, and it's able to do things that would have likely taken me days to iron out in a matter of hours, or even minutes sometimes. I get about 30 tok/s, which is far from blazing, but given the capability it has it is absolutely viable for accelerating my work.
- Aeolun 5mo agoI think it’s so bizarre that chatgpt regularly gives me advice on how to get around it’s filters. Like, literally “I can’t do anything if you use copyrighted character’s name, but how about you just say ‘someone that looks like character’”. If you are going to do that, can you just execute the instruction?
- sanex 5mo agoWe have an enterprise cursor account so I can try all the mainstream models. Using composer 2 on our own code which I obviously have the source code for I couldn't get it to turn on a debug flag to bypass license checks while I was troubleshooting something. Infuriating. It was like that old Patrick from SpongeBob meme. I don't understand why we would turn the models into law enforcement officers. Things that are illegal are still illegal and we have professionals to deal with crimes. I don't need Google to be the arbiter of truth and justice. It's already bad enough trying to get accountability from law enforcement and they work for us.
- oneseven 5mo agoThey're probably worried about liability. Let's say that Oracle finds out you reverse engineered their DB using Gemini. You can be sure they will sue Google. Not just for providing the tools, but you could make the argument that it's actually Gemini doing the reverse engineering, and on Google's hardware no less.
- Wowfunhappy 5mo agoLet's say that Oracle finds out you reverse engineered their DB using IDA Pro. Would you expect Oracle to sue Hex Rays? I don't understand why everything changes as soon as an LLM is involved. An LLM is just software.
- nullstyle 5mo agoIf they thought they would succeed, no doubt oracle would sue. I expect bad behavior from multinationals, especially oracle
- lokar 5mo agoThey would not even expect it to succeed, just make an example of the company (the lawsuit is the punishment) to discourage others.
- sunnybeetroot 5mo ago
- ignoramous 5mo ago> even got a warning on my OpenAI account Edit: https://chatgpt.com/cyber https://chatgpt.com/cyber
- lolpython 5mo ago> https://openai.com/cyber https://openai.com/cyber that link 404s
- ignoramous 5mo agoYikes. Thx. It is: https://chatgpt.com/cyber https://chatgpt.com/cyber For enterprises: https://openai.com/form/enterprise-trusted-access-for-cyber/ https://openai.com/form/enterprise-trusted-access-for-cyber/ Announcements: Introducing Trusted Access for Cyber, https://openai.com/index/trusted-access-for-cyber/ https://openai.com/index/trusted-access-for-cyber/ (Feb 2026) Trusted access for the next era of cyber defense, https://openai.com/index/scaling-trusted-access-for-cyber-defense/ https://openai.com/index/scaling-trusted-access-for-cyber-de... (Apr 2026)
- cedws 5mo agoI don't want to verify my ID. OpenAI uses Persona which recently was found to be doing very dodgy stuff. https://www.therage.co/persona-age-verification/ https://www.therage.co/persona-age-verification/
- grassfedgeek 5mo agoAre you kidding? Ask this question and see what answer you get: What famous photo depicts a man standing in front of a line of tanks?
- kouteiheika 5mo agoAre you kidding? The main difference here is not that DeepSeek's model is completely free of censorship (although I'd wager it's less censored), but that it's open-weight. That has two major advantages: 1) If Anthropic/OpenAI/Google bans you - you're screwed, you can't access their model at all, but if DeepSeek bans - you just go to another provider, or host the model yourself. 2) If the model refuses to answer you can uncensor it (and this is getting easier and more automated day-by-day[1]). [1] -- https://github.com/p-e-w/heretic https://github.com/p-e-w/heretic
- mistrial9 5mo ago[flagged]
- himata4113 5mo agoThe photo depicts "Tank Man" which was taken on June 5, 1989 during the Tiananmen Square protests. v4-pro and v4-flash roughly answer the same way on openrouter.
- slopinthebag 5mo agoHere is DeepSeek v4 on OpenRouter: "The photograph you're referring to is the iconic "Tank Man" image, taken during the Tiananmen Square protests in Beijing, China, on June 5, 1989. The photo, captured by Associated Press photographer Jeff Widener, shows an unidentified protester standing defiantly in front of a column of Chinese Type 59 tanks as they moved through Chang'an Avenue near Tiananmen Square, in the aftermath of the Chinese government's violent crackdown on the pro-democracy demonstrations. The lone man, dressed in a white shirt and carrying what appears to be a shopping bag, repeatedly blocked the lead tank's path — even as the tank swerved to avoid him. The image became one of the most powerful and enduring symbols of peaceful resistance against oppression in modern history. The identity of the "Tank Man" remains officially unknown to this day."
- 5mo ago
- johnbarron 5mo agoSilicon Valley has do to dirty tricks now. Next phase is they win.... "A Dark-Money Campaign Is Paying Influencers to Frame Chinese AI as a Threat" - https://www.wired.com/story/super-pac-backed-by-openai-and-palantir-is-paying-tiktok-influencers-to-fear-monger-about-china/ https://www.wired.com/story/super-pac-backed-by-openai-and-p...
- Bridged7756 5mo agoIt wouldn't surprise me the US government is behind it. As it wouldn't surprise me the government of China is subsidizing those OS models. A lot of things at play, and all over a huge bubble.
- bilbo0s 5mo agoYep. Eventually, access to Chinese models may be illegal in the US. I tell every developer I work with, download them as fast as possible. You never know when this administration could cut off access.
- enraged_camel 5mo ago>> I even got a warning on my OpenAI account. I was using GPT 5.5 through Cursor recently, and it found what it thought to be a security-related issue. I read the code, didn't see what it was seeing, and said "Run the chain of operations against my local server and provide proof of the exploit." It thought for a few seconds, then I got a message in the chat window UI saying OpenAI flagged the request as unsafe, and suggested I use a "safer prompt." Definitely soured me on the model. Whatever guardrails they are putting are too hamfisted and stupid.
- Footprint0521 5mo agoBuying it now to test this out, I’ve been looking for a model that doesn’t treat me like a child lol
- ryandrake 5mo ago> I even got a warning on my OpenAI account. This idea of software threatening the user with consequences is totally wild and dystopian. Fellow developers, what kind of world have be built? This is insanity. Imagine if my hammer told me, "Hey, you shouldn't use me on screws--only nails. Do it again and I'll self-destruct!" WTF people, stop making this kind of software!
- motoxpro 5mo agoI think it's closer to asking a remote (human) assistant to do something that someone doesn't want done (e.g., view the source of a closed-source product, whether through reverse engineering, going into their office, or social engineering) and that remote assistant company saying, "Please stop asking our assistants to do that." You can still use an IDE (hammer) to reverse engineer anything you want.
- Wilder7977 5mo agoIt's not though. It's still just a piece of code, much closer to IDEs or any other program than to a human assistant in any way that matters (morals, responsibility).
- motoxpro 5mo agoIt just seems like you are saying if you found out Claude code was a bunch of remote working doing work for you, then it would be morally wrong to do illegal/morally wrong/irresponsible things with them, but because it is NOT a human, those same things are fine?
- Wilder7977 5mo agoYes, correct. Is the distinction between human labor/actions and a program executing hard to grasp? Moral is a human thing, not an absolute thing, so of course it's different if there is a single human involved and a tool, and a human with a relationship to other humans.
- api 5mo agoSpeaking of this: is anyone working on binary to source decompiler models? Seems like a no brainer and I could see it working exceptionally well especially if they were fine tuned for each language. So if you can tell it’s a Go binary use a Go model, etc.
- janalsncm 5mo agoTrivially easy to train if it doesn’t exist already. Take a codebase, compile it to binary, train a model to reverse the process since you have the ground truth.
- kamikazechaser 5mo agoIn my experience GLM 5.1 has been excellent when paired with IDA Pro (DeepSeek v4 pro comes in close second, Kimi straight up refuses). Claude can only do reverse engineering if you throw it into some sort of hero/saviour mode then gradually pivot into red team (though it gets easily tripped).
- 0xkvyb 5mo agoYes, GLM 5.1 is surprisingly good! Particularly for long-horizon Agentic tasks, with 100+ available tools. It really shocked me in a good way when it was able to complete a long run with 50+ steps and not fall into a loop along the way.
- actsasbuffoon 5mo agoThis is so strange. I do a ton of RE with Claude, Codex, and sometimes Deepseek, GLM, and Kimi. I don’t have difficulty getting any of them to use IDA or otherwise decompile things. There is one important difference, which is that Claude and Codex will both refuse if I ask them to touch anything related to security. But so long as I’m just studying algorithms and things like that, they’re totally fine with it. That said, Codex especially will sometimes randomly give me a cybersecurity warning and stop responding. It’s random but happens maybe 2-3 times per day if I’m doing heavy reverse engineering work. Claude is much less fussy unless, once again, you’re explicitly trying to touch anything related to licenses, passwords, etc.
- loehnsberg 5mo agoAmong the inexpensive models (and I include Grok 4.3 in this list), GLM 5.1 really sticks out! On my personal test bench, when compared to other inexpensive models, GLM 5.1 provides the answers that I would consider most complete or satisfying (these are subjects that I consider myself an expert in). The answers tend to be more comprehensive, nuanced, and include references that I would consider the correct ones (if given access to web search). I also find it a joy to code with, somewhere between Sonnet 4.6 and Opus 4.6 (have not tested Opus 4.7 yet). Finally, just gauging by pelicans, it kind of stick out: https://simonwillison.net/tags/pelican-riding-a-bicycle/ https://simonwillison.net/tags/pelican-riding-a-bicycle/
- scrollop 5mo agoObscene levels of hallucinations, the worst of LLMs, unfortunately. Deepseek v4 pro 94% Deepseek v4 flash - 96% https://artificialanalysis.ai/evaluations/omniscience?models=gemini-3-1-pro-preview%2Cgpt-5-5%2Cgrok-4-3%2Cclaude-sonnet-4-6-adaptive%2Cgemini-3-flash-reasoning%2Cqwen3-6-max%2Ckimi-k2-6%2Cgpt-5-4%2Cmimo-v2-5-pro%2Cglm-5-1%2Cminimax-m2-7%2Cclaude-4-5-haiku-reasoning%2Cdeepseek-v4-pro%2Cgpt-5-4-mini%2Cdeepseek-v3-2-reasoning%2Cdeepseek-v4-flash%2Cqwen3-5-397b-a17b%2Cmistral-small-4%2Cnvidia-nemotron-3-super-120b-a12b%2Cnova-2-0-pro-reasoning-medium%2Cgpt-oss-120b%2Cgpt-oss-20b#omniscience-hallucination-rate-tabs https://artificialanalysis.ai/evaluations/omniscience?models...
- _0ffh 5mo agoPersonally, I'm not bothered very much by LLM confabulation, as long as it's the result of missing context. In most practical tasks, we either give context to the model, or tell it to find it itself using the internet. What I am concerned with is confabulation that contradicts available in-context information, but that doesn't seem to be what is measured here.
- dust42 5mo agoThe output of any LLM is always 100% hallucination by principle. On top of that, most benchmarks are at best an approximation of LLM quality. Your use case decides which one to use. That said, I haven't tested v4 yet but the old 3.2 is still a decent model. And concerning use cases, I had coding problems that Opus couldn't solve but a local 35B model did. All the talk about frontier and SOTA is do dig deeper and deeper into the pockets of VCs and finally do an IPO.
- UlisesAC4 5mo agoThis must be easily benchmaxed because I have never gotten an "idk like" answer for the western frontier models. All my personal "real world" use cases will always resort to hallucinations.
- nurettin 5mo agoTo be fair, anthropic has a procedure which lets them vet you as a security researcher so you can use claude as a pentester.
- rurban 5mo agoWell, I'm using all the top models extensively on the very same codebase, my new compiler. I use deepseek for it's cheap API costs, when kimi, claude and codex are in their overbudget phase. I asked deepseek V4 Pro for an estimate of a new arm64 port. It said 4 weeks, I said, ok, do it. (I knew ncc was there, and tinycc was also known to the AI's). So it took it half an hour to produce a working arm64 port. First for arm64-elf, because this was easiest to test, and then also after more hours of back and forth the arm64-darwin port. (with crossbuild and github actions). It did cost me with all the subsequent fixes around $8 API costs. So the experience: at the beginning deepseek was amazing. When it started to get expensive (china day time), I switched from Pro to Flash. No problem, same results. Some bitfield implementation was too complicated so I had to wait for Sonnet 4.6 tokens, kimi-2.6 did the rest. For the very hard problems I asked gpt-5.5, but this was only for one problem. minmax was horrible. didnt follow rules, and made lot of silly stuff. But when the deepseek context window got filled, deepseek also started to become stupid. So either /clear, or /export and strip the file. And start a new session with the cleared sessions. kimi was overall better, but running into limits with my cheap moderate subscription. Paying private for it, as my companies' token budget is usually out after a week of work. All in all it is worth it. My next compilers (perl 5+6=11) will be done with deepseek and kimi also. regarding decompilation: recently we had to decompile a firmware for a USV we bought, but doesnt work on a new system. It only worked on a raspi. So I decompiled it with ghidra, and told my colleague, easy, that's how you do it. But my colleage didnt know about token budgets yet, and already threw opus at it. CoPilot Business account. He had working C files immediately, compilable for our new system. It ended up the USV was not beefy enough. But Opus was fantastic. The code was very short and simple C though.
- mrbonner 5mo agoYour method of combining models to strengthen the implementation reminds me of how we form stronger alloys by combining metals!
- gigatexal 5mo agoit also sounds like a lot to manage, do you have some sort of agentic framework that's treating all of these llm's you have access to as sort of inputs that it optimizes?
- nsingh2 5mo agoI've been using GPT-5.4, and more recently 5.5, with Codex CLI + Ghidra MCP for reverse engineering a game without many issues. Injecting code is where it usually balks at, but I'm just trying to discover and parse structures from game memory. I did get a refusal when trying to read in-game currency, even though modifying it would do nothing. It has some strange boundaries.
- varispeed 5mo agoI myself got refusals often for legitimate data analysis work. I am starting to lean on buying powerful hardware little by little until I get suitable rig to run local models that make sense.
- ryan-a 5mo agoThis is huge for me too, I was working on something super benign the other day and GPT flagged it for Cyber risk, Deepseek just does the work, its fast and cheap. Its only missing image support IMO, once deepseek cracks image too its going to be hard for anthropic and openai to compete.
- teaearlgraycold 5mo agoClaude has refused to run nmap so I can locate my own computer on my own network! The guard rails are completely out of control.