17 ms·
Exposed DeepSeek database leaking sensitive information, including chat history
- maitola 2y agoHow do we know for sure that DeepSeek is not actually trained on Nvidia chips? Did someone outside of China replicated the training from scratch (Spending $6M)?
- haeffin 2y agoThey themselves said it was trained on NVIDIA chips, so I’m not sure where you got that it wasn’t. It was trained on the less capable versions sold for the Chinese market.
- deleted 2y ago[deleted]
- maitola 2y agoI see, thank you for pointing that out. Then I’d rephrase, how do we know for sure that it wasn’t trained on the most advanced Nvidia chips? Did anyone outside of China replicated the training?
- NathanKP 2y agoAnd that's why you run models locally. Or if you want a remote chat model, use something stateless like AWS Bedrock custom model import to avoid having stored chats on the server.
- dotancohen 2y agoNot many non-gamers have hardware capable of running such a model locally - never mind the skills. For most people, bash is not a tool for interacting with the computer, it is how they express their frustration with the computer (sometimes leaving damaged keyboards).
- loloquwowndueo 2y agoWow all the gamers with mad LLM skillz.
- 0x457 2y agoPretty sure gamers are mentioned because those are the usual demo that has GPUs with enough memory outside of people in the ML industry.
- loloquwowndueo 2y agoSo you’re in the demo scene as well? Yay
- 0x457 2y agoYou were not able to use context clues to figure that "demo" in this case is short for "demographics"? Sad
- razster 2y agoI have DeepSeek-R1 1.5b running on a Raspberry Pi 5. I have DS-R1 14b Q6 running on my old AM4 Ryzen with a AMD GPU, without issues. My primary workstation is running 32B Q8 and without issues. And it's simple!
- smallerize 2y agoThat's not the DeepSeek R1 model that they're offering via the API on these servers. That's a Qwen model that's been fine-tuned on output from the big R1 model.
- xinayder 2y agoSource?
- 2y ago
- tonygiorgio 2y agoYou could also use models that run on nvidia’s trusted execution environment.
- janalsncm 2y agoNvidia naming it “trusted” doesn’t mean I trust it.
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- rvz 2y ago> More critically, the exposure allowed for full database control and potential privilege escalation within the DeepSeek environment, without any authentication or defense mechanism to the outside world. Not only that, this was a "production-grade" database with millions of users using it and the app was #1 on the app store and ALL text sent there in the prompts was logged in plain-text? Unbelievable.
- byearthithatius 2y agoI agree this is really bad but far from unbelievable. I am only 23 and already my SSN and even my freaking DNA have both been leaked by major publicly traded companies.
- sho_hn 2y agoPlus Volkswagen and Subaru in the last few weeks ...
- reaperducer 2y agoPlus Volkswagen and Subaru in the last few weeks Both Volkswagen and Subaru have leaked his DNA in the last few weeks? Dude gets around.
- dotancohen 2y agoVW is the people's wagon - where do you think those people come from?
- nicce 2y agoOn top of SSN and DNA, also the location of the DNA has been leaked.
- byearthithatius 2y agoheeheehee;) i REALLY like cars
- 2y ago
- hdlothia 2y agoThis kinda does support the 'DeepSeek is the side project of a bunch of quants' angle. Seems like the kind of mistake you would make if you are not used to deploying external client facing applications.
- sailingparrot 2y ago> This kinda does support the 'DeepSeek is the side project of a bunch of quants' angle Can we stop with this nonsense ? The list of author of the paper is public, you can just go look it up. There are ~130 people on the ML team, they have regular ML background just like you would find at any other large ML labs. Their infra cost multiple millions of dollar per month to run, and the salary of such a big team is somewhere in the $20-50M per year (not very au fait of the market rate in china hence the spread). This is not a sideproject. Edit: Apparently my comment is confusing some people. Am not arguing that ML people are good at security. Just that DS is not the side project of a bunch of quant bros.
- weird-eye-issue 2y agoNone of that has anything to do with "deploying external client facing applications"
- Dylan16807 2y agoYou're right. It has nothing to do with the second sentence of the two sentence post it replies to.
- islewis 2y agoA bunch of ML researchers who were initially hired to do quant work published their first ever user facing project. So maybe not a side project, but if you have ever worked with ML researchers before, lack of engineering/security chops shouldn't be that surprising to you.
- spoaceman7777 2y ago
- nico 2y agoSo much effort in trying to tarnish DeepSeek the last 24hrs
- deleted 2y ago[deleted]
- mandmandam 2y agoYep. Kinda like how your comment was grey within 1 minute, despite stating an objective truth. Sure, this is to be expected given the billions and billions of dollars at stake but like - that money is gone lol. DeepSeek isn't going back in the bottle, nor is open source AI in general.
- IncreasePosts 2y agoThe comment isn't wrong, but the implication is at deep seek is somehow special and is getting undue attention from hackers. Any app that skyrockets from nowhere to number one in the app store overnight will have the attention of probably hundreds of thousands of hackers.
- mandmandam 2y ago> the implication is at deep seek is somehow special and is getting undue attention from hackers ... It is special. There was over a trillion dollars wiped from tech stocks lol, in a massive win for consumers and the planet. You can't say this is like Flappy Bird or something.
- IncreasePosts 2y agoThe stock market does all sorts of silly things. If the stocks recover in 2 weeks to where they were, will that be deep seek erasing $1T and tech re-earning $1T? Or deep seek doing nothing?
- codr7 2y ago
- tomlockwood 2y agoThis doesn't look like a responsible disclosure, at all. ed: I was wrong!
- CamelCaseName 2y agoWho's going to go after them? Heck, they may get an award for this.
- krick 2y agoUh, I don't know, but cannot DeepSeek do that, for starters? Being located in a different country than the service you are attacking doesn't really make you immune to being sued.
- megous 2y agoIf random security researcher does this kind of disclosure, fine. But if serious company that seems to offer services to seemingly plenty of serious customers acts this way, I'd not want to be their customer, if they seem to have such a cavalier attitude, disclosing stuff without even a sniff of "we notified the company about the breach".
- immibis 2y agoIt was fixed. Disclosing it after it's fixed is responsible.
- varenc 2y agoFrom the article: > The Wiz Research team immediately and responsibly disclosed the issue to DeepSeek, which promptly secured the exposure. Assuming everything mentioned in the article was fixed before publication, I don’t see an issue with it.
- tomlockwood 2y agoYeah my bad I missed that and edited OP.
- mrbungie 2y ago[edit: Nevermind, see below] The direct disclosure of urls and ports is insane. Wonder if they would be as irresponsible if it was MSFT, OpenAI, Anthropic, etc. PS: Not defending DeepSeek for bad practices, but still. Nothing irresponsible here. PS2: It is marked as resolved, I went directly to the vulns due to the title of the post.
- bberenberg 2y agoIt’s been disclosed and resolved. What’s the concern here?
- nyclounge 2y agoWhy is ClickHouse exposing unauthenticated database access at port 9000 to the public? Is this the default behavior or did DeepSeek open it up for dev purposes?
- ceejayoz 2y agoThat used to be the default setup for Redis, too. Might still be. You aren’t supposed to have it on a public subnet.
- SahAssar 2y ago> You aren’t supposed to have it on a public subnet. That's an incredibly bad assumption. To have defaults assume that you are on a protected network (what does that even mean? like what permissions are assumed just because you are on the same network? admin?) is just bad practice.
- j45 2y agoA data point on self-hosting being preferable, or using an alternate gpu cloud host who can run the model privately/semi-privately for you.
- danielodievich 2y agoopen exposed clickhouse is this decade's open exposed elasticsearch so common in the past
- bearjaws 2y agoWhich was originally the open exposed mongo server, then mysql/phpmyadmin, then exposed ftp, and then exposed telnet.
- hmmm-i-wonder 2y agoWe move on and upwards, but never really stop making the same mistakes do we.
- astrea 2y agoShows how old I am. Thought we were still in the "exposed ElasticSearch" era.
- kdmtctl 2y agoI was sure this was Elastic, you are not alone.
- ebfe1 2y agoAFAIK, Opensource Elasticsearch does not offer any form of authentication upon installation for many years but ClickHouse does and in fact I'm often surprised at how many authentication mechanisms were introduced over the years and can be easily configured: - Password authentication (bcrypt, sha256 hashes) - Certificate authentication (Fantastic for server to server communication) - SSH key authentication (Personally, this is my favourite - every database should have this authentication mechanism to make it easy for Dev to work with) Not very popular but LDAP and Http Authentication Server are also great options. I also wonder how DeepSeek engineers deployed their ClickHouse instance. When I deployed using yum/apt install, the installation step literally ask you to input a default password. And if you were to set it up manually with ClickHouse binary, the out-of-the-box config seal the instance from external network access and the default user is only exposed to localhost as explained by Alex here - https://news.ycombinator.com/item?id=42871371#42873446 https://news.ycombinator.com/item?id=42871371#42873446.
- b3ing 2y agoIt seems fair since all the other AI's scraped copyrighted information, images, video online and from pirated books, etc. without ever asking anyone first.
- samedev 2y agoMan! I used deepseek.com luckily I didn't use the same password as I use. :) Time to use ollama!
- caust1c 2y agoInteresting to note: - Dev infra, observability database (open telemetry spans) - Logs of course contain chat data, because that's what happens with logging inevitably The startling rocket building prompt screenshot that was shared is meant to be shocking of course, but most probably was training data to prevent deepseek from completing such prompts, evidenced by the `"finish_reason":"stop"` included in the span attributes. Still pretty bad obviously and could have easily led to further compromise but I'm guessing Wiz wanted to ride the current media wave with this post instead of seeing how far they could take it. Glad to see it was disclosed and patched quickly.
- pedrovhb 2y ago> but most probably was training data to prevent deepseek from completing such prompts, evidenced by the `"finish_reason":"stop"` included in the span attributes As I understand, the finish reason being “stop” in API responses usually means the AI ended the output normally. In any case, I don't see how training data could end up in production logs, nor why they'd want to prevent such data (a prompt you'd expect to see a normal user to write) from being responded to. > [...] I'm guessing Wiz wanted to ride the current media wave with this post instead of seeing how far they could take it. Security researchers are often asked to not pursue findings further than confirming their existence. It can be unhelpful or mess things up accidentally. Since these researchers probably weren't invited to deeply test their systems, I think it's the polite way to go about it. This mistake was totally amateur hour by DeepSeek, though. I'm not too into security stuff but if I were looking for something, the first thing I'd think to do is nmap the servers and see what's up with any interesting open ports. Wouldn't be surprised at all if others had found this too.
- caust1c 2y agoSeems that you're right! Also, not that I doubted they were using OpenAI, but searching for `"finish_reason"` on the web all point to openai docs. Personally, I wouldn't say it's a very common attribute to see in logs generally. https://platform.openai.com/docs/api-reference/introduction https://platform.openai.com/docs/api-reference/introduction Right there in the docs: > Now that you've generated your first chat completion, let's break down the response object. We can see the finish_reason is stop which means the API returned the full chat completion generated by the model without running into any limits. Regarding how training data ends up in logs, it's not that far fetched to create a trace span to see how long prompts + replies take, and as such it makes sense to record attributes like the finish_reason for observability purposes. However the message being incuded itself is just amateur, but common nonetheless.
- jvansc 2y agoThis is probably an incredibly stupid, off-topic question, but why are their database schemas and logs in English? Like, when a DeepSeek dev uses these systems as intended, would they also be seeing the columns, keys, etc. in English? Is there usually a translation step involved? Or do devs around the world just have to bite the bullet and learn enough English to be able to use the majority of tools? I'm realizing now that I'm very ignorant when it comes to non English-based software engineering.
- ceejayoz 2y agoThe languages and frameworks and documentation are often in English. The code has a good chance of also being in English as a result. See also: aviation.
- colordrops 2y agoI worked at a Chinese company for a while and they used Chinese in meetings but English in the code base.
- deleted 2y ago[deleted]
- rcruzeiro 2y agoSomeone who worked on a non-English environment years ago here: sometimes you do use the local language in some contexts, but, more often than not, you end up using English for the majority of stuff since it's a bit off-putting to mix another language with the English of programming languages and APIs.
- sghiassy 2y agoDumb question, but it would then seem that you have to know English to program??
- creakingstairs 2y agoIt’s harder to learn for sure. Majority of the resources are in English and it’s harder to internalise the keywords. But it’s definitely possible to program without knowing English.
- dotcoma 2y agoIt’s a feature, not a bug !
- lysace 2y ago[flagged]
- nostradumbasp 2y agoBecause you can download the models yourself and run them on your own hardware bypassing all of those concerns. Also, I would be far more worried what my government might do with my chat logs then a foreign one.
- lysace 2y agoThis post is about the service, not the model(s).
- fulladder 2y agoTheir models are open weights, and they are supported in ollama. You can run locally if you have sufficient hardware.
- lysace 2y agoWell, thanks for not calling it open source! I did run it locally. It behaved as you would expect when asked about things the CCP cares about. https://hongkongfp.com/wp-content/uploads/2021/11/brave_udRsuvasM7.jpg https://hongkongfp.com/wp-content/uploads/2021/11/brave_udRs...
- ryouna 2y agoVery easy to circumvent if you run the raw model locally, no different from how western models are lobotomized for liberalist/nationalist reasons. In fact these models are less lobotomized than western open source models.
- billyjmc 2y agoWell, it is technically open source, open weights, and closed training set, right? (My recollection is that the training code is MIT licensed.)
- bryan_w 2y agoThis is totally expected when you use AI to build your infrastructure.
- ripped_britches 2y agoI was going to say the opposite, ironic because an LLM would have told them not to do that if they were working closely with one.
- bryan_w 2y agoWith the various ways people setup their dev environment + docker, I imagine there are probably a lot of guides that show you how to set it up insecurely (because it assumes you're connecting from your local network) with a small asterisk at the end saying not to set it up in production like that. Very easy for an LLM to misunderstand.
- mr90210 2y agoPoorly secured or not it still managed to hit your favourite stock. The execs at NVIDIA still haven’t recovered from the bloodbath.
- Etherlord87 2y agoThis argument seems fallacious: the stock was "hit" before security issues emerged. It's hard to say how this recent news will affect the stock, directly and indirectly through eventual damage that follows. Imagine that the security issues were 1000 times worse: you could still write the same comment, but the reputation of DeepSeek and by extension all Chinese AI and software would be so badly hurt that long term any Chinese success would be downplayed, having lesser effect on stock market. If Nvidia stock would recover or not is more nuanced, because market is speculative, and if a bubble is burst, even if what pierced it turns out to be fake/irrelevant, the bubble is no longer there (a new one may need a lot of time and effort to grow).
- Havoc 2y agoUgh. I know I’ve got at least some keys in those logs. Thankfully nothing too intense
- danparsonson 2y agoHopefully this is a lesson not to trust your sensitive private data with a public service?
- sd9 2y agoI've been redacting my keys before sending config to chatgpt, it's a pain but I guess this shows it's worth the effort.
- Havoc 2y agoYeah I avoid it too but I know I missed some during rapid copy pasting.
- ripped_britches 2y agoIronic - I bet if you ask deepseek r1 how to set up clickhouse it would tell you the right way to do it.
- hi_hi 2y agoI don't get the discussions around side project and they're ML engineers, not security experts. Why are you excusing a company for a serious security leak. If you're releasing a major project into the wild, expect serious attention and have the money, you get third parties involved to test for these things before you launch. Now can we get back to discussing the real conspiracy theories. This is clearly a disinformation piece by BigAI to add FUD around the Chinese challenger :-)
- throwaway314155 2y ago> I don't get the discussions around side project and they're ML engineers, not security experts. Why are you excusing a company for a serious security leak. No one is here as far as I can tell. But if you've ever been a software engineer who is required to work with someone purely from an ML lab and/or academia, you'll quickly discover that "principled software engineering" just isn't really something they consider an important facet of software. This is partly due to culture in academia, general inexperience (in the software industry) and deeply complicated/mathematical code really only needing to be read by other researchers who already "get it", to a degree. Not an excuse but rather an explanation for _why_ such an otherwise impressive team might make a mistake like that.
- hi_hi 2y agoYeah, you're right, I was conflating the excusing bit. I haven't worked with serious ML engineers, but having worked in large webdev there's usually a team involved in these projects, including senior none devs who would ensure the correct checks and balances are in place before go live. Does this not happen in ML projects? (of course there are always exceptions and unknowns that will slip through, I don't know if that was the case here, or something else)
- throwaway314155 2y ago> Yeah, you're right, I was conflating the excusing bit. No worries. :) > Does this not happen in ML projects? Consistently? No. At the level of e.g. OpenAI/Anthropic? It is mandatory. These are not just research labs, they're product (ChatGPT, Claude) companies. These American companies have done a reasonable job at hiring for all sorts of skillsets to keep things well rounded. Perhaps DeepSeek hasn't learned this lesson yet... Or, well - it could be far more complicated than that. Speculating is only so useful with so little information.
- seeknotfind 2y agoWhere's the download link?
- mmaunder 2y agoDoes DeepSeek have a bug bounty program I'm not aware of with a clearly defined scope? It appears that Wiz took it upon themselves to probe and access DeepSeek's systems without permission and then write about it. If you do this and the company you're conducting your "research" on hasn't given you permission in some form, you can get yourself in a lot of hot water under the CFAA in the USA and other laws around the world. Please don't follow this example. Sign up for a bug bounty program or work directly with a company to get permission before you probe and access their systems, and don't exceed the access granted.
- tevon 2y agoThey left open a publicly exposed database... I'm sure they informed the company about this before publishing their post. Why are you blaming Wiz for this?
- pinoy420 2y agoYes but they’re chinese so it’s okay /s They are getting DoS’d by us gov too so they were only trying to help /s
- SomeRainIsGood 2y agolol
- SomeRainIsGood 2y agowritten like someone who has never litigated even a traffic light
- ziddoap 2y agoThey're publicly accessible URLs. DeepSeek & users that had data exposed here should be thanking Wiz.
- soulofmischief 2y agoYour posturing is unwarranted. Literally in the first paragraph: > The Wiz Research team immediately and responsibly disclosed the issue to DeepSeek, which promptly secured the exposure
- suraci 2y agothat's why i never use my strong passwords in many chinese websites(in fact, i tend not to use passwords in any website) i suggest you guys don't do that also this industry in china is so young, many devs and orgs don't understand what will happened if they shutdown the firewall or expose their database on the internet without a password they just, can't think of it, need someone to remind them
- gitaarik 2y agoI didn't understand your comment first, because I use a password manager which generates a unique and complicated password for each website I setup an account for. So I never reuse any password. So if one one those sites gets hacked and my password is potentially exposed, it doesn't matter, because I only use that password there. I would recommend that. Bitwarden is a pretty good open-source password manager. You can install it as a plugin in your browser, so it can fill out your password for you so you don't have to manually copy and paste.
- mmaunder 2y agoThe amount of vitriol in these comments is the really surprising data. I've seen the same on Twitter. I can only put it down to the financial pain DeepSeek inflicted on many US retail investors by wiping almost $700 billion off NVidia's stock price. I think a lot of folks didn't see it coming and it hurt them right where it matters most: In the wallet. The anger out there is very real.
- gerdesj 2y ago"The amount of vitriol in these comments is the really surprising data" No it isn't (well it probably is too). This is the rather naff nation state bollocks in play. You have either or both of "some bigger boys found a more efficient way of doing something I thought I was good at" and "I've wet myself".
- gerdesj 2y agosigh
- bobxmax 2y agoIt's also deeply damaging to the western ego, especially one rooted in American exceptionalism. But also one those of us actually working on foundational AI saw coming a mile away when most of the top research of late has been happening in Chinese labs, not American or European ones. Can't wait to see what this boneheaded President's tarrif on TSMC does to this situation.
- hsuduebc2 2y agoWell to be honest most of this on start came from US so the general surprise is understendable. But of course it would be foolish and arrogant assume that for whole progress forever. I don't understand the rage. This is good for everyone. Competition is what drives innovation and they even open sourced it! If you want to outdo them, learn from them. Don't just try to cry louder, it's embarrassing for everyone.
- awesomeMilou 2y ago
- lexandstuff 2y agoAnother example of DeekSeek copying straight from OpenAI's playbook [1] [2] [1] https://www.reuters.com/technology/cybersecurity/openais-internal-ai-details-stolen-2023-breach-nyt-reports-2024-07-05/ https://www.reuters.com/technology/cybersecurity/openais-int... [2] https://openai.com/index/march-20-chatgpt-outage/ https://openai.com/index/march-20-chatgpt-outage/
- nialv7 2y agoI wonder if this is the "cyberattack" DeepSeek was talking about?
- juliuskiesian 2y agoYou are wondering wrong. This is a security hole and data leak. A large scale DDos is being directed against deepseek. US big tech wants to quench the competition.
- anhldbk 2y agoGood finding. I don't see its timeline usually discussed in other Ethical hacking and responsible disclosures.
- semking 2y agoCan you imagine executing arbitrary SQL queries via your web browser? :D Complete database control and potential privilege escalation within the DeepSeek environment without ANY authentication...
- galnagli 2y agoThank you everyone, this was responsibly disclosed to DeepSeek and published after the issue was remediated, we got acknowledgment from their team today on our contribution.
- leftcenterright 2y agowere these "dev" domains holding real production data? the blog post does not clear it for me.
- sylware 2y agoThe second Big Tech was threatened by significant competition (DeepSeek), this competition is "stealing"(lol), and is under heavy hacking attacks (main online inference portal). There you have, the real face of Big Tech. Extinguishing the competition by locking a service behind a portal provided for free, then starting to milk the users, is not enough for them... they will also fight dirty, really dirty.
- SebFender 2y agoNever forget honeypots.