7 ms·
I initially thought that maybe they were concerned about license violations, copyright infringement, etc. because you can't always be sure where the llms are so
by CSMastermind 3y ago
I initially thought that maybe they were concerned about license violations, copyright infringement, etc. because you can't always be sure where the llms are sourcing their information from.
But rather it seems like they're warning against information leaking out?
> The Google parent has advised employees not to enter its confidential materials into AI chatbots, the people said and the company confirmed, citing long-standing policy on safeguarding information.
Is there any example of this happening in the real world? I've never heard of one.
Even if you were to give me access to some google source code free and clear I'm not sure what I could do with it. It's often only useful with the build systems, other services, external documentation, and tribal knowledge contained within the company.
Sure if I had unfettered access to all of the Google source for a product for long enough I might be able to find a security vulnerability or something but random chunks of code pieced together from the inputs to some chat bot?
Just seems like an impractical attack vector.
- extr 3y agoI think it's way more banal than that. Think someone in HR asking it for format a list of salaries. Boring stuff.
- lkjdsklf 3y ago> Just seems like an impractical attack vector. It's not really about security in terms of attacking the site. It's purely about IP protection and potentially leaking user/employee data. Any data you give one of these chat bots can potentially be regurgitated by the bot later. We have the same kind of policies at my megacorp company. We can't even use things like chrome extensions that will let you paste in data to parse json or things like that.
- capitainenemo 3y agoParse json? Like... the built-in JSON formatting that Firefox has when hitting a JSON url or local json file? I'm rather surprised those are considered an attack vector, but I guess the concern is the extension might be running remotely?
- TeMPOraL 3y agoThey're especially good attack vector, because a lot of seemingly smart people may not realize that the little tool for pretty-printing JSON or changing User-Agent header could also be recording and sending out everything that goes through it.
- capitainenemo 3y agoAh. Well, when phrased that way, that's more a general problem with installing extensions, especially ones where the recent version hasn't yet gone through Mozilla review (or whatever the equivalent is for Chrome). Welp, fortunately, I've never had the slightest desire to install a json parser, the Firefox built-in formatting/parsing is quite adequate for my needs. Anything really complex filter-wise I just turn to jq.
- TeMPOraL 3y agoI'm speaking from experience. Many years ago I did install a couple random convenience extensions like this on my work computer - including a very nice JSON pretty printer, back before such things were included by default. It took me couple of months before something in my head clicked, and I realized with horror what attack vector I've been enabling at work. Fortunately, this was also before buying off extensions for malicious purposes became a thing, and as far as I could verify, I got lucky and all those extensions were legit and did not exfiltrare anything.
- capitainenemo 3y agoOuch. That must've been a long time ago, the formatter appears to have first added in Firefox 44 dev edition. But. Yeah. Unfortunately we do require extensions sometimes. For most users they are blocked.. I do stick to well known, Firefox recommended ones, where the extension has gone through their review. Certainly extensions can be bought out, but they'd still have to submit a new version for review with their malicious code in it, which hopefully would help. But yeah, that's not 100%. Major extensions have had issues in the past, thus the Stylish/Stylus fork.
- rootshellz 3y agoFor an example: https://openai.com/blog/march-20-chatgpt-outage https://openai.com/blog/march-20-chatgpt-outage
- grumbel 3y agoI think that depends on how big and smart the AIs grow going forward. All the stuff you type into a chatbot today could be used to train the chatbot from tomorrow. And that chatbot might than be able to make connections between all the different bits of information it has available that could reveal information that weren't directly obvious from just a single snippet posted into the chatbot. Furthermore that information might then become available to the public, so everybody that asks the right questions could snoop around Google's private data. Keep in mind, this isn't one employee asking one question to a chatbot, this will be tens of thousands of employees typing in multiple prompts each day.
- TeMPOraL 3y agoMore immediately though, if that data is stored for future use in training, or even recorded in logs, it can leak. Not just randomly - OpenAI in particular is sitting on a treasure trove of corporate secrets, medical data, PII, possibly even government secrets - all submitted by people unaware or dismissive of data security concerns. It's only a matter of time before someone deems it worth it to invest substantial resources into an attack, and if they succeed, lots of people and organizations will have a bad day.
- hunter2_ 3y agoThe same could be said for typical search engines, although I must admit the amount of data a typical user submits (and/or that the input form accepts) is lower for those. Not to mention URL en/decoders, base64 en/decoders, prettifiers, format converters, ... All of those should be included when training users not to divulge secrets, with chatbots just being one more added to the list. So I think this recent buzz must be more about the model getting trained, than about discrete leaks.
- dmonitor 3y agoA lot of those decoders can just run things on the client, or are unlikely to keep data in the server. Chatbots are explicit in that they use the data submitted to train their models. They pose a much bigger risk.
- luckydata 3y agoI have a few colleagues at Google that were literally copy and pasting long internal documents in ChatGPT to get them summarized.
- vipermark7 3y agoSounds like they're being overly reliant on the tool. I know work life at Google is probably fairly intense and stressful, but this seems pretty odd
- pessimizer 3y ago> Is there any example of this happening in the real world? I've never heard of one. 1) How would you know if your idea or code were ripped off, unless it was an exact bug-for-bug copy? 2) These things have only existed for 5 minutes. 3) Google knows exactly how these models and their interfaces work, and has probably speculatively designed many ways to mine inputs for interesting items. > Even if you were to give me access to some google source code free and clear I'm not sure what I could do with it. They're not afraid of you. I wouldn't know what to do with five pounds of gold or a bag full of stolen jewels. Doesn't mean they're not valuable to someone.
- x86x87 3y agoYes. It has happened and there are examples https://www.techradar.com/news/samsung-workers-leaked-company-secrets-by-using-chatgpt https://www.techradar.com/news/samsung-workers-leaked-compan...
- ted_dunning 3y agoI don't think that this is thought of as an attack vector as much as a leak or a crap vector. If your confidential and clever trade secret technique leaks into somebody else's code (or could have leaked) you may lose trade secret status. If you have a corporate policy to point to, you can claim that any leaks are due to unauthorized activity and thus the trade secret still stands. On the other hand, if you import somebody else's proprietary code, or GPL'ed code, or buggy code without realizing it, you face different risks. Being able to point to an established policy against using code from these tools helps with the defense against trolls who assert some vague kind of theft.
- deleted 3y ago[deleted]
- madeofpalk 3y ago"Don't enter internal secret data into third party websites" seems like a pretty standard, run of the mill recommendation.
- fragmede 3y agoBut "don't enter internal secrets into a public website the company you work for runs" is a different one. How locked down is your Jira access?
- sp332 3y agoWe know it's not private because they're considering a private version for 10x the price. https://arstechnica.com/information-technology/2023/05/report-microsoft-plans-privacy-first-chatgpt-for-businesses-with-secrets-to-keep/ https://arstechnica.com/information-technology/2023/05/repor...
- stOneskull 3y agowhy not just let staff use that?
- sawyna 3y agoI suppose it's an over-arching policy created that makes it easier to mandate. If not, Google would have to say - you can enter some stuff into chatbots, but not sensitive stuff. Then it would have to classify sensitive and non-sensitive, keep track of this, have this up to date etc and etc - you get the idea. By sensitive stuff, I'd say entering youtube's content moderation algorithm package for instance.
- kortilla 3y agoThere are trade secrets in code. Most people vastly overestimate the value of their secrets, but they definitely exist. Pasting your code into OpenAI is definitely giving people outside of the company access who have not entered into agreements to leak it (intentionally or not).
- bfuller 3y agoNot only people, couldn't the LLM itself distribute secrets by being prompted to write similar code or something? I don't know if they end up training the model on the data it has received/generated.
- SkyPuncher 3y agoGenerally, this is more a concern with data than source code. This type of stuff happens all the time and is a big problem for companies. Hence a category of Data Loss Prevention. Most of the concern is going to be on customer data, business relationship data, and confidential internal data. Companies don't want you running your quarterly business review through an unapproved tool. Even though, it's unlikely to happen. It does happen.
- deleted 3y ago[deleted]
- bertday 3y agoThe model can return the internal details of a potential future product. For example, the next Pixel phone.
- lowbloodsugar 3y agoSomeone has an update on a mission critical project that competes with Microsoft, wants to get the wording just right, and sends it to ChatGPT to tart it up. Someone at Microsoft is archiving any and all input from google IP space and making it available for inspection. Doesn't seem that far fetched. The threat isn't AI. The threat is sending intellectual property to competitors.
- mywittyname 3y ago> Is there any example of this happening in the real world? I've never heard of one. This has happened to Samsung. Some of their codebase and meeting notes were leaked by CGPT. The company I work for has the exact same policy for the exact same reasons. It can be difficult to tell who has been effected by this because you might need to know what to put into the prompt in order to have the leaked data leveraged. So this can go undetected by the victim.
- hgsgm 3y agoIt's just corpos being self important and avoiding anything that could become a lawsuit hook. A few megabytes of source code is nothing compared to letting Levandowski walk out with all the self driving car tech.
- verdverm 3y agoGoogle is less worried about internal code leaks, they train bard on part of their internal code. The point is likely more... biz people, don't put in confidential financial & usage numbers when using the writing assistant, don't use code produced by them because it is not trustworthy, and might contain GPL stuff which would infect our code.
- grepfru_it 3y agoI was an early Bard user. Bard allowed me to browse internal google code repositories. I was able to view all of the YouTube code. With the recently announced battle against inviticus, it is very relevant. There is a reason private repos are private Bonus: the code was pretty tough to understand alone, but with public api documentation it gave a lot of insight to how google works. Sure I wasn’t going to find a buffer overflow or a way into their network at quick glance, but enough time and I can siphon ideas, data, or full chunks of code to do what I want with. I reported the vulnerability through a colleague and they have since patched it.
- ipsum2 3y agoThat sounds highly implausible. Bard would not be trained on internal Google code, nor would have access to it. Most likely it was hallucinating, based off of public APIs its seen. Did you ever get confirmation that that code actually existed?
- grepfru_it 3y agoRemember when it was implausible for a website to drop all of its database tables based on the name you input into a form? Get your popcorn ready, all the silly vulnerabilities of the 90s are coming back again
- ShamelessC 3y agoBullshit.