6 ms·
Self-hosting keeps your private data out of AI models
- deleted 2y ago[deleted]
- itronitron 2y ago>> “To develop AI/ML models, our systems analyze Customer Data (e.g. messages, content, and files) submitted to Slack.” — Slack’s privacy principles,
- simonw 2y agoAs far as I can tell that's the copy they've had in place since 2016 - a year before even the original Transformers paper that kicked off today's LLMs - to cover their own tiny old-school ML models for things like channel recommendations.
- jmclnx 2y ago2FA pushed me out of github, but M/S Copilot in github created the road out. Now seems people's chats and posts are being used to train AI. I wonder how long before Cell Phone Providers start using Text Messages to train AI (or sell to AI people).
- giancarlostoro 2y agoGiven how awful some text messages I get from relatives read, I really hope not. The worst types of typos. The only thing I want them to train properly is voice to text. It cannot for the life of itself ever get anything right. I have to scream at Siri a dozen times.
- eesmith 2y ago2FA also pushed me out, and Copilot locked the door ... to be fair, unless a client comes with bucket full of money with a key on top.
- bluish29 2y ago> 2FA pushed me out of github I don't understand why would something like that be a reason to avoid github (or any service) in itself? In this era, 2FA is important security measurement and it is not like they enforce only SMS as the only 2FA form.
- marcosdumay 2y agoModern 2FA is "something you know" and "something you subscribe for".
- cuu508 2y agoWhat do you mean? GitHub supports TOTP, Webauthn on free accounts.
- cuu508 2y agoTo answer myself, from reading the other comments, it sounds like GitHub started to require 2FA at some point, and some people refused to set it up. The problem was not some inadequacy of GitHub's 2FA implementation, but the fact it is mandatory. (I had 2FA set up a long time ago, so didn't notice the policy change)
- eesmith 2y agoTo start with, I don't have two devices, only a laptop. (And a backup laptop and desktop at home, but I typically don't work there.) Correct, I have no smart phone. Second, it's not an important security measurement for me. I didn't go into FOSS to be part of someone's supply chain. I do it as a way to share my knowledge of how one might solve a problem. If you want use my code, then inspect it to make sure it does what you want, or pay me for commercial support. Neither require 2FA. Might someone take over the account? Sure, I suppose. But I'm not into "community building" or GitHub's gamification, and my primary repos are all local, so if that happens and GitHub's support didn't help, I could start a new account. Again, don't depend on me for your supply chain without a commercial support agreement. When Microsoft switched GitHub to require 2FA I concluded it was because they wanted to assure their corporate and government clients that it was "safe" for them. Those profits subsidize Microsoft's free hosting plans, so my presence there was helping contribute to Microsoft's already excessive market power. Third, the change was driven from on high, with no chance for me to decide what was appropriate for my projects. I concluded Microsoft was so powerful they could make such paternalistic changes because they knew network effect was on their side that they could have little concern about the small number of people leaving or getting upset. Fourth, my FOSS projects on GitHub were labors of love that were a net negative on my income. I was not going to spend any money on new hardware or waste my time figuring out how to get things working under a new system when I was already hosting most of my work on Sourcehut, which is much more aligned to my ethical and moral views. I still don't know how many security keys I'm supposed to have (how often should I expect to lose one? should I store the backups off-site at a friend's place?), or how often am I supposed to test they work? And then I hear about issues about lock-in and how attestation requirements might prevent FOSS solutions ad prevent people from backing up one's own security keys, and issues with resident vs. non-resident keys, and being able to register multiple keys. It's all learnable, but I simply don't care enough. And I don't see why I should care about all this when the paying customers of my software have all been fine with only a tar.gz, license agreement, and support contract.
- alright2565 2y ago> 2FA pushed me out of github I've seen several folks have this opinion... and it just doesn't make any sense? SMS for 2FA is problematic for many reasons, but Github offers TOTP & U2F/passkey 2FA, which have zero privacy implications.
- ProllyInfamous 2y ago2FA is why I can no longer utilize online banking (which is frustrating).
- leobg 2y agoLooks like content marketing to me. Both in purpose and motivation.
- oidar 2y agoSelf-hosting is just responsible computing now. The big companies are just too big to care about small businesses, and will use your data in any way that they please - take it or leave it. And it's cheaper to boot. A synology NAS or a raspberry pi 3 could cover 90% of what most internet services offer the average consumer/small business right now.
- throwanem 2y agoYou're leaving out "and having someone on staff or contract who can administer it for you".
- orev 2y agoBecause with cloud providers you don’t need anyone on staff to understand the product, keep up with the constant changes and predatory features they slip in, manage the cost overruns, or help to migrate away from their proprietary data formats once you’ve had enough of their “whaddaya gonna do about it?” treatment of customers?
- mikestew 2y agoTrue or not, it doesn't invalidate GP's statement. You don't just throw a NAS in Mom's garage and call it a day.
- Cheer2171 2y agoYes, it isn't bare metal vs cloud, it is being your own sysadmin vs relying on someone to sysadmin for you
- dvfjsdhgfv 2y agoI don't get your point. It doesn't matter if you host your stuff on AWS, Hetzner, or locally on a Pi, you need someone to manage that. If you are this someone, great, if not, you need to pay someone else. And contrary to what private cloud providers advertise,[0] the amount of work is similar in each case. [Or used to, around 2006 or so - I'm not sure if they still claim managing public cloud resources is means less work.]
- matchagaucho 2y agoTo conflate Slack's T&C faux pas with "self host your own LLM" seems like a stretch. IT Sec and Compliance must read the T&Cs and make better vendor selection choices.
- sneak 2y agoThe recommendation is “self host your own collaboration tools”, not “self host your own LLM”.
- deleted 2y ago[deleted]
- sneak 2y agoAnother offender: codegpt.io ToS grants them an irrevocable perpetual sublicenseable license to all code they see from you. It’s insane what rights companies claim to your data.
- udev4096 2y agoThis will continue to occur and may come as a "shock" to some companies that will ignorantly persist in using proprietary services unless a significant change in data collection from the service itself occurs, which should not be the primary motivation to switch to a self-hosted version in the first place
- codegeek 2y ago"We don’t train LLMs on Zulip Cloud customer data, and we have no plans to do so. Should we decide that training our own LLMs is necessary for Zulip to succeed, we promise to do so in a responsible manner" :). What a clever way to say that even though we don't do it today, we cannot guarantee that we will never do it on our cloud service. At least they are honest I guess.
- simonw 2y agoI'm amused to note that those are effectively the exact same policies that Slack have in place!
- tabbott 2y agoZulip project leader here. I think you're missing a big part of the point of the post: Which is that if you self-hosting, nobody can train models on your data, if you're going to use a cloud service, you should use one where you can move data to self-hosting, and where you can trust the vendor. We do basically guarantee that you won't have your organization's data included in AI training in Zulip Cloud without consent. But yes, we're not ruling out the possibility of some sort of opt-in feature that might be useful in Zulip Cloud. I'm humble enough to not pretend I know what will be possible/expected in terms of AI technology in 5-10 years, but one could easily imagine some sort of tool trained on web-public channel data in open Zulip communities being a thing that could make sense if done with appropriate consent. If such a thing were desired, I don't think it most of the concerns related to the slack controversy would apply.
- codegeek 2y agoI respect the self hosting part of course. All I am saying is that the premise of the post is that you guys take privacy seriously but you still left that door opened that someday you may use LLMs.
- itronitron 2y ago>> we're not ruling out the possibility of some sort of opt-in feature that might be useful in Zulip Cloud. Zulip's weasel wording indicates they are nerfing a great (maybe their best) opportunity to stand out from the herd. How about this for a mind-blowing concept (/s) ... If a web-public channel wants to add some sort of useful feature based on a technology trained on their data then let the owners/administrators of that channel flip that switch on. Zulip should have no involvement in that decision, period.
- simonw 2y agoSince this seems to be written partly in response to (and honestly, to take advantage of) the recent Slack AI training panic, I took a look to see how Slack have updated their materials in response to that panic. These documents are new in the last few days: https://slack.com/blog/news/how-we-built-slack-ai-to-be-secure-and-private https://slack.com/blog/news/how-we-built-slack-ai-to-be-secu... https://slack.com/intl/en-gb/blog/news/how-slack-protects-your-data-when-using-machine-learning-and-ai https://slack.com/intl/en-gb/blog/news/how-slack-protects-yo... I think these updates are really good - Slack's previous messaging around this (especially the way they unclearly conflated older machine learning models with new policies for generative AI) was confusing and it wasn't surprising it caused a widespread panic. It's now very clear what Slack were trying to communicate: they have older ML models for features like channel recommendations which work how you would expect such models to work. They have a separate "Slack AI" addon uou can buy that adds RAG features powered by a foundation model that is never further trained on user data. I expect nobody will care. Once someone has decided that a company might "train AI" on private data you've already lost that person's trust. It's not clear to me if any company has figured out how they can overcome one of these AI training panics at this point. I wrote a bit about this back in December when it happened to Dropbox - there is an AI trust crisis at the moment: https://simonwillison.net/2023/Dec/14/ai-trust-crisis/ https://simonwillison.net/2023/Dec/14/ai-trust-crisis/
- Root_Denied 2y ago>I expect nobody will care. Once someone has decided that a company might "train AI" on private data you've already lost that person's trust. It's not clear to me if any company has figured out how they can overcome one of these AI training panics at this point. I think it goes beyond a single company, or rather a single incidence of this panic. You're looking at each time this happens as an independent coin flip instead of a series of dominoes that trigger a reaction in multiple directions. What I mean by that is there's a counterculture sentiment building based off the idea that people have seen this same pattern enough times at this point that they're distrustful of large scale systems by default. It's happening with government institutions, politics, economics, and individual industries like gaming and streaming. To that end the "panic" is not just a reaction to Slack's (perceived) actions, but an expectation that Slack will be yet another domino in that line of companies that have done the same. It's also difficult to prove a negative (that Slack isn't using private data for training purposes even if they say they're not) so the messaging is up against a very solid wall. The result here is that public announcements and messaging related to data are under heavy scrutiny, and the media is incentivized to try and make their reporting go viral (ironically for the ad revenue) at the expense of actual journalistic reporting. I'm not sure what the solution to this problem is, or if there even is one, but promoting self hosting seems like an indicator that the default assumption is that data collected will be abused in some way. Honestly based on the last couple of years it's not an unreasonable assumption either.
- WhackyIdeas 2y agoConsidering Microsoft are bringing in a ‘feature’ to record your desktop, I wouldn’t be surprised that an additional ‘feature update’ further down the line will simply take all those chats with your self-hosted AI models to train AI models. So in my opinion, it just doesn’t matter if you are using self-hosted AI, the weakest link in your chain for keeping your data private is the very OS’s that you’ll be interacting with said self-hosted AI. And with all the manufactured fear mongering going on around AI, that data will -already- be deliciously irresistible for prism-participating, lovable, trustable companies like Microsoft. Sorry to burst some pretty bubbles for the lovely naive people.
- rpcope1 2y agoIt's never felt better to just be running Debian on everything.
- WhackyIdeas 2y agoOpenBSD and Llamafile go hand in hand. https://github.com/Mozilla-Ocho/llamafile https://github.com/Mozilla-Ocho/llamafile
- deleted 2y ago[deleted]
- airpoint 2y agoUpgrading the UI from year 2000 and improving the UX would incite me to consider Zulip as a potential alternative. Bashing a competitor in a blog post does not.