7 ms·
It does not really matter anymore how bulk usage data collection is called or whether it is "privacy-preserving". Looking at the current developments in AI, I
by dschuetz 3y ago
It does not really matter anymore how bulk usage data collection is called or whether it is "privacy-preserving".
Looking at the current developments in AI, I am concerned that AI models can easily de-anonymize and guess end point users when being fed with "telemetry data" of hundreds of thousands clients.
I hear a lot and read a lot of software and hardware vendors saying that "telemetry" is supposed to somehow magically improve the user experience in the long run, but in actuality software tends to get worse, unstable and less useful.
So, I would like to know how exactly any telemetry data from Fedora Linux clients is going to help them, or how is it going to improve anything.
- musicale 3y ago> So, I would like to know how exactly any telemetry data from Fedora Linux clients is going to help them, or how is it going to improve anything. It won't improve anything for users. It might improve something for IBM.
- marginalia_nu 3y agoYou really don't need AI to do this. Collect enough data points and you can fingerprint basically anyone using very old fashioned techniques. AI doesn't really bring anything new to the table.
- dschuetz 3y agoAI would need far less data points to achieve an acceptable result, and maybe even be able of fingerpointing, not just fingerprinting. Just wait and see how the data broker bros start selling AI assisted data mining for ads. Big tech is doing that already. It's mind boggling how everyone seems to be just OK with that.
- didntcheck 3y agoWhat AI can sometimes add is automatic feature extraction. Rather than having a to have a human explicitly think "I bet we can identify people via $x" and writing bespoke code to do it. E.g. this has been the big deal with AI in medical diagnostics, that it can look at the same data a very skilled human doctor sees, and still manage to discover something that the human couldn't, due to unnoticed features
- marginalia_nu 3y agoYeah but we've got old techniques for that too, Latent Dirichlet Allocation etc.
- deleted 3y ago[deleted]
- MauranKilom 3y ago"Only 5% of users use this feature, so we will remove it to save development efforts." As seen in Firefox..
- marginalia_nu 3y agoIt's getting off topic, but the irony of a browser with a 2-4% user base pulling this shit can't be overstated.
- mhluongo 3y agoWhy? They have a fixed budget after all.
- suprfsat 3y agoAnd those GPT-3.5 calls ain't cheap.
- marginalia_nu 3y agoBecause they're very much the victims of web designers with the same mindset, and if anyone should be the ones to recognize how unfair throwing minority user groups under the bus can be. It also doesn't really make strategic sense to focus on the lowest common denominator. Chrome already has that group. The one place they could eke out a loyal userbase is specifically the users that Chrome fails to capture because they have unusual needs or requirements.
- marginalia_nu 3y agoActually when I think about it, it's even worse than this. Firefox has been on the receiving end of this type of discrimination more or less for as long as it's existed. It was the state of affairs when IE was the challenger too. You have to be just mindbogglingly oblivious to not see how this has been one of their biggest problems the last 20 years.
- rcxdude 3y ago
- bo1024 3y agoDifferential privacy techniques are provably impossible to de-anonymize, if implemented correctly. It is possible. But fraught with possibility for error or manipulation.
- Aachen 3y agoThis is the answer. The person above can speculate and fearmonger what magic "AI" is going to be able to do, but if there is no personal data there to begin with, or if you use math like in differential privacy, there's not going to be a way to identify individuals. That is, if you suspect they'd change their minds and start trying to deanonymise previously collected data anyway — remember that open source distributions (I don't know fedora-the-organisation specifically) are generally made up of volunteers like you and me. Notable exceptions obviously exist, like for-profit Canonical; that's not the org type I mean or trust.
- jacquesm 3y agoIn practice though none of that matters because they'll do a slipshod job of implementing it.
- Aachen 3y agoWith that logic, we shouldn't have police or a judicial system either If we can't trust people ever, what's the point in doing anything?
- jacquesm 3y agoNo, obviously that's not even remotely the same. You need police and a judicial system and you fix them whenever they break. But you don't need telemetry, it's entirely optional and shoddy implementations translate into unnecessary risk. Also: https://lwn.net/ml/fedora-devel/H5JEXR.LLU011IQ4I6K@redhat.com/ https://lwn.net/ml/fedora-devel/H5JEXR.LLU011IQ4I6K@redhat.c...
- JohnFen 3y ago
- agloe_dreams 3y ago> Looking at the current developments in AI, I am concerned that AI models can easily de-anonymize and guess end point users when being fed with "telemetry data" of hundreds of thousands clients. I can almost guarantee you that the US government has a tool where you can input a few posts from a person on an anonymous network and get back all of their public profiles elsewhere. Fingerprinting tools beat all forms of VPNs and the like. Our privacy and anonymity died like maybe two years ago, there is no stopping it.
- JohnFen 3y agoThat may be true, but it doesn't mean there's no value in protecting your privacy from others anyway. Personally, I'm much more worried about private entities collecting information about me than I am about the government doing so.
- techwizrd 3y agoI disagree wholeheartedly. Privacy-preserving technologies like including privacy-preserving AI (e.g., federated learning, homomorphic encryption) and privacy-preserving data linkage/fusion are really important. They're crucial in my day-to-day work in aviation safety, for example. And telemetry is important. We have limited resources. How do we determine the number of users impacted by a bug or security vulnerability? Do we have a bug in our updater or localization? Are we maintaining code paths that aren't actually used? Telemetry doesn't magically improve user experience, but I'd rather make decisions based on real data rather than based on the squeakiest wheel in the bug tracker. We can certainly make flawed decisions based on data, but I'd argue that we're more likely to make flawed decisions with no data.
- m463 3y agoThere must be meaningful consent.
- JohnFen 3y ago> We can certainly make flawed decisions based on data, but I'd argue that we're more likely to make flawed decisions with no data. What I've seen in practice so far is that the use of telemetry has harmed software quality more than helped. It often leads developers to optimize for the wrong things and make poor design decisions. This happens because they tend to think that "the data never lies", ignoring the fact that telemetry always gives a skewed and incomplete picture.
- haswell 3y agoDo you have some specific examples of this playing out? I’ve been a product manager for products that had no telemetry, and that can be a rather undesirable place to operate, especially if you’re in the enterprise space where product changes can impact the operations of businesses. I think it’s certainly possible to focus on the wrong things, but I don’t see that as an outcome of telemetry itself as much as an outcome of a product team that doesn’t understand the problem space or customer base. The attributes to capture are presumably based on what teams understand to be key indicators about their app/service. I think confident incorrectness armed with bad data is just a slightly different version of a complete lack of data. Such a team was operating on whatever they imagined to be important before, and they continue to do so after, albeit with greater conviction. But good telemetry in the hands of a good product team can be immensely beneficial for decision making and can protect customers from bad decisions. Anecdotally, my ability to pull numbers about certain attributes has been key to my ability to shut down executive pressure to make changes that would have drastically impacted customers if not for the direct evidence that it would. I’m also not claiming that downsides don’t exist, and privacy is always my primary concern, but there are a range of outcomes based on the maturity of a team/company, and as long as the PM understands that data is not an alternative to having a relationship with customers, I think data is pretty important.
- JohnFen 3y ago> I am concerned that AI models can easily de-anonymize and guess end point users when being fed with "telemetry data" of hundreds of thousands clients. You don't need AI for this. This is done by real humans right now, using data points correlated from multiple sources.
- dschuetz 3y agoWell, it's simple: AI is getting cheaper, humans are getting more expensive.
- barbariangrunge 3y agoPeople keep saying "you don't need ai for this." Sure. But to do it at scale, and to intelligently connect disparate kinds of data contextually? That's time consuming and expensive without ai, so you can't do it at scale to a comprehensive degree. That hasn't been practical until now. It still isn't quite cost effective to do this for every human, everywhere, but soon it will be. Give it 5-10 years Thanks to ai
- didntcheck 3y agoThat's definitely a legitimate fear, as seen with the AOL controversy [1], but if they're just collecting aggregate statistics it's much less of a risk. I.e. User ANON-123 with default font x and locale y and screen resolution z installed package x1 Is clearly a big hazard, but statistics on what fonts, locales, and resolutions have is not really. Even combinations to answer questions like "what screen resolutions and fonts are most used in $locale?" should be safe as long as the entropy is kept low. It is less useful, since you have to decide on your queries a priori rather than being able to do arbitrary queries on historical data, but ethics and safety > convenience [1] https://en.wikipedia.org/wiki/AOL_search_log_release https://en.wikipedia.org/wiki/AOL_search_log_release
- Espressosaurus 3y agoCombine ANON-123 with information from their browser, which has default font x, locale y, screen resolution Z, and package x1, and that anonymous data just became much more rich. It doesn't take very many bits of information to deanonymize someone once you start combining databases.
- jacquesm 3y ago33 to be precise.