3 ms·
> Look at Google's features like spelling correction or search term suggestions - they likely used huge troves of semi-anonymous user input to develop and suppo
by danielrhodes 4y ago
> Look at Google's features like spelling correction or search term suggestions - they likely used huge troves of semi-anonymous user input to develop and support them.
There's a big difference between data which is directly tied to PII and data which is held in aggregate in terms of privacy. I'm not arguing there can't be leakage here, but it certainly blunts many of the more severe privacy implications. Conflating these two is more sensationalist rather than useful in terms of honing in what is ok vs what is not.
> which means you can store a megabyte of data for every human on earth for a bit under $300k a year
Sure, I mean as a slippery slope you could also write this data to paper and keep it indefinitely at a very cheap price. My point here is more: if you plan on using the data, it becomes more and more expensive as your access patterns change. This also has privacy implications because I would imagine the easier the data is to access in raw form the higher the potential privacy cost to the user. If all they are doing with these key strokes is recording somewhere that you might be interested in Corgis and German Shepards based on your keystrokes, as opposed to something more detailed like an accidental paste of your password, I think that changes the conversation.
- aeturnum 4y ago> My point here is more: if you plan on using the data, it becomes more and more expensive as your access patterns change. I don't think this is true either. Google does not need to keep much of its data in hot storage to use it effectively: their ML products can be periodically trained / updated, their search can be iteratively updated with each crawl, etc. Sure, it would be expensive to keep all user data from all sources in hot storage all the time - but it's not needed. The idea that you...would happen upon some new question you hadn't though of before and need to get the answer immediately is just false. Instead, you make regular updates to a model and periodically run your corpus through that model.