6 ms·
Hi guys! So as basis for my thesis on AI and NLP I've been working on a RRN-based text classifier that basically reads and analyzes privacy policies. It unders
by rameerez 7y ago
Hi guys!
So as basis for my thesis on AI and NLP I've been working on a RRN-based text classifier that basically reads and analyzes privacy policies. It understands that "we don't share your data with third parties" is privacy friendly while "we may share your data with anyone" is a potential threat.
I've then created this website with a bunch of analyzed services to showcase the most relevant info about each service along with other interesting stuff like recent data breaches or instructions to delete your account in said service.
Happy to answer Qs about the tech behind, it'd also be great to hear your feedback on what the site lacks and possible improvements!
- Pete-Codes 7y agoVery nice! The fact you mentioned Tinder's T+C on Twitter got my attention.
- jsingleton 7y agoLooks cool. It's a small point, but not pluralising words when the number is 1 shows attention to detail (e.g. 1 scandals for instagram).
- rameerez 7y agoThanks for the heads up!
- thecleaner 7y agoOh my god thank you so much for doing this ! I think its better (for me as a user) if you don't boil things down to a score as different people expect different things when talking privacy. It would help if you could simply highlight the potential problematic clauses in different privacy statements along with some reason why it might be problematic.
- nickodell 7y agoI don't agree. I just looked in my password manager, and I have roughly ~220 accounts across the web. If I want to go through that list and see which website rank well and which rank poorly, and I want to do that in under two hours, that gives about 30 seconds per service. In other words, giving a single score plus a two-sentence highlight is probably about the right amount of information.
- m-p-3 7y agoOr make the rank adjustable to some personal criteria that matches different privacy expectations.
- drusepth 7y agoThis would also be helpful in determining how to weight (or not) user feedback in the training portion. I just tried it out (the 10 questions) and there were at least a few I thought, "huh, I know some others would disagree with me on this" because I value X and they value Y more. Having scores that weight X more than Y would give me more accurate scores, while seemingly also giving other people more accurate scores at the same time.
- crusty 7y agoHow about a compromise - not a score for each policy but for "each" individual, or tranche of similarly concerned individuals. I go through a list of privacy options (maybe just once) and the point at which my "okay" becomes "not okay" determines my score. Then each policy simply passes or fails based on my score. And if you want more detail, the list of failure items for a given policy can be bulleted.
- wtvanhest 7y agoA good compromise would be a chrome extension that shows a 1-10 score. You click the extension to see clauses
- trickstra 7y agoNot for everyone, as thecleaner pointed out. You are assuming your requirements are universal among other users. Also, you are assuming the policies can be simplified to a weighted average of their parts, which is not necessarily the case.
- blondin 7y agowell, just checked and i already see the majority of the web gathering around "C". so really, we can argue both ways about this score thing...
- rameerez 7y agoThe scoring definitely needs some work – I think some factors should have more weight. Also services' data vary a lot so it's difficult to come up with a good measure for everyone. Ex: I try to take into account whether the service has had any recent data breach, so it penalizes a lot if it has but also scores rather low if it hasn't; privacy policies' length vary wildly and I think that also plays a large role... It needs some tweaking but I think with some improvements I'll reach a more accurate scoring
- epoch_100 7y agoFor a site trying to fight for privacy, don't you think it would be better to not use Google Analytics to track the people who visit your site?
- rameerez 7y agoYes, definitely. I mention it in Guard's own privacy policy, I don't like using it either, but reasons are: (a) it's the simplest and as far as I know one of the few free analytics tools available, (b) not having a measure of the website activity will make me effectively blind and unable to make decisions, (c) I don't send any personally identifiable event (and, for this matter, I don't send any events apart from page loaded events). I'm also open to suggestions to replace GA.
- epoch_100 7y agoYou could just not use user-level analytics; I run the infrastructure for PrivacySpy.org, and using CloudFlare's aggregate analytics has served us perfectly well.
- djsumdog 7y agoHave you considered hosting your own Matomo server? You could just scrape logs or use a Javascript tracker. But the data will stay local to you.
- burnaway 7y ago+1 for Matomo, it's a no brainer for any 'privacy service' to stay far away from Google Analytics.
- digitalengineer 7y agoI’m on mobile, so I can’t check if you’re already using this technique, but you could always use the anonimize IP function. This way the last 3 numbers of the IP data will not be send through the Google Analytics script. More info: https://www.jeffalytics.com/gdpr-ip-addresses-google-analytics/ https://www.jeffalytics.com/gdpr-ip-addresses-google-analyti...
- brianberns 7y agoHow did you create a data set large and accurate enough to be useful in training a model?
- rameerez 7y agoSome friends run an AI bootcamp and helped me finding the initial set of users to help me with labelling. Initial labelled data was generated mostly through them, both manually labelling and with the approach described in https://useguard.com/experiment https://useguard.com/experiment Also, the model I'm using relies heavily in transfer learning and achieves very reasonable results with few labelled items (the paper in which the technique is described actually maintains that with only 100 labelled examples they reached comparable results to using 10x that data in models that use older approaches)
- fny 7y agoWhat paper is this? Is it the UDA paper?
- burtonator 7y agoHey... can you reach out to me... I'm kevin ... at datastreamer.io (trying to hide that from spam but I think you can avoid that). I'm working on something similar at the moment for a client. Right now I'm just starting out but I've built a privacy policy classifier that is an RNN classifier based on TensorFlow that is just an 'is this a privacy policy' classifier. I have about 650MB of privacy policies at the moment which I fetched via a crawler. I'm just about to classify the rest of them. I'm trying to automate the whole thing so that we have a full workflow. Anyway... ping me to discuss.
- domnomnom 7y agoHow do we know it isn't just people doing the analysis if we can't actually use the AI ourselves?
- rameerez 7y agoI'm most probably publishing a paper later this year detailing the process. Also, 80 pages of my thesis would have loved this wouldn't have involved AI to make the whole thing simpler :)
- hanoz 7y ago> It understands that "we don't share your data with third parties" is privacy friendly while "we may share your data with anyone" is a potential threat. Does it understand "we don't share your data with just any old third parties", or "we're not like our competitors who may share your data with anyone"?
- davidkuhta 7y agoNot OP, but curious, are those quotes actually from privacy policies or just hypotheticals?
- drusepth 7y agoWith enough sites / privacy policies out there, even hypotheticals will end up in a real policy at some point (if not already). Is there something significant in the distinction or is it just curiosity? If the latter, I also share it. :)
- goldemerald 7y agoDoes your work integrate pretrained LMs like BERT or GPT2?
- rameerez 7y agoYes, but not Transformer-based like these two, rather LSTM-based like ULMFiT
- jph00 7y agoInteresting! Did you have to make any substantive changes to the ULMFiT approach to make it work for this problem? Did you use the fastai implementation, or write something from scratch? (Disclaimer: I'm a co-author on the ULMFiT paper.)
- naveen99 7y agooff topic, but i saw in one of your fast.ai videos, you were running an autohotkey script. Made my day !
- alexcnwy 7y agoGreat work - would love to read your thesis if it’s available online?
- joshspankit 7y agoImportant work, and this seems to be doing a decent job already. Cheers. One thing about the teaching: some sentences don’t have anything to do with privacy, so there might be a button to train the AI on that.
- godelmachine 7y ago>>”along with other interesting stuff like recent data breaches” Do you have an integration with HaveIBeenPwned?
- crusty511 7y ago> Happy to answer Qs about the tech behind [...] What tech is involved to get something like this on the web?