5 ms·
Or, we could stop sending people programs to track themselves on their own computers. Just use your server logs. I don't care how much of an inconvenience it is
by nathcd 7y ago
Or, we could stop sending people programs to track themselves on their own computers. Just use your server logs. I don't care how much of an inconvenience it is, you should never have started doing any client-side tracking in the first place.
- smashingfiasco 7y agoI concur.
- ohithereyou 7y agoBut if I do that then how am I going to convince investors to give me so much funding I can sell my company and retire?
- pythux 7y agoGreat point! We wrote about how we're doing it on whotracks.me a while ago, if anyone is wondering how that could work: https://whotracks.me/blog/private_analytics.html https://whotracks.me/blog/private_analytics.html
- marcus_holmes 7y agothis is useful, thanks... couple questions: 1. Do you grab the country from the IP address before anonymising it? (I know VPN's are a thing, but we would like to know where our customers are coming from) 2. Why encrypt with a daily key instead of just hashing the IP address and storing the hash? Do you ever decrypt the IP address? 3. How have you found CloudWatch? I'm logging all requests to our own database, because it's so simple to do. What's the bonus for Cloudwatch?
- kkm 7y agoHi Marcus, Thank you for your questions. Disclaimer: I also work on WTM from time to time, so I hope can answer these questions to some degree :) 1. WhoTracksMe is served via CloudFront, it only get's the IP that CF reports. To get the country, there are two ways that can be used. IP => Country mapping via a Database or additionally, because CF has multiple edge locations, if needed the you could see which edge location served the request, that will help you narrow down the location at a country level. Ofcourse, in both cases if the user is using VPN, you would only records the VPN location. 2. The problem we want to avoid is same IP with same hashing algorithm will produce same value. Which can then be used to co-relate user activity across days. Second, depending on what hashing scheme one is using, rainbow tables can be used to get back the original value. Therefore, to avoid, we used the approach of daily key. Now that I write, we could also use some hash + daily_salt. This should give the same guarantees. - Needs to be checked. 3. In this setup we are not using Cloudwatch but Cloudfront, which is the CDN provider from AWS.
- marcus_holmes 7y ago1. Ah, OK, that's useful. I was thinking that dropping the last byte of the IP address would still be able to give us the country via database lookup, but anonymise enough. But getting it via CF (or similar) would be a good alternative. 2. Chances of a hash collision on IP address is pretty small (i.e. statistically insignificant). But adding a daily salt would decrease the odds, for sure. 3. Sorry, my bad, I meant Cloudfront, but it's Sunday night here and there has been wine ;) Same question: how has Cloud{{thing}} worked out for you? Is it worth the extra complexity?
- FabHK 7y ago> 2. Chances of a hash collision on IP address is pretty small (i.e. statistically insignificant) The problem, as I understand GP, is not so much the fear that two different IPs might collide. The problem is that seeing the same hash, you know that it refers to the same IP, and as such you can correlate users/sites over time (a privacy invasion).
- marcus_holmes 7y agoThat makes sense, kinda. Though I'm a little unsure why correlating http requests over an hour is not invading privacy, but doing the same over a week is?
- marcus_holmes 7y agoWhat's the huge objection to client-side tracking over server-side tracking?
- salawat 7y ago20+ years ago, we wouldn't discussing the ethics of server/client based tracking with a straight face, because that shit is creepy as fuck. Now, everyone takes tracking for granted as if the option of not being tracked is some sort of unachievable goal. That's the issue. It isn't tracking A as opposed to tracking B. It's tracking period.
- marcus_holmes 7y agoI need to understand why my customers are not clicking on the "subscribe" button. I'm not selling adverts to them. I'm not capturing personal information. all I'm doing is trying to understand who my customer is and why they're not buying my product. Why is this unethical? And why is it OK to analyse server logs, but not OK to use javascript, to do this?
- nathcd 7y ago> And why is it OK to analyse server logs, but not OK to use javascript, to do this? The data in your server logs is data that people sent to you on purpose (by visiting a page on your site, or filling out a form on your site). If you collect data about people by sending them code to execute and send data back to you, they're not meaning to send you that data; you're doing it behind their back. I'm glad you try to use data you collect responsibly, but if people aren't meaning to send you data, you're not entitled to it.
- mojuba 7y agoUsers are free to block the execution of client-side scripts though we all know what that means: much poorer to non-functional UX in many cases. Plus, to an extent browsers can control what can be collected on the client side, e.g. location, microphone input etc. (I wish there were similar restrictions on the keyboard and mouse, maybe some day we'll see that too) But I'd agree that letting the users use a public resource on a condition that they should allow arbitrary scripts to be executed on their side can be considered unethical.