4 ms·
I'm one of Parse.ly's co-founders. This post was written by one of our product managers about a project and investigation we've been doing for the past few mont
by pixelmonkey 6y ago
I'm one of Parse.ly's co-founders. This post was written by one of our product managers about a project and investigation we've been doing for the past few months. It first got on my team's radar when I posted this set of tweets back in 2019:
https://twitter.com/amontalenti/status/1165262620959617025 https://twitter.com/amontalenti/status/1165262620959617025
Specifically: I noticed a huge difference between the metrics we were reporting on my blog post in Parse.ly, and the metrics being reported by my personal blog's Cloudflare CDN (caching the content).
Ironically enough, this traffic was all coming from HN and the post was itself about modern JavaScript[1].
Since then, we've also been hearing from a lot of customers about various scenarios where traffic is either under-counted or mis-counted. For example, something that has been tripping us up lately is that our Twitter integration relies (partially) upon the official t.co link shortener[2], and yet, due to modern browser rules related to W3C Referrer Policy[3], the t.co link's path segment is often not transmitted to the analytics provider, and thus the source tweet for traffic cannot be easily ascertained.
I firmly believe in privacy and analytics without compromise[4], so the team is trying to come up with ways to at least quantify shadow traffic at an aggregate level, and to ensure legitimate user privacy interests are honored, while making sure they don't break legitimate privacy-safe first-party analytics use cases.
As a developer, something that concerned me recently was realizing that Sentry, the open source error tracking tool with a SaaS reporting frontend and a JavaScript SDK[5], gets blocked in many conservative browser privacy setups. Though the interest to user privacy is legitimate, I think we can all agree it'd be better for site/app operators to know when certain browsers are hitting JavaScript stack traces.
[1]: https://news.ycombinator.com/item?id=20785616 https://news.ycombinator.com/item?id=20785616
[2]: https://help.twitter.com/en/using-twitter/url-shortener https://help.twitter.com/en/using-twitter/url-shortener
[3]: https://www.w3.org/TR/referrer-policy/ https://www.w3.org/TR/referrer-policy/
[4]: https://blog.parse.ly/post/3394/analytics-privacy-without-compromise/ https://blog.parse.ly/post/3394/analytics-privacy-without-co...
[5]: https://sentry.io/for/javascript/ https://sentry.io/for/javascript/
- MauranKilom 6y agoSo what do you actually (want to) do in regards to measuring shadow traffic? The blog post tries to convince the reader they should care about shadow traffic and then handwaves away existing solutions, but "we're developing a solution" is as concrete as the post gets as to why I should turn to parse.ly. Now you say "the team is trying to come up with ways to at least quantify shadow traffic at an aggregate level", so it appears that you don't even have a solution yet. In light of that, presenting "existing analytics services like parse.ly" as one of three solutions on "how to measure shadow traffic" seems borderline disingenuous. If you can do it, why not say so plain and clear? If you can't do it, why do you mention yourself as a solution? Or is it only other services like parse.ly that can do it? It also rubs me the wrong way how both the blog post and your comment has an undertone of "it would be better if users didn't have as much tracking protection". Just take the framing of your last sentence as an example...
- pixelmonkey 6y agoActually we have a few accidental "solutions" to this problem already in production and we are just trying to figure out which one meets the right overlap of respecting privacy preferences and providing site/app owners with visibility into shadow traffic. Here they are: - Server-side proxy: The blog post I linked about The Intercept uses this setup. Basically, a web server run by our customer captures all the traffic inside their cloud or hosting environment. That traffic is then logged and proxied to our data capture server, with some data scrubbed before we receive it (e.g. IP address removed), with data sent via our server-side protocol. - First-party custom domain: We spin up a server and HTTPS certificate and the customer points their own subdomain (via a DNS A or CNAME record) to that server, which serves as a proxy. We originally built this facility to clarify data ownership in a GDPR context -- where the customer is a data controller and we are a data processor, so the controller owns the domain where data ingest happens. This and the prior solution have the side-benefit that Parse.ly couldn't do any cross-site linking even if third-party cookies were enabled in the browser. We never do this anyway, but both of these setups make it technically impossible due to the browser rules around cookies and domains, which is a nice security improvement. But it also raises other issues, like the fact that the customer setup is more complex, with more moving parts. - API connection to CDN. This one is not actually productionized but was merely prototyped. We'd pull basic per-day and per-page CDN server request logs, and compare that to our pageview counts to understand the delta, which is likely mostly shadow traffic. The upside of this solution is that it might be pretty easy to setup for customers, the downside is that we'd have to build connectors for a lot of CDNs, and through market research we have learned that larger customers might use multiple CDNs at once (believe it or not). - "Fallback" logging of blocked page loads. This one was also just prototyped, but the idea is that some JavaScript code would detect whether Parse.ly JavaScript SDK was blocked from loading, and if so, a basic privacy-safe "this page's analytics were blocked" event would be sent to a domain owned by the customer, perhaps one that ensured scrubbing of all details other than the "fact" that a block event happened at a particular timestamp. We actually prototyped this particular idea on our own marketing site because we ran into issues with Marketo & Parse.ly data vs our server logs and even our lead capture forms. (That is, situations where a lead was captured for someone with "zero pageviews", because their session was shadow traffic but their form fill nonetheless happened.) Re: your comment that you sense an undertone of, "it would be better if users didn't have as much tracking protection", I have no such personal or professional belief, and I can assure it isn't a view held by our company/team. I understand the motivation for tracking protection and we even suggest use of Mozilla Firefox's tracking prevention option in our privacy policy. But there's no doubt that it is leading to confusing data discrepancies for site/app owners, and I think site owners have a right to a basic understanding of how much of the traffic they are paying the hosting bills to serve is actually perceptible to their observability/reporting, even if the only detail they get about that visit is "the visit happened", similar to the level of detail they get from server logs or CDN logs as a matter of course.
- inetknght 6y ago> Though the interest to user privacy is legitimate, I think we can all agree it'd be better for site/app operators to know when certain browsers are hitting JavaScript stack traces. I don't want site owners to see that something has failed on my machine -- especially if it's something unique about my setup. To suggest otherwise is to miss the point about privacy. So no, I don't agree.