9 ms·
Hello HN! We started developing Plausible early last year, launched our SaaS business and you can now self-host Plausible on your server too! The project is ba
by markosaric 6y ago
Hello HN!
We started developing Plausible early last year, launched our SaaS business and you can now self-host Plausible on your server too! The project is battle-tested running on more than 5,000 sites and we’ve counted 180 million page views in the last three months.
Plausible is a standard Elixir/Phoenix application backed by a PostgreSQL database for general data and a Clickhouse database for stats. On the frontend we use TailwindCSS for styling and React to make the dashboard interactive.
The script is lightweight at 0.7 KB. Cookies are not used and no personal data is collected. There’s no cross-site or cross-device tracking either.
We build everything in the open with a public roadmap so would love to hear your feedback and feature requests. Thank you!
- bishalb 6y agoHow do you track sessions without cookies? IP address?
- markosaric 6y agoWe generate a daily changing identifier using the visitor’s IP address and User Agent. To anonymize these datapoints, we run them through a hash function with a rotating salt. hash(daily_salt + website_domain + ip_address + user_agent) This generates a random string of letters and numbers that is used to calculate unique visitor numbers for the day. Old salts are deleted to avoid the possibility of linking visitor information from one day to the next. See full details here: https://plausible.io/data-policy https://plausible.io/data-policy
- benbro 6y agoIs it possible to track users over a week or month? Often the same user return to our website several times before buying and it's important to learn about this behavior.
- markosaric 6y agono. that's the decision we made in order to be as privacy focused as possible. there's no way to know whether the same person comes back to a site the day after or later. they will always be counted as a new visitor after the first day.
- propogandist 6y agothis will limit your solution from many deployments. Knowing returning visitors from first time visitors is quite important and helps to asssess if viewership, audience and customer base is growing over time. For startups the "how many unique visitors do you get in a month" may be an important KPI and you're saying your solution cannot answer this question, so another solution will be needed to be deployed. Unique visitor data's also needed to assess effectiveness of campaigns and run e-commerce operations. There's often campaigns to bring back a user who previously didn't buy (email, ads etc). It's important to measure the effectiveness of these investments separately in web analytics given the campaigns will be different for new and recurring visitors.
- markosaric 6y agoyeah, i understand. we're trying to have a balance between these: 1. privacy of site visitors 2. compliance with privacy regulations 3. useful and actionable data for site owners it's difficult to track people from visit to visit or from one device to another without breaking the first two (cookies, browser fingerprinting...) so we had to make some decisions. in general sites that try to get visitor consent to cookies and/or to tracking realise that majority of them don't give it, so even the data that may not be as accurate as full on tracking becomes very valuable.
- gabriel34 6y agoCorrelation will hold between unique hashes and unique visitors to access increase in unique visitors for your campaign. Even the percentage of these accesses that are returning is likely constant. All you would be giving up is the ability to measure variation in the returning percentage across several days (even so you could probably modify the code to change the salt every x days without losing much of the privacy benefits) There will always be value to be extracted from the invasion of your users' privacy, but you also hit diminishing returns over this increasingly invasive probing. Plausible is aiming for "good enough" whilst respecting people's privacy, and that is a good compromise IMHO. There is a trade-off. You will never get 100% of the information without all the tracking, but there is information that represents more bang for the privacy buck. Would you not have a acceptable error increase in your decisions with a bit less information and a lot less privacy invasion? EDIT> I think more control to the user is better, so instead of canvas fingerprinting, shady cross-site tracking and all, I would rather have a uuid that my browser informed, but that I controlled, so I could be anonymous when I want to and be tracked when I don't care, or when I genuinely agrees it adds value.
- Hitton 6y agoThat seems pretty robust, but I have question about the salt. How does the daily salt part work? How is it stored during the day? Are historic salt values stored?
- ukutaht 6y agoThe salt is stored in a Postgres database and we run a daily job that generates a new salt and deletes old ones.
- rkwz 6y ago> hash(daily_salt + website_domain + ip_address + user_agent) That's a pretty interesting technique! I built a privacy friendly analytics project https://github.com/sheshbabu/freshlytics https://github.com/sheshbabu/freshlytics and I've been wondering how to correctly count unique visitors to a website. I don't store cookies or any PII data at all so by definition, it's hard to distinguish between two different visits - are they from same person or different people? An alternative approach is used by Simple Analytics - https://docs.simpleanalytics.com/uniques https://docs.simpleanalytics.com/uniques where they use referrer header to derive unique visits. They mention that they don't use IP addresses as they're considered fingerprinting. But it looks like a hash function (whose salt gets rotated daily) strikes a good balance between fingerprinting while maintaining user privacy. Any downsides to this approach?
- markosaric 6y agodownsides are basically that in order to gain the benefits for the user privacy and compliance with regulations, you lose a bit of accuracy depending on the situation. we cannot see whether the same person returns to a site on a different day so count them as a new unique visitor. i assume that using the referrer header to count uniques has even more downsides as i imagine the number of unique visitors with that method would be much higher than it actually is.
- JimDabell 6y agoDoesn’t this mean that if you accidentally serve your website on two domains (e.g. example.com and example.net with no redirect) it will count one visitor twice if they visit both domains? What is the benefit to including the domain in the hash? Not really keen on the use of the IP address. I’ve been behind load balancing proxies and weird mobile networks often enough to know that I can appear from a dozen different IPs in the space of an hour just by browsing the web normally with default settings. Have you considered requesting a 24 hour private cacheable resource and counting the requests on the server? Or is the browser cache too unreliable?
- ukutaht 6y agoMore accurate would be that we use the site_id in the hash. You can serve your website from many different domains but as long as the site_id in your script tag is the same, the hash remains the same. The site id is included in the hash to prevent cross-site tracking. Otherwise the hash would act almost like a third-party cookie and people could be tracked across different sites. > Not really keen on the use of the IP address. Yeah, it's not ideal. We do check the X-Forwarded-For header so as long as the proxies are being good citizens, the client IP is present in that header.
- dustinmoris 6y ago> Have you considered requesting a 24 hour private cacheable resource and counting the requests on the server? Or is the browser cache too unreliable? Nice idea.
- youngtaff 6y agoDoesn't that depend on visitors not sharing an IP address i.e. not behind something that does NAT?
- markosaric 6y agoyes, that's a trade-off in this privacy first approach. if several people are on the same ip address, visiting the same website and having the same user agent on the same day, they look the same to us.
- jakemal 6y agoLove the product! Since you asked, one request that I would find helpful is having the stat comparisons from the previous day based on the same time the previous day rather than the total. The comparisons almost always show some -X% because you are comparing a 24 hour period yesterday to a 24 - X hour period today.
- markosaric 6y agothanks! makes sense and i'd like to see that change too! i've added the request to our github now so we'll see what can be done: https://github.com/plausible/analytics/issues/344 https://github.com/plausible/analytics/issues/344
- baobabKoodaa 6y agoHey markosaric! Very interesting product and I'm really tempted to try it out. One suggestion to add to your documentation: you should mention drawbacks in your "comparison" pages. For example, the page where you compare Plausible to Google Analytics lists only advantages and doesn't indicate that Plausible could in any way be worse. But based on the description of how Plausible works, I'm fairly certain that people behind VPNs will not be properly counted (everyone with the same user-agent will be counted as if they were the same person). Use of VPNs is fairly widespread, so lack of fingerprinting will actually prevent you from accurately assessing how many visitors a site has.
- markosaric 6y agoThanks! Depends really on the perspective you look at it from. That post is from the visitor privacy / privacy regulations perspective. If you use GA full on, you will get some advantages such as storing cookies on devices but if you either make GA "privacy friendly" (enable IP anonymization, disable user-ID, disable cookies etc) or ask visitors to consent to being tracked you will get less accurate numbers than our method. We haven't noticed any issues with VPN usage as still you need people to use the same IP, same user agent and visit the same domain all on the same day to be counted as one with us.
- vimes656 6y agoIt's so great the whole code is available under the MIT license. After seeing many open source projects with paid services going for more protective licenses, like the AGPL or the newer eventually open licenses (i.e. Business Source License), I wonder what was your rationale to still pick the MIT.
- markosaric 6y agothanks! we're big fans of open source so wanted as permissive licence as possible. one of the things we spend a lot of time thinking about is how to make our project sustainable while still being as open as possible in everything we do. a bit more details in this post "How to pay your rent with your open source project" https://plausible.io/blog/open-source-funding https://plausible.io/blog/open-source-funding
- tommoor 6y agoIt's all fun and games until someone spins up a competitive hosted service using your MIT licensed codebase.
- ponderingfish 6y agoAnd ..... they've switched to AGPL.
- malisper 6y agoQ: Since users of Plausible are sending you their visitors IP addresses, doesn't that mean GDPR applies? GDPR mentions "disclosure by transmission" as an example of processing[0]. Since that's the case, doesn't that mean you need a legal basis for processing the data[1], you need a way to allow users to opt out of collection of their data[2], and your clients need to sign data protection agreements[3] with Plausible? I tried looking for Plausible's point of view on this and how they address the above, but I couldn't find anything. [0] https://gdpr-info.eu/art-4-gdpr/ https://gdpr-info.eu/art-4-gdpr/ [1] https://gdpr-info.eu/art-6-gdpr/ https://gdpr-info.eu/art-6-gdpr/ [2] https://gdpr-info.eu/art-21-gdpr/ https://gdpr-info.eu/art-21-gdpr/ [3] https://gdpr-info.eu/art-28-gdpr/ https://gdpr-info.eu/art-28-gdpr/