5 ms·
For those not familiar with Posthog, it's open an open source product analytics tool. That means you install some javascript on your website. It then collects i
by malisper 6y ago
For those not familiar with Posthog, it's open an open source product analytics tool. That means you install some javascript on your website. It then collects information on what features people are using on your website and let's people run charts and graphs to ask questions about feature engagement and product usage.
Product analytics is a pretty large space. The three main existing product analytics tools are Mixpanel[0], Amplitude[1], and Heap[2] (I used to work at Heap). Notably, Mixpanel was valued at $865MM as of 2014 and Amplitude was valued at $1B as of two weeks ago.
Posthog is an open source competitor to these tools that you are able to self host. This is interesting for several reasons. If you don't feel comfortable sending data to third party tools, you control the data yourself. In addition, existing product analytics tools can get expensive. After you get off the public pricing, you will likely find yourself paying >$12k/year for one of these tools.
There are also challenges I see for them going forward. In my experience, the person who makes the decision to buy a product analytics tool isn't in engineering. It's usually a product manager or sometimes a marketer. I imagine the fact that it's open source is less interesting to one of the traditional product analytics purchasers. Compare this to a company like Gitlab where the end users are engineers.
Scaling these tools is also a huge challenge. My job at Heap was to basically scale Heap. If you want to self-host a product analytics tool, you will basically be responsible for scaling the system. On top of that, because each Posthog instance is isolated from each other, Posthog won't be able to take advantage of the shared compute power. Imagine a product analytics tool that has 100 customers and 100 servers. If the compute power was shared, each query would make use of all 100 servers. If the servers were isolated from each other, each query would only be able to make use of 1 server. That's basically a 100x slow down in performance. Posthog does sell a managed version which in theory shouldn't have this problem. It will be interesting to see long term whether that's a big driver of revenue for them.
[0] https://mixpanel.com/ https://mixpanel.com/
[1] https://amplitude.com/ https://amplitude.com/
[2] https://heap.io/ https://heap.io/
- smithmayowa 6y agoAnalysis like this is why I love hackernews, thank you.
- fergie 6y agoStupid question: Can't you do all of this in Google Analytics? What are the advantages of Posthog/Mixpanel/Amplitude/Heap over GA?
- kevsim 6y agoNot a stupid question. These products tend to offer a lot more functionality and flexibility in terms of how you consume your data. Different analysis possibilities, fancy ML models, custom dashboards, etc. Also more control over how the data is ingested (sampling or not, etc.) These products feel more suited to teams building rich apps/webapps where as GA seems to have its roots in sites that are more "content-based" (news sites, etc.). And for many, not sending the data to Google is an advantage in and of itself.
- dotandgtfo 6y agoI'll also add that product analytics software is built more towards measuring depth of engagement: How many times has this user returned to our site? What is the average LTV of this cohort? What is cohort X's n-day retention compared to cohort Y? What features have correlation to those differences? Did this experiment lead to improvements in retention - 30 days later? Google isn't built to go that deep. Sure you can see that 3% of users have used feature X, but you can't really effectively dig in and see how upstream or downstream actions and events influence each other. Sure you can create some custom segments on a sessions/user level - but that quickly turns complex and unwieldy if you have several segments, cohorts and funnels. Also there are a lot of charts that are just plain better in product analytics tools. Retention charts, funnels and path diagrams are obvious examples.
- labelbias 6y agoAt my previous employer we were using GA and it's very flexible. We were able to track all DOM interactions by just configuring the included js. Although, querying that later with tools engineers can use and build is a whole different story.
- malisper 6y agoThere's two issues with GA. First the functionality is somewhat bare bones. It does a really good job out of the box, but it's pretty hard to use beyond that. Take the funnel feature. You need to define your funnels ahead of time before you can start analyzing them. Compare this to any of the other product analytics tool where as long as you were collecting the events up front, you can compute any funnel you want. Second, once you cross the free tier, you will have to pay at least $150k/year. The free tier is 10 million hits per month[0], which is high enough that most businesses never hit it. Other analytics tools often come in at a better price point at that kind of scale. [0] https://marketingplatform.google.com/about/analytics/terms/us/ https://marketingplatform.google.com/about/analytics/terms/u...
- james_impliu 6y ago(I'm the person that wrote the article) A few thoughts: As you say, the thing that's really different about us is exactly that we are focussed on engineers not PMs with our tool. PMs are of course welcome to use it too, but it's a little more technical feeling. We felt the people building the thing should have the context of usage data and this shouldn't only sit in another team - which was a behavior we saw happening quite frequently when we did user interviews early on. We started with the fundamentals of product analytics with some features we wanted ourselves but are now focussed on features that are much engineering specific. For example, we are about to release (an optional) "inspect element" for usage data: https://github.com/PostHog/posthog/issues/870 https://github.com/PostHog/posthog/issues/870, so you can pop that up whilst working on localhost. We had no idea if that hypothesis was correct - that engineers would care. We did a launch HN to find out and got quite a lot of good feedback and growth. Re scaling, you are spot on - this is definitely hard! We have users doing 5 million events/day on Heroku's cheapest standard tier dyno. We offer paid support for people who need something higher volume still if it doesn't work well out of the box. We're working on supporting databases other than Postgres as people need them.
- billme 6y agoCurious, in the article, you mentioned - “we believe that open source will eat SaaS's lunch in many product categories” — what to you is the core filter for “SaaS vs Open Source” product market fit?
- james_impliu 6y agoI think the closer your product is to something that developers can use the stronger the OS proposition. If there are other value props (privacy or cheaper at scale) that can definitely help though.
- malisper 6y ago> We felt the people building the thing should have the context of usage data and this shouldn't only sit in another team - which was a behavior we saw happening quite frequently when we did user interviews early on. What is preventing the engineers from getting access to that data? Existing product analytics tools usually don't charge by seats, so if they wanted to look at the data, they should be able to get access to Mixpanel, Amplitude, etc. Compare that to Posthog's pricing of $25/user/month which incentives you to minimize the number of people at your company that have access to Posthog. > We had no idea if that hypothesis was correct - that engineers would care. We did a launch HN to find out and got quite a lot of good feedback and growth. Heap also found a lot of success on HN when they launched: https://news.ycombinator.com/item?id=5424206 https://news.ycombinator.com/item?id=5424206. While there was a lot of excitement early on, the kind of crowd it attracted wasn't a good demographic to sell to long term. > We have users doing 5 million events/day on Heroku's cheapest standard tier dyno. The challenge isn't ingesting the data. Especially with an analytics tool where you can turn off synchronous_commit so ingesting an event doesn't require writing to disk. The bigger challenge is around processing queries that span years of data. Queries get proportionality slower the more data you've collected. A query over 12 months of data is going to be 4x slower than a query over 3 months.
- mohitmun 6y agoJust wanted to highlight that heap's engineering blog is one of the best I have ever come across. Love all articles by Kamal and Michael malis
- malisper 6y agoI am Michael Malis :)