3 ms·
1) Design the reports you want. Pay special attention to interactive elements like filters and drilldowns. List all dimensions and metrics you need. Think about
by srrr 7y ago
1) Design the reports you want. Pay special attention to interactive elements like filters and drilldowns. List all dimensions and metrics you need. Think about privacy.
2) Find your visualisation tool of choice. This is more important than any architecture choice for the tracking because this makes your data useable. [1]
3) Select your main data storage that is compatible with your visualisation tool, data size, budget, servers, security, ... SQL is always better because it has a schema the vis tools can work with. For a low amount of data you might just want to use your existing database (if you have one) and not build up new infrastructure that has to be maintained.
4) If you need higher availability on the data ingress than your db can provide use a high availability streaming ingress [2] to buffer the data.
5) Design a schema to connect your db to the visualisation tool. Also think about how you will evolve this schema in the future. (Simplest thing in sql is: Add colunms.)
I hope this helps. If you have selected some tools it is fairly easy to search for blog posts and tech talks. But don't think to big (data). "A few thousands users" and "two dozen parameters" may be handled with postgres and metabase. Also in most enterprise enviroments there already exists a data analytics / data science stack that is covered by SLAs and accepted by privacy officers. Ask around.
[1] https://github.com/onurakpolat/awesome-bigdata#business-intelligence https://github.com/onurakpolat/awesome-bigdata#business-inte...
[2] https://github.com/onurakpolat/awesome-bigdata#data-ingestion https://github.com/onurakpolat/awesome-bigdata#data-ingestio...
- ironchef 7y agoCouple bits (good overall): "1) Design the reports you want. Pay special attention to interactive elements like filters and drilldowns. List all dimensions and metrics you need. Think about privacy." I think what you're getting at here is figure out what information you want to get and then work backwards to figure out if you have the data. A couple minor changes I'd make: A) don't just figure out a a report, figure out what actions you'd want to see. If something gets above or below a threshold, who should be doing what? (Reports for the sake of reports is generally bad) B) Are you trying to build things that will push for operational, tactical, or strategic change? The manifestations of those are often very different. Operational bits are often dashboards / KPIs, whereas with strategic changes we often would want to present something more akin to a story. C) Privacy - Think of GDPR / PII _now_. Look at each metric / dimension and understand the data classification of it. 2 and 3 are tied to each other. You could have visualizatoin drive storage or visa versa. Just understand the tradeoffs. I'd suggest for who's done it before / talks, etc. there are a ton out there. There's those chats from the FAANGs and various groups in the valley (Lyft, etc.). Tons of blog posts there. Vendors have (largely predisposed towards them) builds. Finally the talks/slides at datacouncil and strata often contain lots of more .. "pointed" information. The high level bible that lots of folks would say look at is Kleppmann's "Designing Data-Intensive Applications: The Big Ideas Behind Reliable, Scalable, and Maintainable Systems".