4 ms·
I think one of the big problems with statistics for most people, besides the complexity of the math, is understanding when to use which techniques. How will you
by buss 11y ago
I think one of the big problems with statistics for most people, besides the complexity of the math, is understanding when to use which techniques. How will you help the user choose the correct way to analyze their data?
- augb 11y agoThank you for your reply. I agree with you. For the Free Tier the idea is provide some basic descriptive statistics with maybe a histogram overlaid with the normal distribution curve on top. Some example values that might be calculated are the mean, median and mode along with several standard deviations. Down the road a bit, I could envision a feature to suggest further exploration, which would likely guide the user towards some of the Pro features. (Although, I would like to expand what is available in the free tier, where it makes sense.) For the Pro Tier, they would of course, have access to the same results as the Free Tier, but I imagine a wizard like interface (which can be turned off) to help guide the user through choosing what they want done. Initially, the focus would be on helping spot basic trends and more obvious data anomalies. Further down the road, I could see the options expanding. One of the challenges to be handled is the "width" of the data set. For example, if User A has a CSV with 100 columns, the basic calculations may not be a big deal, but do we present a 100 histograms? What if there are 1,000 columns. My current thinking is for the user to prioritize the columns they want analyzed if they exceed a certain threshold (say, 25 columns). The results would be done on the columns up to this threshold. I am open to ideas and suggestions on this, and even on things such as which "basic" descriptive statistics are most important. Edit: Clarification (last sentence)