3 ms·
An interesting statistical example if you haven’t seen it before is the Anscombe Quartet : https://en.wikipedia.org/wiki/Anscombe%27s_quartet https://en.wikipe
by jeffwass 2mo ago
An interesting statistical example if you haven’t seen it before is the Anscombe Quartet :
https://en.wikipedia.org/wiki/Anscombe%27s_quartet https://en.wikipedia.org/wiki/Anscombe%27s_quartet
Four sets of X,Y datapoints that have exactly (or very close) common statistical parameters (mean, variance, correlation, linear regression, R^2), but with vastly different spatial distributions and “behavior” when looked at visually.
- alphabeta3r56 2mo agoI have mostly worked with computational physics with discontinuous polynomial approximations. Stuff like this would be easily detected in those methods as anatomy brcsuse of assumption pf smoothness. Statistics on the other hand is more accepting of discontinuous data due to it's basis in measure spaces. Hence In general, you should know what your data should look like before you aim to detect anamolies. But also, it might be a good practice with new automated research actors to always use both approximations (measure theory based and otherwise), to figure out what's going on.
- touisteur 2mo agoCool that the Wikipedia page links to the "Mean Dinosaur" paper https://dl.acm.org/doi/10.1145/3025453.3025912 https://dl.acm.org/doi/10.1145/3025453.3025912 that I love to get out each time someone sends me mean, median, or stddev to measure processing latency. By all means use stats, but always eyeball the dataset to check assumptions extracted from statistics, I guess, especially in this world of matplotlib and notebooks and agents.