3 ms·
I found the DS100 course textbook and the author describes the phenomenon here: https://www.textbook.ds100.org/ch/02/design_srs_vs_big_data.html https://www.tex
by mendeza 7y ago
I found the DS100 course textbook and the author describes the phenomenon here: https://www.textbook.ds100.org/ch/02/design_srs_vs_big_data.html https://www.textbook.ds100.org/ch/02/design_srs_vs_big_data....
This is insightful, but how does one deal with determing if your dataset is non-random? I can imagine manually inspecting the data and using cross-validation are ways to help identify skewed datasets. Are there any other ways to test if your data is non-random?