6 ms·
Exactly - if I could make any change to my engineering education, it would be to have taken even more statistics.
by outside2344 14y ago
Exactly - if I could make any change to my engineering education, it would be to have taken even more statistics.
- dclusin 14y agoAgreed. The introductory statistics class that I took as a computer science major wasn't enough. It really needs to be integrated into computer science curriculum's with specific attention to how it applies to writing software.
- chubot 14y agoReally? Can you name of some examples of where you've needed it in your work? (honest question) I am a programmer who works with a bunch of statisticians, doing "big data" stuff. What I observed is that most of them don't really spend very much time doing statistics. They spend all their time finding, collecting and massaging data. That generally involves a lot of programming. Once you get the data, the conclusions are fairly obvious without any statistics. Just make some plots and there are glaring orders of magnitude deficiencies. I also wanted to learn more statistics... but basically with software, you are overflowing with data. The challenge in science usually is to gather data. In software the challenge is the oppoosite -- you have so much data and you need to make sense of it. To be concrete I'm talking about stuff like logs from web servers and various other systems. So I learned a lot about sampling algorithms to cut down data, as well as various streaming algorithms. But I haven't actually learned that much about statistics. So I wonder if I am missing something.
- alexchamberlain 14y agoIt sounds like the sampling and streaming algorithms you know are statistical algorithms. Have you got any good references for streaming algorithms?
- chubot 14y agoThey're not really statistical algorithms -- you certainly wouldn't learn them in a statistics class. I actually looked online for some references on streaming algorithms ... but somewhat surprisingly, I couldn't find anything really. I realized a lot of this knowledge has been hard-earned, I guess that is good :) But there really should be a reference. Definitely look up "reservoir sampling", which gives you a fixed size sample of an infinite length stream. This algorithm has come up over and over for me, and I've implemented it in multiple contexts. There is a way to do it in MapReduce which is very useful. You know probably the most condensed version I can think of is quite hidden: see the open source Sawzall code: http://code.google.com/p/szl/source/browse/#svn%2Ftrunk%2Fsrc%2Femitters http://code.google.com/p/szl/source/browse/#svn%2Ftrunk%2Fsr... All those functions aggregate some property of a stream. Someone (maybe me if I finish the 30 projects I've started...) really should write up some real documentation about all those algorithms, because they're not only fun, but useful. A lot of them are (necessarily) approximations. You don't learn that many approximate algorithms in a traditional CS class. None of this will be in any stats class for sure.