4 ms·
I'm not sure which schools you're referencing when you write about that status of academic teaching. If a university has a statistics program they must offer s
by mattrepl 16y ago
I'm not sure which schools you're referencing when you write about that status of academic teaching. If a university has a statistics program they must offer something beyond a course in basic statistics.
Grants for statistics research do seem to be relatively lacking. Possibly the increased emphasis on statistics in data analysis and machine learning will change that.
The listing of good statistic programs is missing at least a handful of schools: UCLA, UMich, UWisc, CMU (machine learning dept), and GMU (computational stats PhD).
I'm coming from computer science, but it appears the state of statistics in industry and academia is improving. Off the top of my head, here are some tech companies that depend on and employ statisticians: Facebook, Google, bit.ly, FlightCaster, Twitter, BackType, OkCupid, and Microsoft. The title might be Data Scientist or Search Engineer, but it's still statistics and probability.
- HilbertSpace 16y agoYes, problems in machine learning, computational statistics, data scientist or search engineer should be in 'statistics'. But your generous attribution that this work now is actually statistics has a serious failing: The backgrounds of both the students and the professors rarely have the prerequisites for making serious progress with statistics. Maybe their work NEEDS 'statistics' but their backgrounds rarely permit them to make progress in 'statistics' even on the problems they are addressing. E.g., I worked in one of the world's better artificial intelligence groups, and coming through were bright graduate students, part time, summer, etc. all awash in 'machine learning', etc. The work was junk, and here's why: There were some people in computer science departments who wanted some progress in something roughly related to computer software that 'learned' in some sense. So, they tried things. Mostly they tried just heuristics, especially ones they got from just maybe guessing at how humans did things. Or for a while they were all hot on 'neural nets' -- promising for simulating a few neurons in an earthworm. The criterion of progress was mostly just, did the resulting software appear to do something good? Basically they were just starting with a blank slate and next to nothing significant in powerful background in anything. About the deepest thing they knew was, maybe, LALR parsing. They knew how to program computers but didn't know what to program. In addition, and much more serious, there was a 'methodological gap' the width of the Pacific Ocean. That criterion of, does the resulting software appear to do something good, is, in the history of statistics, applied math, pure math, and mathematical applications in physical science and engineering, just JUNK. Such software might be this and that, but it's NOT 'statistics'. Finally I had to explain to one of those intuition pushers: Here's how to exploit mathematics to do applied problem solving. In a transportation analogy, we start with a real problem a point A and we want to get to a real solution at point D. For this trip we can start out walking, without a map, across deserts, rivers, swamps, and oceans, and maybe get to D. But here's what we SHOULD do: First get a taxi ride to a local airport Assumptions at B. At B get a plane trip on Mathematics Airline to another airport Conclusions at C and close to D. Then take a taxi to D. The math connects the assumptions to the conclusions. Here the math is based on theorems and proofs. The taxi trips are supposed to be SHORT. Then for the logical validity of the work, we get to check the math and then argue over the taxi trips. The key is the math, and there we look carefully at assumption, theorems, proofs, and conclusions. The 'research' content is new math for new connections for a new pair of assumptions and conclusions and, usually, a new pair of problem and solution. To do such math, one needs a good background in math, typically an undergraduate major in pure math and about two years of graduate school in focused applicable math. Of high importance will be the 'mathematical sciences' with probability (based on measure theory), statistics (also based on measure theory and functional analysis, e.g, for weak convergence and sufficient statistics), stochastic processes, especially Markov processes and martingales, and optimization, including combinatorial and stochastic. E.g., without the first half of Rudin, 'Real and Complex Analysis', just F'GET about it; that material is not sufficient but it is NECESSARY. As an example, once I took an important problem in practical computing and executed the 'methodology' above. I got some nice results for the real problem and, really, a nice step forward for computer science. My paper connected carefully with instances of the real problem and with real data, but the core of the work was the new theorems and proofs. So, I went to get the paper reviewed. I got back from two chaired professors of computer science at famous research universities and editors in chief of top computer science journals essentially the same wording, "Neither I nor anyone on my board of editors has the background to review the math in your paper.". For a third such person, I wrote them tutorial notes for two weeks before they gave up. At one journal, the editor gave up but the editor in chief stepped in and handled the paper himself. Likely the computer science people he asked said, "It's nice for practical computing and computer science, but I can't say if the math is correct.", and likely math professors said, "The math is correct, but I don't know what it means for computer science.". Broadly, so far research computer science can rarely do research in applied math or mathematical statistics. A Ph.D. in computer science just does not have the right prerequisites for such research. The prerequisites are a focused applied math Ph.D. Sorry 'about that. So, yes, for broad areas of needed research progress in computer science, the computer science community is essentially irrelevant. As I said, for 'statistics', the computer science people are limited to intuitive heuristics and picking formulas they don't understand out of elementary stat books, and that is like snake oil cooked up on a wood stove. Again: The reason is, they don't have the prerequisites. Uh, my paper could have been called 'machine learning' or 'artificial intelligence', but I just called it some theorems and proofs for some progress in computer science for an important problem in practical computing. It would also be appropriate to call the work some nice progress in mathematical statistics.