5 ms·
Can someone say in few sentences what Statistics is all about? I can't shake off the feeling that it is just glorified curve fitting. Edit: Please stop the dow
by MichailP 10y ago
Can someone say in few sentences what Statistics is all about? I can't shake off the feeling that it is just glorified curve fitting.
Edit: Please stop the down votes, just an electrical engineer here, with one basic course in Probability and Stat. :)
- jsprogrammer 10y agoStatistics is applied Probability Theory. Statistics tries to find and characterize the probability distribution of 'Random Variables' through series of observations. Basically, you count things and compare that to how many things you think you should have counted given your assumptions.
- johnloeber 10y agoInferring meaning from data.
- DavidSJ 10y agoIt is just glorified curve fitting. Glorified curve fitting is a very rich field. Another way of thinking about it (described in Wasserman's book) is that statistics is the inverse problem of probability. Probability theory asks: given a process, what does its data look like? Statistics asks: given data, what process might have generated it?
- glial 10y ago> Probability theory asks: given a process, what does its data look like? Statistics asks: given data, what process might have generated it? Excellent summary, thank you.
- yoplait_ 10y agoestimating uncertainty. The curve you fit represents what you have in your data. The question is what is out there, in your target population.
- mjw 10y agoPretty much any kind of mathematical modelling that involves uncertainty, really. Making inferences and predictions from data, in the presence of uncertainty. Analysis of the properties of procedures for doing the above. If you want examples that avoid the feel of just "curve fitting" (assume you mean something like "inferring parameters given noisy observations of them") -- maybe look at models involving latent variables. Bayesian statistics has quite a few interesting examples.
- MichailP 10y agoThanks! I had a course at uni named Probability and Statistics, but since it was first (and only) course in EE curriculum it was oriented toward probability, and Statistics was an afterthought (I only remember simple linear and multilinear regression). That is probably the main reason I only see curve fitting everywhere :)
- jeffwass 10y agoAs an EE, how would you explain concepts like a PN junction or field effect transistor without using statistical mechanics? (Ie, expected behaviour for ensembles of huge numbers of particles).
- MichailP 10y agoThe models EE use are simplified, drift and diffusion current and electron and holes with their different mobilities and energy levels. Math apparatus used here, and strictly related to statistics, is limited to averaging, I would dare to say.
- smaddox 10y agoQuantum Monte Carlo simulation is pretty standard for modern semiconductor devices. Most of the effects of interest in highly scaled transistors, for example, cannot be properly accounted for otherwise. On the basic materials level, density functional theory is the current gold standard, and it's extremely statistics heavy. At the systems and architecture levels, you may be right, though.
- jeffwass 10y agoGreat. Those formulae are derived from statistical representation of huge number of particles, eg electrons modelled as a nearly-free gas of fermions obeying Fermi-Dirac statistics. And so too would anything making use of PN junctions, band gaps, especially when considering temperature dependence. I think you can agree now that your original observation of statistics as "glorified curve fitting" as a bit naive.
- theophrastus 10y agoStatistics has, of course, grown from its beginnings as a means to summarize social/population conditions (the median number of serfs per farm, average bushels per acre, etc), yet: "estimating an accurate metric, (for example a 'central tendency'), from incomplete data" remains a central theme.
- hacker42 10y agoStatistics is about inferring probabilities from data such that we can make predictions (where data are discernible differences of some quantities). Inference means finding out what the world is about using some sort of representation (a model). The entire project is basically concerned with (mostly lossy) compression: How to represent the complexity of the world such that we can reason about it with limited resources, i.e. estimate things we can't compute using things we can compute. If our statistic summarizes enough to allow us to make useful predictions, it is called a sufficient statistic. Probability is at the heart of the project: frequencies that summarize reoccurring data. Instead of storing a reoccurring pattern multiple times, we just store it once and record how often it has occurred.
- tlarkworthy 10y agoNeural nets are glorified curve fitting. The are curves parameterized by the weight matrix. The weight matrix is relatively massive (e.g. 1M DOF), which makes the family of curves it generate essentially almost fluid like a piece of yarn. Now given a small amount of data, and a programmable piece of string, how well can you fit the data? Turns out the string is higher dimensional than the data, so you can fit any curve you like. The trick, it avoiding overfitting. Overfitting is the yarn warping its shape to fit noise that has no intrinsic meaning. That's what cross validation prevents ... overfitting. Stop moving the yarn to match the training data better when it fails to improve an independent performance test. Thats what machine learning is... figuring out algorithms that don't overfit and have some ability to generalize onto data not seen before. It's still basically glorified curve fitting.