10 ms·
Ideas in statistics that have powered AI
- hyttioaoa 5y ago"Generalized adversarial networks, or GANs, are a conceptual advance that allow reinforcement learning problems to be solved automatically." - "Generalized" :D Also the description is nonsense. This has nothing to do with reinforcement learning. Makes me wonder about the rest.
- deleted 5y ago[deleted]
- cscurmudgeon 5y agoNow, if a press release from a top univ is so wrong on something that is easily checkable, how accurate are other forms of news?
- nerdponx 5y agoThink of "press release wrongness" with a probability distribution. Some press releases are really good, some are really bad. A sensible prior would be somewhere in the middle. If you start to see a lot of bad press releases, then you can update your posterior towards "I can't trust any of these."
- bee_rider 5y agoIs that a good prior? I expect due to Dunning–Kruger that the willingness to produce an article on a topic would follow a pretty intense bimodal distribution.
- deleted 5y ago[deleted]
- totoglazer 5y agoThe paper has it right, at least.
- vcdimension 5y agoI'm surprised they didn't mention support vector machines and the kernel trick which was discovered by statisticians.
- srean 5y agoAlthough Vapnik's treatise is called Statistical Learning Theory, neither statisticians nor he himself identifies himself as a statistician. In fact his proposals were quite radically different from the established norm in contemporary statistics. The same holds for Corrina Cortes. Kernel 'trick', representer theorem etc are far older and have their origins in functional analysis
- JHonaker 5y agoI highly doubt that a person with a PhD in statistics doesn't identify as a statistician.
- srean 5y agoThe only way to dispute that would be appeal to authority so I will avoid. Perhaps there was a time when he did identify as a mathematical/theoretical statistician, but his contributions were quite a dramatic break from what was the norm in statistical practice at the time. I would argue that his contributions were central to the birth of rigorous machine learning (the non deep learning kind) as a field of its own with a focus that's different from that of statistics. One quantitative test I can suggest -- one can count the number of occurrences of the word 'statistics' in the journals and conferences he has published and compare that with the number of publications he has authored. My sense is that it will be close to 0 and getting closer (if not already there yet).
- dkshdkjshdk 5y ago> One quantitative test I can suggest -- one can count the number of occurrences of the word 'statistics' in the journals and conferences he has published and compare that with the number of publications he has authored. My sense is that it will be close to 0 and getting closer (if not already there yet). I'm not sure this "quantitative test" is the best approach... after all, I'm pretty sure "Biometrica" and "Econometrica" don't have "statistics" in their names.
- sgt101 5y agoHow have they attributed GANs and Deep Learning to Statistics? I thought Goodfellow was doing an AI PhD and that Hinton is a biologically inspired / neuroscience fellow?
- MAXPOOL 5y agoDeep learning machine learning models are statistical and probabilistic models. You can categorize Deep Learning under both computer science and statistics. For example stat.ML and cs.LG in Arxiv. Machine learning and statistics are closely related fields, both historically and in current practice and methodology.
- sgt101 5y agoMy working definition is that statisticians choose and engineer models while machine learning searches a vast space of models.
- nightski 5y agoThat doesn't seem right to me. In both cases you have a model and are just searching for optimal parameters considering the bias/variance tradeoff. There may be a few instances of a neural network or other ML model being set up to dynamically change it's architecture during training but that seems to me like it would not work out well at all. If anything in specific cases the statistical model (if Bayesian) is more comprehensive in that it doesn't try to find a point estimate of the parameters but instead forms a full distribution around the plausibility of the parameters.
- sjg007 5y agoYou can have Bayesian DNNs.
- whimsicalism 5y agoNot easily. I don't think the literature is very incredible on this - how do you define a prior over all of the parameters of an NN?
- bmc7505 5y agohttps://statmodeling.stat.columbia.edu/2020/12/09/what-are-the-most-important-statistical-ideas-of-the-past-50-years/ https://statmodeling.stat.columbia.edu/2020/12/09/what-are-t...
- heinrichhartman 5y agoOut of the 10 papers I am able to download 3 of them freely. - For the papers I am quoted 26EUR - 39EUR - For the books I am quoted 129EUR - 133EUR This is audacious. Some of these papers are form the 70ies. And I highly doubt that the authors get any royalties from those sales.
- the_svd_doctor 5y agoAuthors never get _any_ royalties from paper sales as far as I know :) (for books maybe).
- shakow 5y ago> (for books maybe). We do. I don't know if it's the general rule, but for the one I partook in, we get ~20 €cents.sale-1.author-1.
- nolroz 5y ago<donates to sci-hub>
- ur-whale 5y agosci-hub FTW why would you want to feed the parasites?
- gwern 5y agoI'm not sure what you mean. I went through the list myself, and while the books are obviously only on Libgen, the only one I didn't find readily available in Google Scholar was the AIC paper (https://www.gwern.net/docs/statistics/decision/1998-akaike.pdf https://www.gwern.net/docs/statistics/decision/1998-akaike.p...), and you can safely assume any paper in GS is in SH too.
- bjornsing 5y agoI’m sorely missing Maximum Likelihood Estimation (MLE). It’s a statistical technique that goes back to Gauss and Laplace but was popularized by Fisher. In AI/ML it’s often referred to as “minimizing cross-entropy loss”, but this is just a misappropriation / reinvention of the wheel. The math is the same and MLE is a much more sane theoretical framework.
- MontyCarloHall 5y ago“Cross entropy” specifically refers to the log-likelihood function of a binary random variable, and is only used as the cost function for binary classifiers. It does not refer to likelihood functions in general.
- ansk 5y agoDo people not google terms before trying to speak authoritatively on a topic they aren't familiar with? The original commenter is correct, cross entropy is a generic measure of two probability distributions - in the case of maximum likelihood estimation, these are the data distribution and the distribution of the learned model.
- MontyCarloHall 5y agoYou are incorrect. For a given probability distribution parameterized by θ with probability mass/density p(x|θ), the likelihood of θ given a set of data X = {x_1, …, x_n} (assuming X is independently/identically distributed) is simply the product of independent probabilities, L(θ|X) = Π_i=1^n p(x_i|θ) Maximizing this product with respect to θ yields the maximum likelihood estimate of θ. Since sums are generally easier to work with than products, and log is a monotonic function, we generally work with the log-likelihood function log L(θ|X) = Σ_i=1^n log p(x_i|θ) since the log-likelihood will achieve its maximum for the same value of θ as the likelihood. The cross entropy of two discrete probability distributions p and q is Σ_i=1^n p_i log q_i (For continuous distributions, replace the sum with an integral.) This is completely unrelated to the generic log-likelihood function defined above. The two are only related if p happens to be the probability distribution of a binary random variable x = {0,1}, with probability π of equalling 1: p(x|π) = π^x(1-π)^(1-x) Its log-likelihood is therefore x log π + (1-x) log(1-π) which for this particular case, happens to be a cross entropy. Note that this is the log-likelihood of a single observation in a single class; for multiple observations/multiple classes, we sum across them, e.g. Σ_i=1^n x_i log π_i + (1-x_i)log (1-π_i) for a single observation across n total classes. But again, the relationship to cross entropy only holds for this particular choice of p. It is not generally the case that the generic log-likelihood function, log L(θ|X) = Σ_i=1^n log p(x_i|θ) is a cross entropy!
- andyxor 5y ago..aand on the same page there is a link promoting Critical Race Theory. It’s kind of like coming to listen to math lecture and seeing swastika signs on the walls.
- ehw3 5y ago> 2. John Tukey (1977). Exploratory Data Analysis. > This book has been hugely influential and is a fun read that can be digested in one sitting. Wow. The PDF is over 700 pages. That seems fairly impressive for single-sitting digestion.
- 317070 5y ago> Generative adversarial networks, or GANs, are a conceptual advance that allow reinforcement learning problems to be solved automatically. They mark a step toward the longstanding goal of artificial general intelligence while also harnessing the power of parallel processing so that a program can train itself by playing millions of games against itself. At a conceptual level, GANs link prediction with generative models. What? Every sentence here is so wrong I have a hard time seeing what kind of misunderstanding would lead to this. GAN's are a conceptual advance of generative models (i.e. models that can generate more, similar data). Reinforcement learning is a separate field. Parallel processing is ubiquitous, and has nothing to do with GANs or reinforcement learning (they are both usually pretty parallellized). Self-play sounds like they wanted to talk about the alphago/alphazero papers? And GANs are infamously not really predictive/discriminative. If anything, they thoroughly disconnected predicition from generative models.
- deleted 5y ago[deleted]
- whimsicalism 5y ago> GAN's are a conceptual advance of generative models (i.e. models that can generate more, similar data). This is something I've long had confusion with, coming from a probabilistic perspective. How does a GAN model the joint probability of the data? My understanding was that was what a generative model does. There doesn't seem to be a clear probabilistic interpretation of a GAN whatsoever.
- gyom 5y agoPart of the cleverness of GANs was to have found a way to train a neural network that generates data without explicitly modeling the probability density. In a stats textbook, when you know that your training data comes from a normal distribution, you can maximize the MLE wrt the parameters, and then use that for sampling. That's basic theory. In practice, it was very hard to learn a good pdf for experimental data when you had a training set of images. GANs provided a way to bypass this. Of course, people could have said "hey let's generate samples without maximizing a loglikelihood first", but they didn't know how to do it properly, how to train the network in any other way besides minimizing cross-entropy (which is equivalent to maximizing loglikelihood). Then GANs actually provided a new loss function that could be trained. Total paradigm shift!
- master_yoda_1 5y agohalf of these are relevant to small data problem which is not exactly we mean when we say AI.
- dkshdkjshdk 5y agoWhat do you mean when you say AI? I'm curious. As far as I can tell, most people (e.g., whoever wrote this article) seem to use AI as a synonym for "machine learning", basically.
- master_yoda_1 5y agoAI is very old term. But "new AI" is all based on big data. Most of the small data algorithm is always been used and still being used in medicine etc. But when you talk about "new AI" you mean deep learning which is very data hungry and except last 2 none of the algorithm has anything to do with deep leaning.
- deleted 5y ago[deleted]
- sjg007 5y ago<sarcasm> Psssh.. it's all math. </scarcasm>