8 ms·
Statisticians use a technique that leverages randomness to deal with the unknown
- xiaodai 2y agoI don’t know. I find quanta articles very high noise. It’s always hyping something
- jll29 2y agoI don't find the language of the article full of "hype"; they describe the history of different forms of imputation from single to multiple to ML-based. The table is particularly useful as it describes what the article is all about in a way that can stick to students' minds. I'm very grateful for QuantaMagazine for its popular science reporting.
- billfruit 2y agoThe Quanta articles usually have a gossipy style and are very low information density.
- SAI_Peregrinus 2y agoThey're usually more science history than science. Who did what, when, and a basic ovnrview of why it's important.
- vouaobrasil 2y agoI agree with that. I skip the Quanta magazine articles, mainly because the titles seem to be a little to hyped for my taste and don't represent the content as well as they should.
- MiddleMan5 2y agoCurious, what sites would you recommend?
- light_hue_1 2y agoI wish they actually engaged with this issue instead of writing a fluff piece. There are plenty of problems with multiple imputation. Not the least of which is that it's far too easy to do the equivalent of p hacking and get your data to be significant by playing games with how you do the imputation. Garbage in, garbage out. I think all of these methods should be abolished from the curriculum entirely. When I review papers in the ML/AI I automatically reject any paper or dataset that uses imputation. This is all a consequence of the terrible statics used in most fields. Bayesian methods don't need to do this.
- jll29 2y agoThere are plenty of legit. articles that discuss/survey imputation in ML/AI: https://scholar.google.com/scholar?hl=de&as_sdt=0%2C5&q=%22machine+learning%22+imputation&btnG= https://scholar.google.com/scholar?hl=de&as_sdt=0%2C5&q=%22m...
- light_hue_1 2y agoThe prestigious journal "Artificial intelligence in medicine"? No. Just because it's on Google scholar doesn't mean it's worth anything. These are almost all trash. On the first page there's one maybe legit paper in an ok venue as far as ML is concerned (KDD; an adjacent field to ML) that's 30 years old. No. AI/ML folks don't do imputation on our datasets. I cannot think of a single major dataset in vision, nlp, or robotics that does so. Despite missing data being a huge issue in those fields. It's an antiqued method for an antiqued idea of how statistics should work that is doing far more damage than good.
- disgruntledphd2 2y agoOk that's interesting. I profoundly disagree with your tone, but would really like to hear with you regard as good approaches to the problem of missing data (particularly where you have dropout from a study or experiment).
- 2y ago
- clircle 2y agoDoes any living statistician come close to the level of Donald Rubin in terms of research impact? Missing data analysis, causal inference, EM algorithm, any probably more. He just walks around creating new subfields.
- selimthegrim 2y agoEfron?
- richrichie 2y ago& Tibshirani
- selimthegrim 2y agoStein too
- aquafox 2y agoAndrew Gelman?
- nabla9 2y agoGelman has contributed to Bayesianism, hierarchial models and Stan is great, but that's not even close to what Rubin has done. ps. Gelman was Rubin's doctoral student.
- selectionbias 2y agoAlso approximate Bayesian computation, principal stratification, and the Bayesian Bootstrap.
- j7ake 2y agoMike Jordan, Tibshirani, Emmanuel Candes
- paulpauper 2y agowhy not use regression on the existing entries to infer what the missing ones should be?
- ivansavz 2y agoThat would push things towards the mean... not necessarily a bad thing, but presumably later steps of the analysis will be pooling/averaging data together so not that useful. A more interesting approach, let's call it OPTION2, would be to sample from the predictive distribution of a regression (regression mean + noise), which would result in more variability in the imputations, although random so might not what you want. The multiple imputation approach seems to be a resampling methods of obtaining OPTION2, w/o need to assume linear regression model.
- stdbrouw 2y agoMultiple imputation simply means you impute multiple times and run the analysis on each complete (imputed) dataset so you can incorporate the uncertainty that comes from guessing at missing values into your final confidence intervals and such. How you actually do the imputation will depend on the type of variable, the amount of missingness etc. A draw from the predictive distribution of a linear model of other variables without missing data is definitely a common method, but in a state-of-the-art multiple imputation package like mi in R you can choose from dozens.
- Jun8 2y agoNot one mention of the EM algorithm, which is, as far as I can understand, is being described here (https://en.m.wikipedia.org/wiki/Expectation%E2%80%93maximization_algorithm https://en.m.wikipedia.org/wiki/Expectation%E2%80%93maximiza...). It has so many applications, among which is estimating number of clusters for a Gaussian mixture model. An ELI5 intro: https://abidlabs.github.io/EM-Algorithm/ https://abidlabs.github.io/EM-Algorithm/
- Sniffnoy 2y agoIt does not appear to be what's being described here? Could you perhaps expand on the equivalence between the two if it is?
- miki123211 2y ago> It has so many applications, among which is estimating number of clusters for a Gaussian mixture model Any sources for that? As far as I remember, EM is used to calculate actual cluster parameters (means, covariances etc), but I'm not aware of any usage to estimate what number of clusters works best. Source: I've implemented EM for GMMs for a college assignment once, but I'm a bit hazy on the details.
- fleischhauf 2y agoyou are right you still need the number of clusters
- BrokrnAlgorithm 2y agoI've been out of the loop for stats for a while, but is there a viable approach for estimating ex ante the number of clusters when creating a GMM? I can think if constructing ex post metrics, i.e using a grid and goodness of fit measurements, but these feel more like brute forcing it
- disgruntledphd2 2y agoUnsupervised learning is hard, and the pick K problem is probably the hardest part. For PCA or factor analysis, there's lots of ways but without some way of determining ground truth it's difficult to know if you've done a good job.
- TaurenHunter 2y agoDonald Rubin is kind of a modern day Leibniz... Rubin Causal Model Propensity Score Matching Contributions to Bayesian Inference Missing data mechanisms Survey sampling Causal inference in observations Multiple comparisons and hypothesis testing
- a-dub 2y agolife is nothing but shaped noise
- userbinator 2y agoIt reminds me somewhat of dithering in signal processing.
- bgnn 2y agoWhy not interpolate the missing data points with similar patients data? This must be about the confidence of the approach. Maybe interpolation would be overconfident too.
- SillyUsername 2y agoIsn't this just Monte Carlo, or did I miss something?
- hatmatrix 2y agoMonte Carlo is one way to implement multiple imputation.
- karaterobot 2y agoDoes anyone else find it maddeningly difficult to read Quanta articles on desktop, because the nav bar keeps dancing around the screen? One of my least favorite web design things is the "let's move the bar up and down the screen depending on what direction he's scrolling, that'll really mess with him." I promise I can find the nav bar on my own when I need it.