7 ms·
How to Tell Good Studies from Bad? Bet on Them
- btilly 11y agoI really like the idea at the end of using prediction markets to figure out which studies should be challenged by attempting replication.
- nonbel 11y agoReplication is not supposed to be a "challenge", that is not the point. The first point is ensure you can communicate the methods effectively, this demonstrates the experimental conditions are understood. The other point is to see if the observation is stable to whatever differences arise due to varying time/location. And they should be looking at estimates of effect size, not whether p<0.05. Attempting to replace the role of independent replication with opinion is an awful idea. And that is the goal here, to replace, not to supplement. Of course, it doesn't seem replication attempts are very common in this area to begin with. So this is actually going to be a justification for continuing that pseudoscientific practice and to avoid checking all the previous results.
- yummyfajitas 11y agoIn a Bayesian framework, replication is pretty easy to make predictions on. Suppose there is a parameter - say the probability q of a Bernoulli random variable being true. Based on your past experiments you have a posterior p(q). Then given p(q), you can easily compute a probability distribution on S(N), where S(N) is the number of successes you'd get from a bernoulli random variable with N attempts. In code terms, you are just computing posterior.flatMap(q => Bernoulli(N,q)). (Using the inherent monadic structure of probability.) This actually works in general. If you want to predict the outcome of a later experiment (the replication) you just compute posterior.flatMap(parameter => generateResult(parameter)).
- cba9 11y agoOr more concretely, since these projects are typically running on a hypothesis-testing paradigm, you can do a hybrid power analysis: compute a fully Bayesian analysis of the original experiment using other experiments and informative priors which take into account the true distribution of effects in a particular subfield (eg a broad distribution with many large effects if it's related to IQ, or a narrow distribution around zero if it's related to things like priming or stereotype thread) to generate a posterior distribution of effect sizes. Then you can simulate the results of future experiments with _n_ datapoints: sample 1 effect size from that distribution, generate _n_ datapoints assuming that effect size, run the hypothesis-testing, and return the _p_-value. The fraction of _p_<=0.05 is your best forecast of whether the future experiment will succeed in reproducing it or not.
- btilly 11y agoNo, the main purpose of attempting replication is to challenge of the result. "I think you made a mistake, let's try to replicate and show it." And if you fail to replicate, you publish that and show the world that the result shouldn't be believed. If you successfully replicate, then it has survived the challenge. The more challenges it survives, the more confidence we can have that the result is valid. That said, we do not try to make replication challenging. We try to present experiments in a way that makes replication as easy as possible to perform. Exactly so that people don't have to take what we said on faith. The point of the prediction market here is not to replace independent replication with opinion. It is to ensure that the energy that gets spent on replication is more likely to be spent effectively. Ideally, of course, targeting replication efforts more effectively will increase the value of time spent in attempting replication. This should therefore increase how much effort is spent on replication. Which is exactly the opposite of replacing independent replication with opinion!
- nonbel 11y ago> "if you fail to replicate, you publish that and show the world that the result shouldn't be believed." A lone report of some observation should never be believed. It should be verified by others who retrace the steps. >"If you successfully replicate, then it has survived the challenge. The more challenges it survives, the more confidence we can have that the result is valid." If the observation can be independently replicated, it shows that the methods required are understood well enough to communicate and that it is stable in the face of unknown influences. This increases our confidence that we understand what is going on and that the phenomenon is worth theorizing about. It has nothing to do with an observation being true or valid. The observation was made, it happened. It is true. It is valid. (Excepting outright fraud, which is treated equivalently to some severely unstable phenomenon) >"The point of the prediction market here is not to replace independent replication with opinion." They explicitly say that is the goal in the paper. I like this method of eliciting priors, not that the goal is to substitute it for actual replications. >"It is to ensure that the energy that gets spent on replication is more likely to be spent effectively." If a study isn't worth replicating, then it wasn't worth doing and reporting in the first place.
- Fomite 11y ago> And they should be looking at estimates of effect size This. Null hypothesis testing: Just stop it. We've known this is a bad idea for decades
- jimrandomh 11y agoThis is a good answer from the incentives angle - how to motivate people to check whether studies are good or bad. On the object level, the answer is surprisingly simple: actually read the thing. The press is full of stories where a journalist rephrased another journalist's story about a press release about a study, and when you actually go to the study, it says something subtly different. When I see those, I try to jump out of the journalist's summary and get to a PDF as fast as possible, guess what the caveat is going to be, then check for it.
- pavpanchekha 11y agoThere are two stories here. One is, how do you—I'm guessing you're not a scientist—tell which studies are worth a damn. In this case, yes, reading the study is the best thing to do. The second is, how do scientists know which studies are good? This is a harder job than for you¸ because you're only ever made aware of papers that made it through the publication gauntlet. (No, they won't publish "anything" these days, not in a journal that gets any coverage.) For scientists, the task is harder, so prediction markets might be the tool they need.
- mistermann 11y agoI've often though there should be a similar mechanism for solving disagreements in the workplace.
- Rmilb 11y agoThe CIA(1), and other Fortune 500 Companies have internal prediction markets that work quite well. However, internal politics inside of the organization sometimes spell the death of those markets. This is more good data that gives me more faith in the www.Augur.net decentralized prediction market premise. [1] https://www.cia.gov/library/center-for-the-study-of-intelligence/csi-publications/csi-studies/studies/vol50no4/using-prediction-markets-to-enhance-us-intelligence-capabilities.html https://www.cia.gov/library/center-for-the-study-of-intellig...
- ultramancool 11y agoIs there any software to create a simple intranet prediction market?
- adam 11y agoWe make prediction market software for companies to use internally: https://cultivatelabs.com/forecasts https://cultivatelabs.com/forecasts if you want to chat sometime.
- mistermann 11y agoVery interesting....since you're in the business, I wonder if you could share any observations on any difficulty selling into organizations where politically competent people are very much not interested in discovering and publicizing the ability of others in the organization to make correct predictions?
- adam 11y agoWe often have people participate anonymously - either anonymously among their peers, or in certain instances, anonymous from anyone in the company and we serve as the 3rd party arbiter of all the data. It really just depends on the organization's culture and how transparent they are ready to be. Minimally though we're looking to establish an ongoing dialogue between the different levels of an organization. Our belief is that people on the ground building product, interfacing with clients, etc. aren't consulted nearly enough about predictions that inform big strategic decisions. Instead, leaders are making decisions based on input from a limited number of SME's, data analytics, and their own beliefs. None of these are bad per se, but not leveraging your own people we believe is a huge opportunity lost. Happy to follow up live/over email if you'd like. adam at cultivatelabs
- sharp11 11y agoThe problem with this is that it seems likely to be biased against unexpected results or results that contradict the dominant theory. The old saying, "Science advances one funeral at a time," has a lot of truth in it.
- mrjaeger 11y agoI'm sure there is some natural human bias against something totally unexpected being correct, but these people also read the papers. I imagine people would be more concerned with methodological flaws and the like. Also if many unexpected results continue to be verified, the market should shift to properly value those types of studies.
- michael_fine 11y agoThat's a feature, not a bug, I believe. Results that contradict the dominant theory are in general much less likely than one's that confirm it or extend it, and as such will have much lower odds. As a result, if you have significant evidence that such a result is true, then there is a large upside in betting on it, enough to outweigh the unlikelihood.
- sharp11 11y agoThat's interesting. You're looking at it from the point of view of the bettor. My concern is just that non-dominant lines of inquiry are already intensely selected against and this just institutionalizes that bias. But I see your point that the lopsided upside incentivizes contrarians. That's very interesting .. I have to think about that! [Edited as enlightenment dawns ...]
- savanaly 11y agoThat's as it should be, though. Extraordinary claims require extraordinary evidence, and studies should be regarded more skeptically the more surprising their results are.
- yummyfajitas 11y ago
- masonhipp 11y ago"The beauty of the market is that we allow people to be Bayesian" [...] "People come in with some prior belief, but they can also follow prices to see what other people believe and may update their beliefs accordingly [...] participants in the market could focus their bets on the studies they felt most sure of, and as a result, rough guesses didn’t skew the averages as much." It certainly isn't a fool-proof method of increasing accuracy, and it does favor popularity of a theory over other factors, but overall it's probably a nice layer of data to consider adding to the mix.
- smt88 11y agoIt isn't fool-proof, but there is a lot of research into the phenomenon that groups of humans are pretty good at predicting outcomes (much better than most individuals). I forget the math behind it, but it makes a lot of sense mathematically. Here's a book all about it: http://www.amazon.com/The-Wisdom-Crowds-James-Surowiecki/dp/0385721706 http://www.amazon.com/The-Wisdom-Crowds-James-Surowiecki/dp/...
- masonhipp 11y agoVery true. There's another one floating around somewhere about how good we are at estimating the IQ of other people. Pretty interesting.
- cba9 11y ago> "The beauty of the market is that we allow people to be Bayesian" Yes, this is the critical piece. The results of the Reproducibility Project were not remotely a surprise to Bayesian observers. People like Gelman have been pointing out for ages (and I mean back to the 1960s) that the prior probabilities in these fields is low and necessarily a lot of the results were false positives. With the rise of meta-analyses, it is possible to have informative priors for particular fields of psychology or for psychology as a whole, which would let you make much better predictions about whether a result was real. But you can't use these in papers - authors are heavily biased towards using procedures or flat priors which are uninterpretable or grossly overestimate the evidence, and if you try to use any of the informative priors or more advanced models, they'll nag you to death with a thousand objections and complain about double standards and subjectivity and how this time is different and (ironically) bias. So for the most part, there's not much to gain in academic research. But in a prediction market, you don't have to listen to the self-serving excuses or explain your reasoning, and there's something to make it worth your while.
- qznc 11y agoPrediction markets are great tools in general. Unfortunately, incentives are usually against implementing them. Experts are easier to control.
- nemo1618 11y agoThere are a few nascent cryptocurrency-based prediction markets on the horizon, Augur being the most well-known. If one of them takes off, it could influence our economy and society in a big way.
- jerryhuang100 11y agoisn't that just how options or event prediction exchange / markets work?
- evmar 11y agoIt is funny they were worried about whether they just got lucky with their result, then did the prediction market thing, and then didn't worry whether they just got lucky with that result! (At least the article didn't, perhaps the researchers did.) So here are some amateur stats, please check my work. This article says that the prediction market correctly predicted 71% of the replication results of 44 studies, or 31 correct. Assume the studies have a 50% chance of being replicable. Then a random coin would predict a mean of 22 correct with a std dev of sqrt(0.5 * 0.5 * 44) = 3.3. This sample has a z score of 2.72, which means there's a probability of 0.003264 (0.3%) of the random chance approach being correct 71% or better. So the result seems pretty significant. (Changing the assumed 50% to other values makes the probability even more extreme.)
- socrates1998 11y agoThat's exactly what I was thinking. Is their study of "using betting to accurately predict whether studies are reproducible" also reproducible? A very interesting question about stats and studies. Even the p-value of 0.01 seems SUPER strong, but in reality, falls apart in many cases.
- Vraxx 11y agoWouldn't changing the assumed rate of 50% to something higher make the probability less extreme? I know my initial inclination is to believe that less than 50% of the published studies have results that are able to be replicated, but that could be just the cynic in me. I know that the article mentioned that just 39% of the selected studies were able to replicate the results, but that could be caused by a selection bias in which 44 studies were used. I'm inclined to believe this study, especially because the results are rather amusing, but there is definitely room to be more sure.
- evmar 11y agoThank you for thinking critically about my comment! As I mentioned I am an amateur at stats and I now think I modeled the problem wrong. Let's hypothetically suppose that we knew that 75% of the studies were replicable. Can we make a better coin flip prediction? If you had a coin flip that says yes 75% of the time, it isn't necessarily correct at a rate of 75% -- instead it'd be right .75^2+.25^2 = 62.5% of the time. In fact a coin that just predicted "replicable" every time in this scenario would be right 75% of the time. So I think maybe my null hypothesis should've been based on "did they do better than a parrot that always says yes", not a coin flip. I think the math in the original problem stays the same, it's just you change it to "a coin that always predicts it will be replicable". And in that case, if the underlying rate of replicable was 71%, then their prediction market only does as well as the always-yes coin and is in fact not very useful.
- benp84 11y agoSo in other words, a bunch of people guessing which hypotheses were true was more accurate than actual scientific studies of them (71% vs 39%). Great.
- jeffdavis 11y agoThe cost of a contract doesn't represent whether the result is reproducible or not, it predicts the probability that it's reproducible. So what do they mean when they say it correctly predicted the outcome? Are they just saying the odds fell on the same side as the reproduction indicated? If so, that seems arbitrary. If the cutoff for a p-value is 0.05, then shouldn't we say that any contract selling for less than $95 predicts a reproduction failure?
- Houshalter 11y ago>With a p-value [of 0.01], the result hardly screamed “false positive” like a barely significant one of, say, 0.05 might. Is 0.01 that low for such a crazy finding? Let's say you believe that it has a probability of 1 in 10,000. And that result really seemed really really unlikely. 1 in 10,000 might be generous. Then, after this study, the probability that it's true is 1%.