8 ms·
Do not confuse a random variable with its distribution
- baking 2y agoWhy make it complicated? One coin flip, X = Heads and Y = Tails. P(X = Y) = 0.
- magnio 2y agoX and Y in your examples are events, not random variables.
- benkuykendall 2y agoSure, but you can associate events with indicator rvs: X = 1[Heads], Y = 1[Tails]
- mturmon 2y agoDefinitely. Taking this to its logical end leads to a beautiful concise notation that I learned from David Pollard: - You literally identify sets with their indicators: they are the same. - You identify the “P” operator as expectation (integration) with respect to the underlying measure, and next… - You note that integration is linear, so you use linear operator notation everywhere you’d use P. So if “A” is a set, you just write P A = 0.5 This is equivalent to: P A = P 1[ω ∈ A] = ∫ 1[ω ∈ A] dP(ω) in Lebesgue notation. There’s an example on page 2 of http://www.stat.yale.edu/~pollard/Courses/241.fall2014/notes2014/Probability.pdf http://www.stat.yale.edu/~pollard/Courses/241.fall2014/notes... although he’s not using measure theory there. It can be really clean and terse especially when doing bounds for random variables.
- dfeng 2y agoWow, I did not expect to see David's notation here on HN. The only problem with the notation is that it becomes so second nature that you forget it's not standard!
- knappa 2y agoA lot is lost here by using notation that looks like it is rigorous math, but is actually pretty vague. For example, are X and Y indicators for the same flip? If so, they are mutually exclusive, X=Y is contradictory, and hence P(X=Y)=0. If they are samples from different flips (and your coin is the usual idealized one) then X and Y are independent random variables and P(X=Y)=0.25. It's just like if X~N(0,1), Y~N(0,1) and you want to know the distribution of X-Y. You need to know what the PDF of (X,Y) looks like. Well, you don't know. X and Y could be correlated or they might not be. e.g. if could be that (X,Y)~N( (0,0), [(1,0),(0,1)] ) or maybe (X,Y)~N( (0,0), [(1,1/2),(1/2,1)] ). The distribution of X-Y cares how correlated X and Y are.
- setopt 2y ago> If they are samples from different flips (and your coin is the usual idealized one) then X and Y are independent random variables and P(X=Y)=0.25. Isn’t it 0.5?
- MathMonkeyMan 2y agofeels more distributiony if 2n + 1 > 1
- rottencupcakes 2y ago[flagged]
- benkuykendall 2y agoI feel like I'm missing some context here?
- condwanaland 2y agoLove to see things built with bookdown, which is such an awesome R package (although it's successor, Quarto, is much better and simpler)
- clircle 2y agoWhoops, you posted the wrong page. The statistics page that Hacker News needs to read is the one about how the Central Limit Theorem doesn't apply to everything damn thing.
- deleted 2y ago[deleted]
- seanhunter 2y ago...with a sidenote about how no the CLT doesn't actually mean that if you take lots of samples of something the distribution of those samples is Gaussian.
- kqr 2y agoYou mean this one? https://two-wrongs.com/it-takes-long-to-become-gaussian.html https://two-wrongs.com/it-takes-long-to-become-gaussian.html
- 082349872349872 2y agoIn particular for software, instead of gaussians we often have the sort of distribution of completion times where "if the expected completion time is T1, but empirical observation says it never actually got done in between (0,T1], the conditional expected completion time is T2, and T2>>T1": ie, the longer you work on something without success, the further away into the indefinite future the horizon of expected success recedes...
- dwqwdqd 2y agoDoes this work? X = 1 with probability 0.5, 0 with probability 0.5 Y = 0 when X = 1, 1 when X = 0 (for the \omega for which X(\omega) = 1, Y(\omega) = 0). They're both bernoulli distributions with p=0.5 (i.e. they follow the same distribution) and P(X=Y) = 0
- lmm 2y agoYes. The continuous case is more interesting.
- Davidzheng 2y agoHonestly it's fine to confuse a random variable with its distribution if you only are working with a single RV. Changing probability space without changing distribution doesn't really matter much, probability space is more of an abstraction it's not really measurable
- eru 2y ago> A random variable measures a numerical quantity which depends on the outcome of a random phenomenon. Hmm, that sentence at the beginning is already wrong. Random variables can measure anything, not just numbers. Heads or Tails of a coin, or colours of cars etc. It's fine to restrict yourself to numeric random variables only. But if you are writing a rant telling other people to be more careful in their analysis, you better dot your i's and cross your t's yourself.
- zwaps 2y agoI think they mean that a random variable maps to a real number space. If it maps elsewhere, mathematicians like to call it a random element instead. https://en.m.wikipedia.org/wiki/Random_element https://en.m.wikipedia.org/wiki/Random_element
- eru 2y agoInteresting! When I studied math (in Germany) we used the German equivalent of 'random variable' to describe the more generalised concept that English seems to call 'random element'.
- defrost 2y agoIndeed, I suspect those two wikipedia articles ( Random variable | Random element ) have been captured by a particular school of thought, I studied post grad math in Australia and interacted with many mathematicians from a number of backgrounds, all appeared fine with treating (say) a random unit vector ( or point on the surface a sphere ) as a Random variable. I can understand why some might make a cut between pure numbers and other objects, but it's not something that troubles many.
- eru 2y ago> I can understand why some might make a cut between pure numbers and other objects, but it's not something that troubles many. I can see why you would teach the more restricted definition eg in high school.
- jhrmnn 2y agoIn quantum mechanics, the measurement and observation are two sides of the same coin, and the sample space is _defined_ by the random variable (observable) of interest, so it makes a little less sense to separate the two. (There is no hidden observation-independent sample space.)
- trueismywork 2y agoIt still makes sense to separate them when thinking about theory.
- elbear 2y agoCan you explain why? Because I still don't get the point the article is trying to make. To clarify what I do understand: so you have the variable, like height, and all its possible values along with their probabilities (that's the distribution, if I understand things right). The distribution represents a big part of what the variable is, although I realise the variable maybe has other attributes too (none comes to mind right now though).
- jhrmnn 2y agoThe article makes the point that the random variable map and the underlying sample distributions can change independently. That’s just not the case in quantum mechanics.
- glitchc 2y agoThis article is overly complicated. The random variable X is a function mapping the outcome to its probability. The distribution, or the probability density function, or pdf, is the integral of that function. The cumulative density function, the cdf, is in turn the integral of the pdf.
- deleted 2y ago[deleted]
- blt 2y ago> The random variable X is a function mapping the outcome to its probability This is simply wrong. Random variables are not defined this way.
- defrost 2y agoIt would be more constructive if you were to provide an example definition that you consider correct and perhaps even a link to where it is defined and used as you have yet to actually say.
- blt 2y agoThe standard definition is the one given in the two bullet points at the top of the article we are discussing. Or, in more detail, at https://en.wikipedia.org/wiki/Random_variable#Definition https://en.wikipedia.org/wiki/Random_variable#Definition.
- defrost 2y agoglitchc stated: > The random variable X is a function mapping the outcome to its probability. you've stated: "The standard definition is the one given in the two bullet points at the top of the article we are discussing" ie: > The random variable X itself, that is, the function which maps sample space outcomes to numbers. you've also stated: "Or, in more detail, at (wikipedia link)" which has: > A random variable is a measurable function from a sample space as a set of possible outcomes to a measurable space These all appear to be in rough alignment .. all three agree upon a function mapping from outcomes to measure. You've described the first as "This is simply wrong. Random variables are not defined this way." Perhaps you can expand on why this is so wrong compared to the other two definitions.
- dinobones 2y agoThese types of explanations are the reason I dislike school. This is such a stuffy and contrived way to explain things. I’m so glad I have ChatGPT now, I always ask for applied examples and ask it to explain things intuitively. I would’ve been a 4.0 student if I would’ve had ChatGPT as my personal tutor when I was in school.
- auraai 2y agoThis is pretty important in mathematical finance, where one moves from a real-world measure to a risk-neutral measure to make computations feasible. https://en.wikipedia.org/wiki/Girsanov_theorem https://en.wikipedia.org/wiki/Girsanov_theorem https://en.wikipedia.org/wiki/Risk-neutral_measure https://en.wikipedia.org/wiki/Risk-neutral_measure
- panic 2y agoThere’s an interesting connection here to another article on the front page: https://news.ycombinator.com/item?id=40794786 https://news.ycombinator.com/item?id=40794786 In that article, squaring a number in interval arithmetic is different from multiplying two independent numbers with the same interval. Here, squaring a random variable is different from multiplying two independent random variables with the same distribution.
- kazinator 2y agoWho confuses a random variable with its distribution, and what does that mistake look like? I don't get it.
- btown 2y agoFor the code-minded out there, a "random variable" is something of a lazily evaluated value that can be "sampled" and emit a quantity (or a vector/tensor thereof) each time. And the OP article boils down to the fact that it's generally incorrect to assume that any random variable can be represented solely by its unconditional probability distribution; a distribution is more of a visualization than a sufficient definition. Rather, one must track the entire graph of other random variables that may feed the current one (e.g. that the current one is conditional on), akin to how an Excel spreadsheet models all the dependencies of a cell. The fun part comes when you can ask this computation graph: "what parameters for a random variable early on in the chain would be the ones that optimize some function of variables later in the chain?" And, handwaving a ton of nuance here, when those parameters are weights in a neural network, the function is a loss function on the training data, and the optimization is done by automatic differentiation (e.g. https://pytorch.org/tutorials/beginner/introyt/autogradyt_tutorial.html https://pytorch.org/tutorials/beginner/introyt/autogradyt_tu...), you have modern AI. If you're interested in the theoretical underpinnings here, Bishop's PRML is perhaps the classic starting point: https://www.microsoft.com/en-us/research/uploads/prod/2006/01/Bishop-Pattern-Recognition-and-Machine-Learning-2006.pdf https://www.microsoft.com/en-us/research/uploads/prod/2006/0...
- tpoacher 2y agoI've often felt that one of the reasons such warnings are even necessary, is because the notation we use to denote probabilities in the first place is atrocious, and clearly an abuse of notation. A better convention would make clear the distinction between the set of possible outcomes, the act of obtaining a (range of) samples from that set, and the probability that those events match a value range of interest. p(x=X) is not enough to capture all that information. let alone p(x) vs p(X).