In probability theory and statistics, the law of the unconscious statistician, or LOTUS, is a theorem which expresses the expected value of a function g(X) of a random variable X in terms of g and the probability distribution of X. The form of the law depends on the type of random variable X in question. If the distribution of X is discrete and one knows its probability mass function pX, then the expected value of g(X) is
E [ g ( X ) ] = ∑ x g ( x ) p X ( x ) , {\displaystyle \operatorname {E} [g(X)]=\sum _{x}g(x)p_{X}(x),\,}
where the sum is over all possible values x of X. If instead the distribution of X is continuous with probability density function fX, then the expected value of g(X) is
E [ g ( X ) ] = ∫ − ∞ ∞ g ( x ) f X ( x ) d x {\displaystyle \operatorname {E} [g(X)]=\int _{-\infty }^{\infty }g(x)f_{X}(x)\,\mathrm {d} x}
Both of these special cases can be expressed in terms of the cumulative probability distribution function FX of X, with the expected value of g(X) now given by the Lebesgue–Stieltjes integral
E [ g ( X ) ] = ∫ − ∞ ∞ g ( x ) d F X ( x ) . {\displaystyle \operatorname {E} [g(X)]=\int _{-\infty }^{\infty }g(x)\,\mathrm {d} F_{X}(x).}
In even greater generality, X could be a random element in any measurable space, in which case the law is given in terms of measure theory and the Lebesgue integral. In this setting, there is no need to restrict the context to probability measures, and the law becomes a general theorem of mathematical analysis on Lebesgue integration relative to a pushforward measure.
Etymology This proposition is (sometimes) known as the law of the unconscious statistician because of a purported tendency to think of the aforementioned law as the very definition of the expected value of a function g(X) and a random variable X, rather than (more formally) as a consequence of the true definition of expected value. The naming is sometimes attributed to Sheldon Ross' 1972 textbook Introduction to Probability Models, although he removed the reference in later editions. Many statistics textbooks do present the result as the definition of expected value.
Joint distributions A similar property holds for joint distributions, or equivalently, for random vectors. For discrete random variables X and Y, a function of two variables g, and joint probability mass function p X , Y ( x , y ) {\displaystyle p_{X,Y}(x,y)} :
E [ g ( X , Y ) ] = ∑ y ∑ x g ( x , y ) p X , Y ( x , y ) {\displaystyle \operatorname {E} [g(X,Y)]=\sum _{y}\sum _{x}g(x,y)p_{X,Y}(x,y)}
In the absolutely continuous case, with f X , Y ( x , y ) {\displaystyle f_{X,Y}(x,y)} being the joint probability density function,
E [ g ( X , Y ) ] = ∫ − ∞ ∞ ∫ − ∞ ∞ g ( x , y ) f X , Y ( x , y ) d x d y {\displaystyle \operatorname {E} [g(X,Y)]=\int _{-\infty }^{\infty }\int _{-\infty }^{\infty }g(x,y)f_{X,Y}(x,y)\,\mathrm {d} x\,\mathrm {d} y}
Special cases A number of special cases are given here. In the simplest case, where the random variable X takes on countably many values (so that its distribution is discrete), the proof is particularly simple, and holds without modification if X is a discrete random vector or even a discrete random element. The case of a continuous random variable is more subtle, since the proof in generality requires subtle forms of the change-of-variables formula for integration. However, in the framework of measure theory, the discrete case generalizes straightforwardly to general (not necessarily discrete) random elements, and the case of a continuous random variable is then a special case by making use of the Radon–Nikodym theorem.
… excerpt ends here. Continue reading the full article.
