In statistics, the Lehmann–Scheffé theorem provides sufficient conditions for the existence of a best unbiased estimator in a statistical model. The theorem states that any unbiased estimator for a quantity that depends on the data only through a complete, sufficient statistic is the unique uniformly minimum-variance unbiased estimator (UMVUE) of that quantity. The Lehmann–Scheffé theorem is named after Erich Leo Lehmann and Henry Scheffé, given their two early papers.
Introduction Given a vector X = ( X 1 , X 2 , … , X n ) {\displaystyle X=(X_{1},X_{2},\dots ,X_{n})} of random samples from a distribution P θ {\displaystyle \mathbb {P} _{\theta }} for some parameter θ ∈ Θ {\displaystyle \theta \in \Theta } , the goal is to establish sufficient conditions for the existence of an UMVU estimator T ∗ ( X ) {\displaystyle T^{\ast }(X)} for some quantity g ( θ ) {\displaystyle g(\theta )} , that is, E θ [ T ∗ ( X ) ] = g ( θ ) {\displaystyle \mathbb {E} _{\theta }[T^{\ast }(X)]=g(\theta )} and for any unbiased estimator S ( X ) {\displaystyle S(X)} it holds
V a r θ ( T ∗ ( X ) ) ≤ V a r θ ( S ( X ) ) , ∀ θ ∈ Θ . {\displaystyle \mathrm {Var} _{\theta }(T^{\ast }(X))\leq \mathrm {Var} _{\theta }(S(X)),\quad \forall \theta \in \Theta .}
The Rao–Blackwell theorem already shows that, given a sufficient statistic T {\displaystyle T} , the estimator T ∗ ( X ) = E [ S ( X ) ∣ T ( X ) ] {\displaystyle T^{\ast }(X)=\mathbb {E} [S(X)\mid T(X)]} has a uniformly smaller variance than S ( X ) {\displaystyle S(X)} , but it does not guarantee that T ∗ {\displaystyle T^{\ast }} is already UMVU. This is where the Lehmann–Scheffé theorem comes in, if T {\displaystyle T} is also complete, that is, for any real-valued measurable function φ {\displaystyle \varphi } it holds
E θ [ φ ( T ( X ) ) ] = 0 ⟹ φ ( T ( X ) ) = 0 P θ -a.s. {\displaystyle \mathbb {E} _{\theta }[\varphi (T(X))]=0\implies \varphi (T(X))=0\quad \mathbb {P} _{\theta }{\text{-a.s.}}}
Statement As above, let X = ( X 1 , X 2 , … , X n ) {\displaystyle X=(X_{1},X_{2},\dots ,X_{n})} be a vector of random samples from a distribution P θ {\displaystyle \mathbb {P} _{\theta }} for some parameter θ ∈ Θ {\displaystyle \theta \in \Theta } and Θ {\displaystyle \Theta } an arbitrary set. Assume that there exists a complete, sufficient statistic T {\displaystyle T} for the family of distributions ( P θ ) θ ∈ Θ {\displaystyle (\mathbb {P} _{\theta })_{\theta \in \Theta }} . Then, the following two equivalent statements hold:
There exists at most one measurable function h {\displaystyle h} such that T ∗ = h ( T ( X ) ) {\displaystyle T^{\ast }=h(T(X))} is unbiased for g ( θ ) {\displaystyle g(\theta )} and E θ [ T ∗ ( X ) 2 ] < ∞ {\displaystyle \mathbb {E} _{\theta }[T^{\ast }(X)^{2}]<\infty } for all θ ∈ Θ {\displaystyle \theta \in \Theta } , in which case T ∗ ( X ) {\displaystyle T^{\ast }(X)} is the unique UMVUE for g ( θ ) {\displaystyle g(\theta )} . For any unbiased estimator S ( X ) {\displaystyle S(X)} , if it exists, with E θ [ S ( X ) 2 ] < ∞ {\displaystyle \mathbb {E} _{\theta }[S(X)^{2}]<\infty } for all θ ∈ Θ {\displaystyle \theta \in \Theta } , the estimator T ∗ ( X ) = E θ [ S ( X ) ∣ T ( X ) ] {\displaystyle T^{\ast }(X)=\mathbb {E} _{\theta }[S(X)\mid T(X)]} is the unique UMVUE for g ( θ ) {\displaystyle g(\theta )} . In fact, the theorem does not state that unbiased estimators exist in the first place. However, if they do, then there exists a unique square-integrable UMVUE. Moreover, the estimator T ∗ ( X ) = E θ [ S ( X ) ∣ T ( X ) ] {\displaystyle T^{\ast }(X)=\mathbb {E} _{\theta }[S(X)\mid T(X)]} does neither depend on θ {\displaystyle \theta } , since T {\displaystyle T} is sufficient, nor on S {\displaystyle S} , since T {\displaystyle T} is also complete.
Proof In the following, the dependence of an estimator on the data X {\displaystyle X} will not be written out explicitly, i.e., we write S {\displaystyle S} instead of S ( X ) {\displaystyle S(X)} . First of all, if there is no unbiased estimator for g ( θ ) {\displaystyle g(\theta )} , then there is obviously no UMVUE, and if all unbiased estimator are not square-integrable, then their variances are infinity and the statement is trivial. Thus, we focus on the case where a square-integrable unbiased estimator exists. Uniqueness of h {\displaystyle h} : Let S = h 1 ( T ) {\displaystyle S=h_{1}(T)} and R = h 2 ( T ) {\displaystyle R=h_{2}(T)} be unbiased estimators of g ( θ ) {\displaystyle g(\theta )} for some measurable functions h 1 {\displaystyle h_{1}} and h 2 {\displaystyle h_{2}} . For the expectation of the difference it holds E θ [ h 1 ( T ) − h 2 ( T ) ] = g ( θ ) − g ( θ ) = 0. {\displaystyle \mathbb {E} _{\theta }[h_{1}(T)-h_{2}(T)]=g(\theta )-g(\theta )=0.} Since T {\displaystyle T} is complete, this implies h 1 − h 2 = 0 {\displaystyle h_{1}-h_{2}=0} . Thus, T ∗ = h 1 ( T ) =: h ( T ) {\displaystyle T^{\ast }=h_{1}(T)=:h(T)} is the unique unbiased estimator that is a function of T {\displaystyle T} .
h ( T ) {\displaystyle h(T)} is the UMVUE: Let S {\displaystyle S} be any square-integrable unbiased estimator and S ∗ = E θ [ S ∣ T ] {\displaystyle S^{\ast }=\mathbb {E} _{\theta }[S\mid T]} . By the factorization lemma there exists measurable function h ~ {\displaystyle {\tilde {h}}} such that S ∗ = h ~ ( T ) {\displaystyle S^{\ast }={\tilde {h}}(T)} and since S ∗ {\displaystyle S^{\ast }} is unbiased as well, by the above, it must hold S ∗ = h ~ ( T ) = h ( T ) = T ∗ {\displaystyle S^{\ast }={\tilde {h}}(T)=h(T)=T^{\ast }} . Thus, by the Rao–Blackwell theorem, it follows V a r θ ( T ∗ ) = V a r θ ( S ∗ ) = V a r θ ( E θ [ S ∣ T ] ) ≤ V a r θ ( S ) , ∀ θ ∈ Θ . {\displaystyle \mathrm {Var} _{\theta }(T^{\ast })=\mathrm {Var} _{\theta }(S^{\ast })=\mathrm {Var} _{\theta }(\mathbb {E} _{\theta }[S\mid T])\leq \mathrm {Var} _{\theta }(S),\quad \forall \theta \in \Theta .} In other words, T ∗ = h ( T ) {\displaystyle T^{\ast }=h(T)} is an UMVUE and according to the first part it is unique.
Application The Lehmann–Scheffé theorem motivates two general methods to construct UMVU estimators for g ( θ ) {\displaystyle g(\theta )} in models which allow for a complete sufficient statistic T {\displaystyle T} . Method 1: Determining the function h {\displaystyle h}
The UMVUE, if it exists, is the (unique) solution of the equation E θ [ h ( T ) ] = g ( θ ) {\displaystyle \mathbb {E} _{\theta }[h(T)]=g(\theta )} for all θ ∈ Θ {\displaystyle \theta \in \Theta } . If h {\displaystyle h} is, for example, a linear function, then solving this equation is fairly easy. Method 2: Conditioning on an unbiased estimator First, it suffices to find any unbiased estimator S {\displaystyle S} of g ( θ ) {\displaystyle g(\theta )} , which is often easily feasible. The UMVUE can then be determined by evaluating the condition expectation E θ [ S ∣ T ] {\displaystyle \mathbb {E} _{\theta }[S\mid T]} . Since the choice of S {\displaystyle S} is arbitrary, it is preferable to choose it such that the conditional expectation is as simple as possible.
Examples
Bernoulli distribution Let X 1 , … , X n {\displaystyle X_{1},\dots ,X_{n}} be Bernoulli-distributed with probability p ∈ ( 0 , 1 ) {\displaystyle p\in (0,1)} . The joint probability mass function f p {\displaystyle f_{p}} is of the form f p ( x 1 , … , x n ) = ∏ i = 1 n p x i ( 1 − p ) 1 − x i = ( p 1 − p ) ∑ i = 1 n x i ( 1 − p ) n , x i ∈ { 0 , 1 } , {\displaystyle f_{p}(x_{1},\dots ,x_{n})=\prod _{i=1}^{n}p^{x_{i}}(1-p)^{1-x_{i}}=\left({\frac {p}{1-p}}\right)^{\sum _{i=1}^{n}x_{i}}(1-p)^{n},\quad x_{i}\in \{0,1\},} thus, by the Fisher–Neyman factorization theorem, T ( x ) = ∑ i = 1 n x i {\displaystyle T(x)=\sum _{i=1}^{n}x_{i}} is a sufficient statistic. It is also complete: Let φ {\displaystyle \varphi } be any measurable function such that E θ [ φ ( T ( X ) ) ] = 0 {\displaystyle \mathbb {E} _{\theta }[\varphi (T(X))]=0} . Since T {\displaystyle T} is B e r ( n , p ) {\displaystyle \mathrm {Ber} (n,p)} -distributed, this means 0 = E θ [ φ ( T ( X ) ) ] = ∑ k = 0 n φ ( k ) ( n k ) p k ( 1 − p ) n − k = ( 1 − p ) n ∑ k = 0 n φ ( k ) ( n k ) r k , {\displaystyle 0=\mathbb {E} _{\theta }[\varphi (T(X))]=\sum _{k=0}^{n}\varphi (k){\binom {n}{k}}p^{k}(1-p)^{n-k}=(1-p)^{n}\sum _{k=0}^{n}\varphi (k){\binom {n}{k}}r^{k},} with r = p / ( 1 − p ) {\displaystyle r=p/(1-p)} . Since the right hand side is a polynomial in r > 0 {\displaystyle r>0} that is equal to zero, each coefficient must be zero as well, implying φ ( k ) = 0 {\displaystyle \varphi (k)=0} for all k {\displaystyle k} . For the estimation of the parameter p {\displaystyle p} , it is now easy to see that the UMVUE is the sample mean X ¯ = 1 n ∑ i = 1 n X i = 1 n T ( X ) , {\displaystyle {\overline {X}}={\frac {1}{n}}\sum _{i=1}^{n}X_{i}={\frac {1}{n}}T(X),} since it is unbiased and a function of T ( X ) {\displaystyle T(X)} . Finding the UMVUE for the parameter p ( 1 − p ) {\displaystyle p(1-p)} (that is, the variance of the distribution) is less obvious. According to method 1, we seek a function h {\displaystyle h} such that p ( 1 − p ) = E p [ h ( T ( X ) ) ] = ∑ k = 0 n h ( k ) ( n k ) p k ( 1 − p ) n − k . {\displaystyle p(1-p)=\mathbb {E} _{p}[h(T(X))]=\sum _{k=0}^{n}h(k){\binom {n}{k}}p^{k}(1-p)^{n-k}.} By defining ρ = p / ( 1 − p ) {\displaystyle \rho =p/(1-p)} , the above rewrites as ∑ k = 0 n h ( k ) ( n k ) ρ k = ρ ( 1 + ρ ) n − 2 = ∑ k = 1 n − 1 ( n − 1 k − 1 ) ρ k . {\displaystyle \sum _{k=0}^{n}h(k){\binom {n}{k}}\rho ^{k}=\rho (1+\rho )^{n-2}=\sum _{k=1}^{n-1}{\binom {n-1}{k-1}}\rho ^{k}.} Comparing the coefficients shows that h ( k ) = k ( n − k ) / ( n ( n − 1 ) ) {\displaystyle h(k)=k(n-k)/(n(n-1))} , thus, the UMVUE is given by h ( T ( X ) ) = n n − 1 X ¯ ( 1 − X ¯ ) . {\displaystyle h(T(X))={\frac {n}{n-1}}{\overline {X}}(1-{\overline {X}}).} Noting that in the Bernoulli model X i = X i 2 {\displaystyle X_{i}=X_{i}^{2}} and that ∑ i = 1 n ( X i − X ¯ ) 2 = ∑ i = 1 n X i 2 − n X ¯ 2 {\displaystyle \sum _{i=1}^{n}(X_{i}-{\overline {X}})^{2}=\sum _{i=1}^{n}X_{i}^{2}-n{\overline {X}}^{2}} , the UMVUE is, in fact, the unbiased sample variance S n 2 {\displaystyle S_{n}^{2}} .
Uniform distribution Let X 1 , … , X n {\displaystyle X_{1},\dots ,X_{n}} be uniformly distributed on the interval ( 0 , θ ) {\displaystyle (0,\theta )} for some θ {\displaystyle \theta } to be determined. The joint probability mass function f p {\displaystyle f_{p}} is of the form f p ( x 1 , … , x n ) = ∏ i = 1 n 1 θ 1 ( 0 , θ ) ( x i ) = 1 θ n 1 [ 0 , ∞ ) ( min i = 1 , … , n x i ) 1 ( − ∞ , θ ] ( max i = 1 , … , n x i ) , x ∈ R n , {\displaystyle f_{p}(x_{1},\dots ,x_{n})=\prod _{i=1}^{n}{\frac {1}{\theta }}1_{(0,\theta )}(x_{i})={\frac {1}{\theta ^{n}}}1_{[0,\infty )}\left(\min _{i=1,\dots ,n}x_{i}\right)1_{(-\infty ,\theta ]}\left(\max _{i=1,\dots ,n}x_{i}\right),\quad x\in \mathbb {R} ^{n},} thus, by the Fisher–Neyman factorization theorem, T ( x ) = max i = 1 , … , n x i {\displaystyle T(x)=\max _{i=1,\dots ,n}x_{i}} is a sufficient statistic. To see the completeness of T {\displaystyle T} , the distribution of T ( X ) {\displaystyle T(X)} is
P θ ( T ( X ) ≤ t ) = P θ ( X i ≤ t , ∀ i = 1 , … , n ) = ∏ i = 1 n P θ ( X i ≤ t ) = ∏ i = 1 n t θ = t n θ n . {\displaystyle \mathbb {P} _{\theta }(T(X)\leq t)=\mathbb {P} _{\theta }(X_{i}\leq t,\forall i=1,\dots ,n)=\prod _{i=1}^{n}\mathbb {P} _{\theta }(X_{i}\leq t)=\prod _{i=1}^{n}{\frac {t}{\theta }}={\frac {t^{n}}{\theta ^{n}}}.}
Thus, for any measurable function φ {\displaystyle \varphi } such that E θ [ φ ( T ( X ) ) ] = 0 {\displaystyle \mathbb {E} _{\theta }[\varphi (T(X))]=0} it holds for any θ > 0 {\displaystyle \theta >0}
0 = E θ [ φ ( T ( X ) ) ] = ∫ 0 θ φ ( x ) n θ n x n − 1 d x ⟹ ∫ 0 θ φ ( x ) x n − 1 d x = 0 , {\displaystyle 0=\mathbb {E} _{\theta }[\varphi (T(X))]=\int _{0}^{\theta }\varphi (x){\frac {n}{\theta ^{n}}}x^{n-1}\mathrm {d} x\implies \int _{0}^{\theta }\varphi (x)x^{n-1}\mathrm {d} x=0,}
and consequently for any b > a > 0 {\displaystyle b>a>0}
∫ a b φ ( x ) x n − 1 d x = ∫ 0 b φ ( x ) x n − 1 d x − ∫ 0 a φ ( x ) x n − 1 d x = 0. {\displaystyle \int _{a}^{b}\varphi (x)x^{n-1}\mathrm {d} x=\int _{0}^{b}\varphi (x)x^{n-1}\mathrm {d} x-\int _{0}^{a}\varphi (x)x^{n-1}\mathrm {d} x=0.}
This implies that φ = 0 {\displaystyle \varphi =0} (almost everywhere) on [ 0 , ∞ ) {\displaystyle [0,\infty )} . So T {\displaystyle T} is also complete. Its expectation is
E θ [ T ( X ) ] = ∫ 0 θ x n θ n x n − 1 d x = n θ n θ n + 1 n + 1 = n n + 1 θ . {\displaystyle \mathbb {E} _{\theta }[T(X)]=\int _{0}^{\theta }x{\frac {n}{\theta ^{n}}}x^{n-1}\mathrm {d} x={\frac {n}{\theta ^{n}}}{\frac {\theta ^{n+1}}{n+1}}={\frac {n}{n+1}}\theta .}
Thus, the rescaled estimator
T ∗ ( X ) = n + 1 n T ( X ) = n + 1 n max i = 1 , … , n X i {\displaystyle T^{\ast }(X)={\frac {n+1}{n}}T(X)={\frac {n+1}{n}}\max _{i=1,\dots ,n}X_{i}}
is unbiased and since it is a function of T {\displaystyle T} , it is the UMVUE for θ {\displaystyle \theta } .
Exponential distribution Let X 1 , … , X n {\displaystyle X_{1},\dots ,X_{n}} be exponentially distributed with parameter λ {\displaystyle \lambda } and suppose the parameter g ( λ ) = P λ ( X 1 > t ) {\displaystyle g(\lambda )=\mathbb {P} _{\lambda }(X_{1}>t)} should be estimated for some fixed t > 0 {\displaystyle t>0} . A straightforward unbiased estimator for
