V-statistics are a class of statistics named for Richard von Mises who developed their asymptotic distribution theory in a fundamental paper in 1947. V-statistics are closely related to U-statistics (U for "unbiased") introduced by Wassily Hoeffding in 1948. A V-statistic is a statistical function (of a sample) defined by a particular statistical functional of a probability distribution.
Statistical functions Statistics that can be represented as functionals T ( F n ) {\displaystyle T(F_{n})} of the empirical distribution function ( F n ) {\displaystyle (F_{n})} are called statistical functionals. Differentiability of the functional T plays a key role in the von Mises approach; thus von Mises considers differentiable statistical functionals.
Examples of statistical functions
The k-th central moment is the functional T ( F ) = ∫ ( x − μ ) k d F ( x ) {\displaystyle T(F)=\int (x-\mu )^{k}\,dF(x)} , where μ = E [ X ] {\displaystyle \mu =E[X]} is the expected value of X. The associated statistical function is the sample k-th central moment,
T n = m k = T ( F n ) = 1 n ∑ i = 1 n ( x i − x ¯ ) k . {\displaystyle T_{n}=m_{k}=T(F_{n})={\frac {1}{n}}\sum _{i=1}^{n}(x_{i}-{\overline {x}})^{k}.}
The chi-squared goodness-of-fit statistic is a statistical function T(Fn), corresponding to the statistical functional
T ( F ) = ∑ i = 1 k ( ∫ A i d F − p i ) 2 p i , {\displaystyle T(F)=\sum _{i=1}^{k}{\frac {(\int _{A_{i}}\,dF-p_{i})^{2}}{p_{i}}},}
where Ai are the k cells and pi are the specified probabilities of the cells under the null hypothesis.
The Cramér–von-Mises and Anderson–Darling goodness-of-fit statistics are based on the functional
T ( F ) = ∫ ( F ( x ) − F 0 ( x ) ) 2 w ( x ; F 0 ) d F 0 ( x ) , {\displaystyle T(F)=\int (F(x)-F_{0}(x))^{2}\,w(x;F_{0})\,dF_{0}(x),}
where w(x; F0) is a specified weight function and F0 is a specified null distribution. If w is the identity function then T(Fn) is the well known Cramér–von-Mises goodness-of-fit statistic; if w ( x ; F 0 ) = [ F 0 ( x ) ( 1 − F 0 ( x ) ) ] − 1 {\displaystyle w(x;F_{0})=[F_{0}(x)(1-F_{0}(x))]^{-1}} then T(Fn) is the Anderson–Darling statistic.
Representation as a V-statistic Suppose x1, ..., xn is a sample. In typical applications the statistical function has a representation as the V-statistic
… excerpt ends here. Continue reading the full article.
