In statistics, the Khmaladze transformation is a mathematical tool used in constructing convenient goodness of fit tests for hypothetical distribution functions. More precisely, suppose X 1 , … , X n {\displaystyle X_{1},\ldots ,X_{n}} are i.i.d., possibly multi-dimensional, random observations generated from an unknown probability distribution. A classical problem in statistics is to decide how well a given hypothetical distribution function F {\displaystyle F} , or a given hypothetical parametric family of distribution functions { F θ : θ ∈ Θ } {\displaystyle \{F_{\theta }:\theta \in \Theta \}} , fits the set of observations. The Khmaladze transformation allows us to construct goodness of fit tests with desirable properties. It is named after Estate V. Khmaladze. Consider the sequence of empirical distribution functions F n {\displaystyle F_{n}} based on a sequence of i.i.d random variables, X 1 , … , X n {\displaystyle X_{1},\ldots ,X_{n}} , as n increases. Suppose F {\displaystyle F} is the hypothetical distribution function of each X i {\displaystyle X_{i}} . To test whether the choice of F {\displaystyle F} is correct or not, statisticians use the normalized difference,
v n ( x ) = n [ F n ( x ) − F ( x ) ] . {\displaystyle v_{n}(x)={\sqrt {n}}[F_{n}(x)-F(x)].}
This v n {\displaystyle v_{n}} , as a random process in x {\displaystyle x} , is called the empirical process. Various functionals of v n {\displaystyle v_{n}} are used as test statistics. The change of the variable v n ( x ) = u n ( t ) {\displaystyle v_{n}(x)=u_{n}(t)} , t = F ( x ) {\displaystyle t=F(x)} transforms to the so-called uniform empirical process u n {\displaystyle u_{n}} . The latter is an empirical processes based on independent random variables U i = F ( X i ) {\displaystyle U_{i}=F(X_{i})} , which are uniformly distributed on [ 0 , 1 ] {\displaystyle [0,1]} if the X i {\displaystyle X_{i}} s do indeed have distribution function F {\displaystyle F} . This fact was discovered and first utilized by Kolmogorov (1933), Wald and Wolfowitz (1936) and Smirnov (1937) and, especially after Doob (1949) and Anderson and Darling (1952), it led to the standard rule to choose test statistics based on v n {\displaystyle v_{n}} . That is, test statistics ψ ( v n , F ) {\displaystyle \psi (v_{n},F)} are defined (which possibly depend on the F {\displaystyle F} being tested) in such a way that there exists another statistic φ ( u n ) {\displaystyle \varphi (u_{n})} derived from the uniform empirical process, such that ψ ( v n , F ) = φ ( u n ) {\displaystyle \psi (v_{n},F)=\varphi (u_{n})} . Examples are
… excerpt ends here. Continue reading the full article.
