In mathematics, specifically functional analysis, Mercer's theorem is a representation of a symmetric positive-definite function on a square as a sum of a convergent sequence of product functions. This theorem, presented in (Mercer 1909), is one of the most notable results of the work of James Mercer (1883–1932). It is an important theoretical tool in the theory of integral equations; it is used in the Hilbert space theory of stochastic processes, for example the Karhunen–Loève theorem; and it is also used in the reproducing kernel Hilbert space theory where it characterizes a symmetric positive-definite kernel as a reproducing kernel.
Introduction To explain Mercer's theorem, we first consider an important special case; see below for a more general formulation. A kernel, in this context, is a symmetric continuous function
K : [ a , b ] × [ a , b ] → R {\displaystyle K:[a,b]\times [a,b]\rightarrow \mathbb {R} }
where K ( x , y ) = K ( y , x ) {\displaystyle K(x,y)=K(y,x)} for all x , y ∈ [ a , b ] {\displaystyle x,y\in [a,b]} . K is said to be a positive-definite kernel if and only if
∑ i = 1 n ∑ j = 1 n K ( x i , x j ) c i c j ≥ 0 {\displaystyle \sum _{i=1}^{n}\sum _{j=1}^{n}K(x_{i},x_{j})c_{i}c_{j}\geq 0}
for all finite sequences of points x1, ..., xn of [a, b] and all choices of real numbers c1, ..., cn. Note that the term "positive-definite" is well-established in literature despite the weak inequality in the definition. The fundamental characterization of stationary positive-definite kernels (where K ( x , y ) = K ( x − y ) {\displaystyle K(x,y)=K(x-y)} ) is given by Bochner's theorem. It states that a continuous function K ( x − y ) {\displaystyle K(x-y)} is positive-definite if and only if it can be expressed as the Fourier transform of a finite non-negative measure μ {\displaystyle \mu } :
K ( x − y ) = ∫ − ∞ ∞ e i ( x − y ) ω d μ ( ω ) {\displaystyle K(x-y)=\int _{-\infty }^{\infty }e^{i(x-y)\omega }\,d\mu (\omega )}
This spectral representation reveals the connection between positive definiteness and harmonic analysis, providing a stronger and more direct characterization of positive definiteness than the abstract definition in terms of inequalities when the kernel is stationary, e.g., when it can be expressed as a 1-variable function of the distance between points rather than the 2-variable function of the positions of pairs of points. Associated to K is a linear operator (more specifically a Hilbert–Schmidt integral operator when the interval is compact) on functions defined by the integral
[ T K φ ] ( x ) = ∫ a b K ( x , s ) φ ( s ) d s . {\displaystyle [T_{K}\varphi ](x)=\int _{a}^{b}K(x,s)\varphi (s)\,ds.}
We assume φ {\displaystyle \varphi } can range through the space of real-valued square-integrable functions L2[a, b]; however, in many cases the associated reproducing kernel Hilbert space can be strictly larger than L2[a, b]. Since TK is a linear operator, the eigenvalues and eigenfunctions of TK exist. Theorem. Suppose K is a continuous symmetric positive-definite kernel. Then there is an orthonormal basis {ei}i of L2[a, b] consisting of eigenfunctions of TK such that the corresponding sequence of eigenvalues {λi}i is nonnegative. The eigenfunctions corresponding to non-zero eigenvalues are continuous on [a, b] and K has the representation
K ( s , t ) = ∑ j = 1 ∞ λ j e j ( s ) e j ( t ) {\displaystyle K(s,t)=\sum _{j=1}^{\infty }\lambda _{j}\,e_{j}(s)\,e_{j}(t)}
where the convergence is absolute and uniform.
Details We now explain in greater detail the structure of the proof of Mercer's theorem, particularly how it relates to spectral theory of compact operators.
… excerpt ends here. Continue reading the full article.
