In probability theory, the theory of large deviations concerns the asymptotic behaviour of remote tails of sequences of probability distributions. While some basic ideas of the theory can be traced to Laplace, the formalization started with insurance mathematics, namely ruin theory with Cramér and Lundberg. A unified formalization of large deviation theory was developed in 1966, in a paper by Varadhan. Large deviations theory formalizes the heuristic ideas of concentration of measures and widely generalizes the notion of convergence of probability measures. Roughly speaking, large deviations theory concerns itself with the exponential decline of the probability measures of certain kinds of extreme or tail events.
Introductory examples Any large deviation is done in the least unlikely of all the unlikely ways!
An elementary example Consider a sequence of independent tosses of a fair coin. The possible outcomes could be heads or tails. Let us denote the possible outcome of the i-th trial by X i {\displaystyle X_{i}} , where we encode head as 1 and tail as 0. Now let M N {\displaystyle M_{N}} denote the mean value after N {\displaystyle N} trials, namely
M N = 1 N ∑ i = 1 N X i {\displaystyle M_{N}={\frac {1}{N}}\sum _{i=1}^{N}X_{i}} . Then M N {\displaystyle M_{N}} lies between 0 and 1. From the law of large numbers it follows that as N grows, the distribution of M N {\displaystyle M_{N}} converges to 0.5 = E [ X ] {\displaystyle 0.5=\operatorname {E} [X]} (the expected value of a single coin toss). Moreover, by the central limit theorem, it follows that M N {\displaystyle M_{N}} is approximately normally distributed for large N {\displaystyle N} . The central limit theorem can provide more detailed information about the behavior of M N {\displaystyle M_{N}} than the law of large numbers. For example, we can approximately find a tail probability of M N {\displaystyle M_{N}} – the probability that M N {\displaystyle M_{N}} is greater than some value x {\displaystyle x} – for a fixed value of N {\displaystyle N} . However, the approximation by the central limit theorem may not be accurate if x {\displaystyle x} is far from E [ X i ] {\displaystyle \operatorname {E} [X_{i}]} and N {\displaystyle N} is not sufficiently large. Also, it does not provide information about the convergence of the tail probabilities as N → ∞ {\displaystyle N\to \infty } . However, the large deviation theory can provide answers for such problems. Let us make this statement more precise. For a given value 0.5 < x < 1 {\displaystyle 0.5<x<1} , let us compute the tail probability P ( M N > x ) {\displaystyle P(M_{N}>x)} . Define
… excerpt ends here. Continue reading the full article.
