ArticleslgStudy

mathematics

Statistical distance

Statistical distance is a mathematics topic covered in the lgStudy science library. This page brings together a partial reference excerpt, illustrations, worked examples, real-world applications and a short study plan, so you can understand Statistical distance rather than just read about it. In short: In statistics, probability theory, and information theory, a statistical distance quantifies the distance between two statistical objects, which can be two random variables, or two probability distributions or samples, or the distance can be between an individual sample point and a population or a wider sample of points. A distance between populations can be interpreted as measuring the distance between two probabil…

Key takeaways

  • Statistical distance belongs to mathematics; place it in that map before memorising details.
  • Learn the definition first, then one example that makes the definition concrete.
  • Connect Statistical distance to a quantity you can measure, compute or draw — that is where exam questions come from.
  • Reproduce the core statement of Statistical distance from memory before moving on to harder problems.

Reference excerpt

In statistics, probability theory, and information theory, a statistical distance quantifies the distance between two statistical objects, which can be two random variables, or two probability distributions or samples, or the distance can be between an individual sample point and a population or a wider sample of points. A distance between populations can be interpreted as measuring the distance between two probability distributions and hence they are essentially measures of distances between probability measures. Where statistical distance measures relate to the differences between random variables, these may have statistical dependence, and hence these distances are not directly related to measures of distances between probability measures. Again, a measure of distance between random variables may relate to the extent of dependence between them, rather than to their individual values. Many statistical distance measures are not metrics, and some are not symmetric. Some types of distance measures, which generalize squared distance, are referred to as (statistical) divergences.

Terminology Many terms are used to refer to various notions of distance; these are often confusingly similar, and may be used inconsistently between authors and over time, either loosely or with precise technical meaning. In addition to "distance", similar terms include deviance, deviation, discrepancy, discrimination, and divergence, as well as others such as contrast function and metric. Terms from information theory include cross entropy, relative entropy, discrimination information, and information gain.

Distances as metrics

Metrics A metric on a set X is a function (called the distance function or simply distance) d : X × X → R+ (where R+ is the set of non-negative real numbers). For all x, y, z in X, this function is required to satisfy the following conditions:

d(x, y) ≥ 0 (non-negativity) d(x, y) = 0 if and only if x = y (identity of indiscernibles. Note that condition 1 and 2 together produce positive definiteness) d(x, y) = d(y, x) (symmetry) d(x, z) ≤ d(x, y) + d(y, z) (subadditivity / triangle inequality).

Generalized metrics Many statistical distances are not metrics, because they lack one or more properties of proper metrics. For example, pseudometrics violate property (2), identity of indiscernibles; quasimetrics violate property (3), symmetry; and semimetrics violate property (4), the triangle inequality. Statistical distances that satisfy (1) and (2) are referred to as divergences.

Statistically close The total variation distance of two distributions X {\displaystyle X} and Y {\displaystyle Y} over a finite domain D {\displaystyle D} , (often referred to as statistical difference or statistical distance in cryptography) is defined as

Δ ( X , Y ) = 1 2 ∑ α ∈ D | Pr [ X = α ] − Pr [ Y = α ] | {\displaystyle \Delta (X,Y)={\frac {1}{2}}\sum _{\alpha \in D}|\Pr[X=\alpha ]-\Pr[Y=\alpha ]|} . We say that two probability ensembles { X k } k ∈ N {\displaystyle \{X_{k}\}_{k\in \mathbb {N} }} and { Y k } k ∈ N {\displaystyle \{Y_{k}\}_{k\in \mathbb {N} }} are statistically close if Δ ( X k , Y k ) {\displaystyle \Delta (X_{k},Y_{k})} is a negligible function in k {\displaystyle k} .

Examples

Metrics Total variation distance (sometimes just called "the" statistical distance) Hellinger distance Lévy–Prokhorov metric Wasserstein metric: also known as the Kantorovich metric, or earth mover's distance Mahalanobis distance Integral probability metrics generalize several metrics or pseudometrics on distributions

Divergences Kullback–Leibler divergence Rényi divergence Jensen–Shannon divergence Ball divergence Bhattacharyya distance (despite its name it is not a distance, as it violates the triangle inequality) f-divergence: generalizes several distances and divergences Discriminability index, specifically the Bayes discriminability index, is a positive-definite symmetric measure of the overlap of two distributions.

See also Probabilistic metric space Randomness extractor Similarity measure Zero-knowledge proof

Notes

External links Distance and Similarity Measures (Wolfram Alpha)

References Dodge, Y. (2003) Oxford Dictionary of Statistical Terms, OUP. ISBN 0-19-920613-9

Worked examples

Example 1 — a first encounter with Statistical distance

Start with the simplest possible case. Write down what Statistical distance claims or describes in one sentence, then invent the smallest concrete situation in which that sentence is true. In mathematics, the smallest case is usually a single object, a single equation or a single measurement. Check that every symbol or term in your sentence has a meaning in that case.

Example 2 — changing one variable

Take the situation from Example 1 and change exactly one quantity: double it, halve it, or set it to zero. Predict what should happen to Statistical distance before you calculate. Comparing your prediction with the result is the fastest way to find out whether you understand the idea or only the words.

Example 3 — an exam-style question

Typical questions about Statistical distance ask you to (a) state it precisely, (b) apply it to given data, and (c) explain a limitation. Practise writing all three answers in under five minutes; the third part is what separates a full-mark answer from an average one.

Applications of Statistical distance

In research
Statistical distance appears in mathematics research whenever the underlying quantities have to be modelled precisely. Papers usually cite it as a starting assumption and then explore where it breaks down.
In technology and industry
Engineering practice reuses Statistical distance in design rules, simulations and safety margins. Knowing the idea lets you read a specification sheet and understand why the numbers look the way they do.
In the classroom
Statistical distance is common in secondary-school and first-year university syllabi. It links to neighbouring topics Statistical distance, so understanding it makes those chapters shorter.
In everyday life
Look for Statistical distance outside the textbook — in sport, cooking, traffic, electronics or the sky above you. An example you found yourself is remembered far longer than one you were given.

Affiliate

Preply — study more efficiently by working with a personal tutor. 50% off.

How to study Statistical distance in 20 minutes

  1. Read the reference excerpt below once, without taking notes.
  2. Close the page and write down what Statistical distance means in your own words.
  3. Compare your version with the excerpt and mark what you missed.
  4. Work through the three examples above with pen and paper.
  5. Explain Statistical distance out loud to somebody else — or to Teacher Smith in the lgStudy chat.

Frequently asked questions

What is Statistical distance in simple terms?

In statistics, probability theory, and information theory, a statistical distance quantifies the distance between two statistical objects, which can be two random variables, or two probability distributions or samples, or the distance can be between an individual sample point and a population or a…

Why does Statistical distance matter?

Because it connects several mathematics ideas at once: it gives you a definition you can apply, a quantity you can calculate, and a way to check whether a result is plausible.

How should I study Statistical distance?

Read the excerpt, restate it from memory, then work through the examples and applications listed on this page. The five-step study plan above takes about twenty minutes.

What does this page cover?

It gives you a compact reference excerpt plus original lgStudy explanations, examples, applications and study material on Statistical distance.

Tags

  • Statistical distance

Keep exploring