Preply — Study more efficiently by working with a personal tutor. Get 50% off.Affiliate

Wikipedia

Pareto distribution

Pareto distribution

The Pareto distribution, named after the Italian polymath Vilfredo Pareto, is a probability distribution in the form of a power law that is used to describe social, quality control, scientific, geophysical, actuarial, and many other types of observable phenomena; the principle originally applied to describing the distribution of wealth in a society, fitting the trend that a large portion of wealth is held by a small fraction of the population. Empirical observation has shown that the Pareto distribution fits a wide range of cases, including natural phenomena and human activities. The Pareto principle or "80:20 rule" stating that 80% of outcomes are due to 20% of causes was named in honour of Pareto. Pareto distributions with a shape value (α) of log 4 5 ≈ 1.16 exhibit the Pareto principle.

Definitions If X is a random variable with a Pareto (Type I) distribution, then the probability that X is greater than some number x, i.e., the survival function (also called tail function), is given by

F ¯ ( x ) = Pr ( X > x ) = { ( x m x ) α x ≥ x m , 1 x < x m , {\displaystyle {\overline {F}}(x)=\Pr(X>x)={\begin{cases}\left({\frac {x_{\mathrm {m} }}{x}}\right)^{\alpha }&x\geq x_{\mathrm {m} },\\1&x<x_{\mathrm {m} },\end{cases}}}

where xm is the (necessarily positive) minimum possible value of X, and α is a positive parameter. The type I Pareto distribution is characterized by a scale parameter xm and a shape parameter α, which is known as the tail index. If this distribution is used to model the distribution of wealth, then the parameter α is called the Pareto index.

Cumulative distribution function From the definition, the cumulative distribution function of a Pareto random variable with parameters α and xm is

F X ( x ) = { 1 − ( x m x ) α x ≥ x m , 0 x < x m . {\displaystyle F_{X}(x)={\begin{cases}1-\left({\frac {x_{\mathrm {m} }}{x}}\right)^{\alpha }&x\geq x_{\mathrm {m} },\\0&x<x_{\mathrm {m} }.\end{cases}}}

Probability density function It follows (by differentiation) that the probability density function is

f X ( x ) = { α x m α x α + 1 x ≥ x m , 0 x < x m . {\displaystyle f_{X}(x)={\begin{cases}{\frac {\alpha x_{\mathrm {m} }^{\alpha }}{x^{\alpha +1}}}&x\geq x_{\mathrm {m} },\\0&x<x_{\mathrm {m} }.\end{cases}}}

When plotted on linear axes, the distribution assumes the familiar J-shaped curve which approaches each of the orthogonal axes asymptotically. All segments of the curve are self-similar (subject to appropriate scaling factors). When plotted in a log–log plot, the distribution is represented by a straight line.

Properties

Moments and characteristic function The expected value of a random variable following a Pareto distribution is E ⁡ ( X ) = { ∞ α ≤ 1 , α x m α − 1 α > 1. {\displaystyle \operatorname {E} (X)={\begin{cases}\infty &\alpha \leq 1,\\{\frac {\alpha x_{\mathrm {m} }}{\alpha -1}}&\alpha >1.\end{cases}}}

The variance of a random variable following a Pareto distribution is Var ⁡ ( X ) = { ∞ α ∈ ( 1 , 2 ] , ( x m α − 1 ) 2 α α − 2 α > 2. {\displaystyle \operatorname {Var} (X)={\begin{cases}\infty &\alpha \in (1,2],\\\left({\frac {x_{\mathrm {m} }}{\alpha -1}}\right)^{2}{\frac {\alpha }{\alpha -2}}&\alpha >2.\end{cases}}} (If α ≤ 1, the variance does not exist.) The raw moments are μ n ′ = { ∞ α ≤ n , α x m n α − n α > n . {\displaystyle \mu _{n}'={\begin{cases}\infty &\alpha \leq n,\\{\frac {\alpha x_{\mathrm {m} }^{n}}{\alpha -n}}&\alpha >n.\end{cases}}}

The moment generating function is only defined for non-positive values t ≤ 0 as M ( t ; α , x m ) = E ⁡ [ e t X ] = α ( − x m t ) α Γ ( − α , − x m t ) {\displaystyle M\left(t;\alpha ,x_{\mathrm {m} }\right)=\operatorname {E} \left[e^{tX}\right]=\alpha (-x_{\mathrm {m} }t)^{\alpha }\Gamma (-\alpha ,-x_{\mathrm {m} }t)} M ( 0 , α , x m ) = 1. {\displaystyle M\left(0,\alpha ,x_{\mathrm {m} }\right)=1.} Thus, since the expectation does not converge on an open interval containing t = 0 {\displaystyle t=0} we say that the moment generating function does not exist. The characteristic function is given by φ ( t ; α , x m ) = α ( − i x m t ) α Γ ( − α , − i x m t ) , {\displaystyle \varphi (t;\alpha ,x_{\mathrm {m} })=\alpha (-ix_{\mathrm {m} }t)^{\alpha }\Gamma (-\alpha ,-ix_{\mathrm {m} }t),} where Γ(a, x) is the incomplete gamma function. The parameters may be solved for using the method of moments.

Conditional distributions The conditional probability distribution of a Pareto-distributed random variable, given the event that it is greater than or equal to a particular number x 1 {\displaystyle x_{1}} exceeding x m {\displaystyle x_{\text{m}}} , is a Pareto distribution with the same Pareto index α {\displaystyle \alpha } but with minimum x 1 {\displaystyle x_{1}} instead of x m {\displaystyle x_{\text{m}}} :

Pr ( X ≥ x | X ≥ x 1 ) = { ( x 1 x ) α x ≥ x 1 , 1 x < x 1 . {\displaystyle {\text{Pr}}(X\geq x|X\geq x_{1})={\begin{cases}\left({\frac {x_{1}}{x}}\right)^{\alpha }&x\geq x_{1},\\1&x<x_{1}.\end{cases}}}

This implies that the conditional expected value (if it is finite, i.e. α > 1 {\displaystyle \alpha >1} ) is proportional to x 1 {\displaystyle x_{1}} :

E ( X | X ≥ x 1 ) ∝ x 1 . {\displaystyle {\text{E}}(X|X\geq x_{1})\propto x_{1}.}

In case of random variables that describe the lifetime of an object, this means that life expectancy is proportional to age, and is called the Lindy effect or Lindy's Law.

A characterization theorem Suppose X 1 , X 2 , X 3 , … {\displaystyle X_{1},X_{2},X_{3},\dotsc } are independent identically distributed random variables whose probability distribution is supported on the interval [ x m , ∞ ) {\displaystyle [x_{\text{m}},\infty )} for some x m > 0 {\displaystyle x_{\text{m}}>0} . Suppose that for all n {\displaystyle n} , the two random variables min { X 1 , … , X n } {\displaystyle \min\{X_{1},\dotsc ,X_{n}\}} and ( X 1 + ⋯ + X n ) / min { X 1 , … , X n } {\displaystyle (X_{1}+\dotsb +X_{n})/\min\{X_{1},\dotsc ,X_{n}\}} are independent. Then the common distribution is a Pareto distribution.

Geometric mean The geometric mean (G) is

G = x m exp ⁡ ( 1 α ) . {\displaystyle G=x_{\text{m}}\exp \left({\frac {1}{\alpha }}\right).}

Harmonic mean The harmonic mean (H) is

H = x m ( 1 + 1 α ) . {\displaystyle H=x_{\text{m}}\left(1+{\frac {1}{\alpha }}\right).}

Graphical representation The characteristic curved 'long tail' distribution, when plotted on a linear scale, masks the underlying simplicity of the function when plotted on a log-log graph, which then takes the form of a straight line with negative gradient: It follows from the formula for the probability density function that for x ≥ xm,

log ⁡ f X ( x ) = log ⁡ ( α x m α x α + 1 ) = log ⁡ ( α x m α ) − ( α + 1 ) log ⁡ x . {\displaystyle \log f_{X}(x)=\log \left(\alpha {\frac {x_{\mathrm {m} }^{\alpha }}{x^{\alpha +1}}}\right)=\log(\alpha x_{\mathrm {m} }^{\alpha })-(\alpha +1)\log x.}

Since α is positive, the gradient −(α + 1) is negative.

Related distributions

Generalized Pareto distributions

There is a hierarchy of Pareto distributions known as Pareto Type I, II, III, IV, and Feller–Pareto distributions. Pareto Type IV contains Pareto Type I–III as special cases. The Feller–Pareto distribution generalizes Pareto Type IV.

Pareto types I–IV The Pareto distribution hierarchy is summarized in the next table comparing the survival functions (complementary CDF). When μ = 0, the Pareto distribution Type II is also known as the Lomax distribution. In this section, the symbol xm, used before to indicate the minimum value of x, is replaced by σ.

The shape parameter α is the tail index, μ is location, σ is scale, γ is an inequality parameter. Some special cases of Pareto Type (IV) are

P ( I V ) ( σ , σ , 1 , α ) = P ( I ) ( σ , α ) , {\displaystyle P(IV)(\sigma ,\sigma ,1,\alpha )=P(I)(\sigma ,\alpha ),}

P ( I V ) ( μ , σ , 1 , α ) = P ( I I ) ( μ , σ , α ) , {\displaystyle P(IV)(\mu ,\sigma ,1,\alpha )=P(II)(\mu ,\sigma ,\alpha ),}

P ( I V ) ( μ , σ , γ , 1 ) = P ( I I I ) ( μ , σ , γ ) . {\displaystyle P(IV)(\mu ,\sigma ,\gamma ,1)=P(III)(\mu ,\sigma ,\gamma ).}

The finiteness of the mean, and the existence and the finiteness of the variance depend on the tail index α (inequality index γ). In particular, fractional δ-moments are finite for some δ > 0, as shown in the table below, where δ is not necessarily an integer.

Feller–Pareto distribution Feller defines a Pareto variable by transformation U = Y−1 − 1 of a beta random variable ,Y, whose probability density function is

f ( y ) = y γ 1 − 1 ( 1 − y ) γ 2 − 1 B ( γ 1 , γ 2 ) , 0 < y < 1 ; γ 1 , γ 2 > 0 , {\displaystyle f(y)={\frac {y^{\gamma _{1}-1}(1-y)^{\gamma _{2}-1}}{B(\gamma _{1},\gamma _{2})}},\qquad 0<y<1;\gamma _{1},\gamma _{2}>0,}

where B( ) is the beta function. If

W = μ + σ ( Y − 1 − 1 ) γ , σ > 0 , γ > 0 , {\displaystyle W=\mu +\sigma (Y^{-1}-1)^{\gamma },\qquad \sigma >0,\gamma >0,}

then W has a Feller–Pareto distribution FP(μ, σ, γ, γ1, γ2). If U 1 ∼ Γ ( δ 1 , 1 ) {\displaystyle U_{1}\sim \Gamma (\delta _{1},1)} and U 2 ∼ Γ ( δ 2 , 1 ) {\displaystyle U_{2}\sim \Gamma (\delta _{2},1)} are independent Gamma variables, another construction of a Feller–Pareto (FP) variable is

W = μ + σ ( U 1 U 2 ) γ {\displaystyle W=\mu +\sigma \left({\frac {U_{1}}{U_{2}}}\right)^{\gamma }}

and we write W ~ FP(μ, σ, γ, δ1, δ2). Special cases of the Feller–Pareto distribution are

F P ( σ , σ , 1 , 1 , α ) = P ( I ) ( σ , α ) {\displaystyle FP(\sigma ,\sigma ,1,1,\alpha )=P(I)(\sigma ,\alpha )}

F P ( μ , σ , 1 , 1 , α ) = P ( I I ) ( μ , σ , α ) {\displaystyle FP(\mu ,\sigma ,1,1,\alpha )=P(II)(\mu ,\sigma ,\alpha )}

F P ( μ , σ , γ , 1 , 1 ) = P ( I I I ) ( μ , σ , γ ) {\displaystyle FP(\mu ,\sigma ,\gamma ,1,1)=P(III)(\mu ,\sigma ,\gamma )}

F P ( μ , σ , γ , 1 , α ) = P ( I V ) ( μ , σ , γ , α ) . {\displaystyle FP(\mu ,\sigma ,\gamma ,1,\alpha )=P(IV)(\mu ,\sigma ,\gamma ,\alpha ).}

Inverse-Pareto Distribution / Power Distribution When a random variable Y {\displaystyle Y} follows a pareto distribution, then its inverse X = 1 / Y {\displaystyle X=1/Y} follows a Power distribution. Inverse Pareto distribution is equivalent to a Power distribution

Y ∼ P a ( α , x m ) = α x m α y α + 1 ( y ≥ x m ) ⇔ X ∼ i P a ( α , x m ) = P o w e r ( x m − 1 , α ) = α x α − 1 ( x m − 1 ) α ( 0 < x ≤ x m − 1 ) {\displaystyle Y\sim \mathrm {Pa} (\alpha ,x_{\mathrm {m} })={\frac {\alpha x_{\mathrm {m} }^{\alpha }}{y^{\alpha +1}}}\quad (y\geq x_{\mathrm {m} })\quad \Leftrightarrow \quad X\sim \mathrm {iPa} (\alpha ,x_{\mathrm {m} })=\mathrm {Power} (x_{\mathrm {m} }^{-1},\alpha )={\frac {\alpha x^{\alpha -1}}{(x_{\mathrm {m} }^{-1})^{\alpha }}}\quad (0<x\leq x_{\mathrm {m} }^{-1})}

Relation to the exponential distribution The Pareto distribution is related to the exponential distribution as follows. If X is Pareto-distributed with minimum xm and index α, then

Y = log ⁡ ( X x m ) {\displaystyle Y=\log \left({\frac {X}{x_{\mathrm {m} }}}\right)}

is exponentially distributed with rate parameter α. Equivalently, if Y is exponentially distributed with rate α, then

x m e Y {\displaystyle x_{\mathrm {m} }e^{Y}}

is Pareto-distributed with minimum xm and index α. This can be shown using the standard change-of-variable techniques:

Pr ( Y < y ) = Pr ( log ⁡ ( X x m ) < y ) = Pr ( X < x m e y ) = 1 − ( x m x m e y ) α = 1 − e − α y . {\displaystyle {\begin{aligned}\Pr(Y<y)&=\Pr \left(\log \left({\frac {X}{x_{\mathrm {m} }}}\right)<y\right)\\&=\Pr(X<x_{\mathrm {m} }e^{y})=1-\left({\frac {x_{\mathrm {m} }}{x_{\mathrm {m} }e^{y}}}\right)^{\alpha }=1-e^{-\alpha y}.\end{aligned}}}

The last expression is the cumulative distribution function of an exponential distribution with rate α. Pareto distribution can be constructed by hierarchical exponential distributions. Let ϕ | a ∼ Exp ( a ) {\displaystyle \phi |a\sim {\text{Exp}}(a)} and η | ϕ ∼ Exp ( ϕ ) {\displaystyle \eta |\phi \sim {\text{Exp}}(\phi )} . Then we have p ( η | a ) = a ( a + η ) 2 {\displaystyle p(\eta |a)={\frac {a}{(a+\eta )^{2}}}} and, as a result, a + η ∼ Pareto ( a , 1 ) {\displaystyle a+\eta \sim {\text{Pareto}}(a,1)} . More in general, if λ ∼ Gamma ( α , β ) {\displaystyle \lambda \sim {\text{Gamma}}(\alpha ,\beta )} (shape-rate parametrization) and η | λ ∼ Exp ( λ ) {\displaystyle \eta |\lambda \sim {\text{Exp}}(\lambda )} , then β + η ∼ Pareto ( β , α ) {\displaystyle \beta +\eta \sim {\text{Pareto}}(\beta ,\alpha )} . Equivalently, if Y ∼ Gamma ( α , 1 ) {\displaystyle Y\sim {\text{Gamma}}(\alpha ,1)} and X ∼ Exp ( 1 ) {\displaystyle X\sim {\text{Exp}}(1)} , then x m ( 1 + X Y ) ∼ Pareto ( x m , α ) {\displaystyle x_{\text{m}}\!\left(1+{\frac {X}{Y}}\right)\sim {\text{Pareto}}(x_{\text{m}},\alpha )} .

Relation to the log-normal distribution The Pareto distribution and log-normal distribution are alternative distributions for describing the same types of quantities. One of the connections between the two is that they are both the distributions of the exponential of random variables distributed according to other common distributions, respectively the exponential distribution and normal distribution. (See the previous section.)

Relation to the generalized Pareto distribution The Pareto distribution is a special case of the generalized Pareto distribution, which is a family of distributions of similar form, but containing an extra parameter in such a way that the support of the distribution is either bounded below (at a variable point), or bounded both above and below (where both are variable), with the Lomax distribution as a special case. This family also contains both the unshifted and shifted exponential distributions. The Pareto distribution with scale x m {\displaystyle x_{\mathrm {m} }} and shape α {\displaystyle \alpha } is equivalent to the generalized Pareto distribution with location μ = x m {\displaystyle \mu =x_{\mathrm {m} }} , scale σ = x m / α {\displaystyle \sigma =x_{\mathrm {m}

Tags

  • Actuarial science
  • Continuous distributions
  • Eponyms in economics
  • Exponential family distributions
  • Management science
  • Power laws
  • Probability distributions with non-finite variance
  • Vilfredo Pareto