In probability theory and statistics, the normal-Wishart distribution (or Gaussian-Wishart distribution) is a multivariate four-parameter family of continuous probability distributions. It is the conjugate prior of a multivariate normal distribution with unknown mean and precision matrix (the inverse of the covariance matrix).
Definition Suppose
μ | μ 0 , λ , Λ ∼ N ( μ 0 , ( λ Λ ) − 1 ) {\displaystyle {\boldsymbol {\mu }}|{\boldsymbol {\mu }}_{0},\lambda ,{\boldsymbol {\Lambda }}\sim {\mathcal {N}}({\boldsymbol {\mu }}_{0},(\lambda {\boldsymbol {\Lambda }})^{-1})}
has a multivariate normal distribution with mean μ 0 {\displaystyle {\boldsymbol {\mu }}_{0}} and covariance matrix ( λ Λ ) − 1 {\displaystyle (\lambda {\boldsymbol {\Lambda }})^{-1}} , where
Λ | W , ν ∼ W ( Λ | W , ν ) {\displaystyle {\boldsymbol {\Lambda }}|\mathbf {W} ,\nu \sim {\mathcal {W}}({\boldsymbol {\Lambda }}|\mathbf {W} ,\nu )}
has a Wishart distribution. Then ( μ , Λ ) {\displaystyle ({\boldsymbol {\mu }},{\boldsymbol {\Lambda }})}
has a normal-Wishart distribution, denoted as
( μ , Λ ) ∼ N W ( μ 0 , λ , W , ν ) . {\displaystyle ({\boldsymbol {\mu }},{\boldsymbol {\Lambda }})\sim \mathrm {NW} ({\boldsymbol {\mu }}_{0},\lambda ,\mathbf {W} ,\nu ).}
Characterization
Probability density function
f ( μ , Λ | μ 0 , λ , W , ν ) = N ( μ | μ 0 , ( λ Λ ) − 1 ) W ( Λ | W , ν ) {\displaystyle f({\boldsymbol {\mu }},{\boldsymbol {\Lambda }}|{\boldsymbol {\mu }}_{0},\lambda ,\mathbf {W} ,\nu )={\mathcal {N}}({\boldsymbol {\mu }}|{\boldsymbol {\mu }}_{0},(\lambda {\boldsymbol {\Lambda }})^{-1})\ {\mathcal {W}}({\boldsymbol {\Lambda }}|\mathbf {W} ,\nu )}
Properties
Scaling
Marginal distributions By construction, the marginal distribution over Λ {\displaystyle {\boldsymbol {\Lambda }}} is a Wishart distribution, and the conditional distribution over μ {\displaystyle {\boldsymbol {\mu }}} given Λ {\displaystyle {\boldsymbol {\Lambda }}} is a multivariate normal distribution. The marginal distribution over μ {\displaystyle {\boldsymbol {\mu }}} is a multivariate t-distribution.
Posterior distribution of the parameters After making n {\displaystyle n} observations x 1 , … , x n {\displaystyle {\boldsymbol {x}}_{1},\dots ,{\boldsymbol {x}}_{n}} , the posterior distribution of the parameters is
( μ , Λ ) ∼ N W ( μ n , λ n , W n , ν n ) , {\displaystyle ({\boldsymbol {\mu }},{\boldsymbol {\Lambda }})\sim \mathrm {NW} ({\boldsymbol {\mu }}_{n},\lambda _{n},\mathbf {W} _{n},\nu _{n}),}
where
λ n = λ + n , {\displaystyle \lambda _{n}=\lambda +n,}
… excerpt ends here. Continue reading the full article.
