ArticleslgStudy

mathematics

Maximum likelihood estimation

Maximum likelihood estimation is a mathematics topic covered in the lgStudy science library. This page brings together a partial reference excerpt, illustrations, worked examples, real-world applications and a short study plan, so you can understand Maximum likelihood estimation rather than just read about it. In short: In statistics, maximum likelihood estimation (MLE) is a method of estimating the parameters of an assumed probability distribution, given some observed data. This is achieved by maximizing a likelihood function so that, under the assumed statistical model, the observed data is most probable.

Maximum likelihood estimation — main illustration
Maximum likelihood estimation — illustration

Key takeaways

  • Maximum likelihood estimation belongs to mathematics; place it in that map before memorising details.
  • Learn the definition first, then one example that makes the definition concrete.
  • Connect Maximum likelihood estimation to a quantity you can measure, compute or draw — that is where exam questions come from.
  • Reproduce the core statement of Maximum likelihood estimation from memory before moving on to harder problems.

Reference excerpt

In statistics, maximum likelihood estimation (MLE) is a method of estimating the parameters of an assumed probability distribution, given some observed data. This is achieved by maximizing a likelihood function so that, under the assumed statistical model, the observed data is most probable. The point in the parameter space that maximizes the likelihood function is called the maximum likelihood estimate. The logic of maximum likelihood is both intuitive and flexible, and as such the method has become a dominant means of statistical inference. If the likelihood function is differentiable, the derivative test for finding maxima can be applied. In some cases, the first-order conditions of the likelihood function can be solved analytically; for instance, the ordinary least squares estimator for a linear regression model maximizes the likelihood when the random errors are assumed to have normal distributions with the same variance. From the perspective of Bayesian inference, MLE is generally equivalent to maximum a posteriori (MAP) estimation with a prior distribution that is uniform in the region of interest. In frequentist inference, MLE is a special case of an extremum estimator, with the objective function being the likelihood.

Principles We model a set of observations as a random sample from an unknown joint probability distribution which is expressed in terms of a set of parameters. The goal of maximum likelihood estimation is to determine the parameters for which the observed data have the highest joint probability. We write the parameters governing the joint distribution as a vector θ = [ θ 1 , θ 2 , … , θ k ] T {\displaystyle \;\theta =\left[\theta _{1},\,\theta _{2},\,\ldots ,\,\theta _{k}\right]^{\mathsf {T}}\;} so that this distribution falls within a parametric family { f ( ⋅ ; θ ) ∣ θ ∈ Θ } , {\displaystyle \;\{f(\cdot \,;\theta )\mid \theta \in \Theta \}\;,} where Θ {\displaystyle \,\Theta \,} is called the parameter space, a finite-dimensional subset of Euclidean space. Evaluating the joint density at the observed data sample y = ( y 1 , y 2 , … , y n ) {\displaystyle \;\mathbf {y} =(y_{1},y_{2},\ldots ,y_{n})\;} gives a real-valued function,

L n ( θ ) = L n ( θ ; y ) = f n ( y ; θ ) , {\displaystyle {\mathcal {L}}_{n}(\theta )={\mathcal {L}}_{n}(\theta ;\mathbf {y} )=f_{n}(\mathbf {y} ;\theta )\;,}

which is called the likelihood function. For independent random variables, f n ( y ; θ ) {\displaystyle f_{n}(\mathbf {y} ;\theta )} will be the product of univariate density functions:

f n ( y ; θ ) = ∏ k = 1 n f k u n i v a r ( y k ; θ ) . {\displaystyle f_{n}(\mathbf {y} ;\theta )=\prod _{k=1}^{n}\,f_{k}^{\mathsf {univar}}(y_{k};\theta )~.}

The goal of maximum likelihood estimation is to find the values of the model parameters that maximize the likelihood function over the parameter space, that is:

θ ^ = a r g m a x θ ∈ Θ L n ( θ ; y ) . {\displaystyle {\hat {\theta }}={\underset {\theta \in \Theta }{\operatorname {arg\;max} }}\,{\mathcal {L}}_{n}(\theta \,;\mathbf {y} )~.}

… excerpt ends here. Continue reading the full article.

Illustrations

Maximum likelihood estimation: Likelihood function for proportion value of a binomial process (n = 10)
Likelihood function for proportion value of a binomial process (n = 10)
Maximum likelihood estimation: Ronald Fisher in 1913
Ronald Fisher in 1913

Worked examples

Example 1 — a first encounter with Maximum likelihood estimation

Start with the simplest possible case. Write down what Maximum likelihood estimation claims or describes in one sentence, then invent the smallest concrete situation in which that sentence is true. In mathematics, the smallest case is usually a single object, a single equation or a single measurement. Check that every symbol or term in your sentence has a meaning in that case.

Example 2 — changing one variable

Take the situation from Example 1 and change exactly one quantity: double it, halve it, or set it to zero. Predict what should happen to Maximum likelihood estimation before you calculate. Comparing your prediction with the result is the fastest way to find out whether you understand the idea or only the words.

Example 3 — an exam-style question

Typical questions about Maximum likelihood estimation ask you to (a) state it precisely, (b) apply it to given data, and (c) explain a limitation. Practise writing all three answers in under five minutes; the third part is what separates a full-mark answer from an average one.

Applications of Maximum likelihood estimation

In research
Maximum likelihood estimation appears in mathematics research whenever the underlying quantities have to be modelled precisely. Papers usually cite it as a starting assumption and then explore where it breaks down.
In technology and industry
Engineering practice reuses Maximum likelihood estimation in design rules, simulations and safety margins. Knowing the idea lets you read a specification sheet and understand why the numbers look the way they do.
In the classroom
Maximum likelihood estimation is common in secondary-school and first-year university syllabi. It links to neighbouring topics M-estimators, Maximum likelihood estimation, Probability distribution fitting, so understanding it makes those chapters shorter.
In everyday life
Look for Maximum likelihood estimation outside the textbook — in sport, cooking, traffic, electronics or the sky above you. An example you found yourself is remembered far longer than one you were given.
Ask Teacher Smith questions about this articleOpens your AI tutor with a question about “Maximum likelihood estimation” →

Affiliate

Preply — study more efficiently by working with a personal tutor. 50% off.

How to study Maximum likelihood estimation in 20 minutes

  1. Read the reference excerpt below once, without taking notes.
  2. Close the page and write down what Maximum likelihood estimation means in your own words.
  3. Compare your version with the excerpt and mark what you missed.
  4. Work through the three examples above with pen and paper.
  5. Explain Maximum likelihood estimation out loud to somebody else — or to Teacher Smith in the lgStudy chat.

Frequently asked questions

What is Maximum likelihood estimation in simple terms?

In statistics, maximum likelihood estimation (MLE) is a method of estimating the parameters of an assumed probability distribution, given some observed data. This is achieved by maximizing a likelihood function so that, under the assumed statistical model, the observed data is most probable.

Why does Maximum likelihood estimation matter?

Because it connects several mathematics ideas at once: it gives you a definition you can apply, a quantity you can calculate, and a way to check whether a result is plausible.

How should I study Maximum likelihood estimation?

Read the excerpt, restate it from memory, then work through the examples and applications listed on this page. The five-step study plan above takes about twenty minutes.

What does this page cover?

It gives you a compact reference excerpt plus original lgStudy explanations, examples, applications and study material on Maximum likelihood estimation.

Tags

  • M-estimators
  • Maximum likelihood estimation
  • Probability distribution fitting

Keep exploring