ArticleslgStudy

mathematics

Zipf's law

Zipf's law is a mathematics topic covered in the lgStudy science library. This page brings together a partial reference excerpt, illustrations, worked examples, real-world applications and a short study plan, so you can understand Zipf's law rather than just read about it. In short: Zipf's law () is an empirical law stating that when a set of measured values is sorted in decreasing order, the value of the n-th entry is often approximately inversely proportional to n. The best-known instance of Zipf's law applies to the frequency distribution of words in a text or corpus of natural language: w o r d f r e q u e n c y ∝ 1 w o r d r a n k . {\displaystyle \ {\mathsf {word\ frequency}}\ \propto \ {…

Zipf's law — main illustration
Zipf's law — illustration

Key takeaways

  • Zipf's law belongs to mathematics; place it in that map before memorising details.
  • Learn the definition first, then one example that makes the definition concrete.
  • Connect Zipf's law to a quantity you can measure, compute or draw — that is where exam questions come from.
  • Reproduce the core statement of Zipf's law from memory before moving on to harder problems.

Reference excerpt

Zipf's law () is an empirical law stating that when a set of measured values is sorted in decreasing order, the value of the n-th entry is often approximately inversely proportional to n. The best-known instance of Zipf's law applies to the frequency distribution of words in a text or corpus of natural language:

w o r d f r e q u e n c y ∝ 1 w o r d r a n k . {\displaystyle \ {\mathsf {word\ frequency}}\ \propto \ {\frac {1}{\ {\mathsf {word\ rank}}\ }}~.}

It is usually found that the most common word occurs approximately twice as often as the next common one, three times as often as the third most common, and so on. For example, in the Brown Corpus of American English text, the word "the" is the most frequently occurring word, and by itself accounts for nearly 7% of all word occurrences (69,971 out of slightly over 1 million). True to Zipf's law, the second-place word "of" accounts for slightly over 3.5% of words (36,411 occurrences), followed by "and" (28,852). It is often used in the following form, called the Zipf-Mandelbrot law:

f r e q u e n c y ∝ 1 ( r a n k + b ) a {\displaystyle \ {\mathsf {frequency}}\ \propto \ {\frac {1}{\ \left(\ {\mathsf {rank}}+b\ \right)^{a}\ }}\ } where a {\displaystyle \ a\ } and b {\displaystyle \ b\ } are fitted parameters, with a ≈ 1 {\displaystyle \ a\approx 1} , and b ≈ 2.7 {\displaystyle \ b\approx 2.7~} . This law is named after the American linguist George Kingsley Zipf, and is still an important concept in quantitative linguistics. It has been found to apply to many other types of data studied in the physical and social sciences. In mathematical statistics, the concept has been formalized as the Zipfian distribution: A family of related discrete probability distributions whose rank-frequency distribution is an inverse power law relation. They are related to Benford's law and the Pareto distribution. Some sets of time-dependent empirical data deviate somewhat from Zipf's law. Such empirical distributions are said to be quasi-Zipfian.

History In 1913, the German physicist Felix Auerbach observed an inverse proportionality between the population sizes of cities, and their ranks when sorted by decreasing order of that variable. Zipf's law had been discovered before Zipf, first by the French stenographer Jean-Baptiste Estoup in 1916, and also by G. Dewey in 1923, and by E. Condon in 1928. The same relation for frequencies of words in natural language texts was observed by George Zipf in 1932, but he never claimed to have originated it. In fact, Zipf did not like mathematics. In his 1932 publication, the author speaks with disdain about mathematical involvement in linguistics, a.o. ibidem, p. 21:

... let me say here for the sake of any mathematician who may plan to formulate the ensuing data more exactly, the ability of the highly intense positive to become the highly intense negative, in my opinion, introduces the devil into the formula in the form of − i . {\displaystyle \ {\sqrt {-i\;}}~.}

The only mathematical expression Zipf used looks like ab2 = constant, which he "borrowed" from Alfred J. Lotka's 1926 publication. The same relationship was found to occur in many other contexts, and for other variables besides frequency. For example, when corporations are ranked by decreasing size, their sizes are found to be inversely proportional to the rank. The same relation is found for personal incomes (where it is called Pareto principle), number of people watching the same TV channel, notes in music, cells transcriptomes, and more. In 1957 George A. Miller proposed that a power law emerges even in randomly generated texts. and in 1992 bioinformatician Wentian Li published a proof that the power law form of Zipf's law was a byproduct of ordering words by rank.

Formal definition

Formally, the Zipf distribution on N elements assigns to the element of rank k (counting from 1) the probability:

… excerpt ends here. Continue reading the full article.

Illustrations

Zipf's law: A plot of the frequency of each word as a function of its frequency rank for two English language texts: Culpeper's Complete Herbal (1652) and H. G. Wells's The War of the Worlds (1898) in a log-log scale. The dashed line is the ideal law 
  
    
      
        y
        ∝
        
          
            1
            x
          
        
      
    
    {\textstyle y\propto {\frac {1}{x}}}
  
.
A plot of the frequency of each word as a function of its frequency rank for two English language texts: Culpeper's Complete Herbal (1652) and H. G. Wells's The War of the Worlds (1898) in a log-log scale. The dashed line is the ideal law y ∝ 1 x {\textstyle y\propto {\frac {1}{x}}} .
Zipf's law illustration
Zipf's law illustration
Zipf's law: Zipf's law plot for the first 10 million words in 30 Wikipedias (as of October 2015) in a log-log scale
Zipf's law plot for the first 10 million words in 30 Wikipedias (as of October 2015) in a log-log scale
Zipf's law illustration

Worked examples

Example 1 — a first encounter with Zipf's law

Start with the simplest possible case. Write down what Zipf's law claims or describes in one sentence, then invent the smallest concrete situation in which that sentence is true. In mathematics, the smallest case is usually a single object, a single equation or a single measurement. Check that every symbol or term in your sentence has a meaning in that case.

Example 2 — changing one variable

Take the situation from Example 1 and change exactly one quantity: double it, halve it, or set it to zero. Predict what should happen to Zipf's law before you calculate. Comparing your prediction with the result is the fastest way to find out whether you understand the idea or only the words.

Example 3 — an exam-style question

Typical questions about Zipf's law ask you to (a) state it precisely, (b) apply it to given data, and (c) explain a limitation. Practise writing all three answers in under five minutes; the third part is what separates a full-mark answer from an average one.

Applications of Zipf's law

In research
Zipf's law appears in mathematics research whenever the underlying quantities have to be modelled precisely. Papers usually cite it as a starting assumption and then explore where it breaks down.
In technology and industry
Engineering practice reuses Zipf's law in design rules, simulations and safety margins. Knowing the idea lets you read a specification sheet and understand why the numbers look the way they do.
In the classroom
Zipf's law is common in secondary-school and first-year university syllabi. It links to neighbouring topics 1949 introductions, Bibliometrics, Computational linguistics, so understanding it makes those chapters shorter.
In everyday life
Look for Zipf's law outside the textbook — in sport, cooking, traffic, electronics or the sky above you. An example you found yourself is remembered far longer than one you were given.
Ask Teacher Smith questions about this articleOpens your AI tutor with a question about “Zipf's law” →

Affiliate

Preply — study more efficiently by working with a personal tutor. 50% off.

How to study Zipf's law in 20 minutes

  1. Read the reference excerpt below once, without taking notes.
  2. Close the page and write down what Zipf's law means in your own words.
  3. Compare your version with the excerpt and mark what you missed.
  4. Work through the three examples above with pen and paper.
  5. Explain Zipf's law out loud to somebody else — or to Teacher Smith in the lgStudy chat.

Frequently asked questions

What is Zipf's law in simple terms?

Zipf's law () is an empirical law stating that when a set of measured values is sorted in decreasing order, the value of the n-th entry is often approximately inversely proportional to n. The best-known instance of Zipf's law applies to the frequency distribution of words in a text or corpus of nat…

Why does Zipf's law matter?

Because it connects several mathematics ideas at once: it gives you a definition you can apply, a quantity you can calculate, and a way to check whether a result is plausible.

How should I study Zipf's law?

Read the excerpt, restate it from memory, then work through the examples and applications listed on this page. The five-step study plan above takes about twenty minutes.

What does this page cover?

It gives you a compact reference excerpt plus original lgStudy explanations, examples, applications and study material on Zipf's law.

Tags

  • 1949 introductions
  • Bibliometrics
  • Computational linguistics
  • Corpus linguistics
  • Discrete distributions
  • Empirical laws
  • Eponymous rules
  • Power laws
  • Quantitative linguistics
  • Statistical laws
  • Tails of probability distributions

Keep exploring