Preply — Study more efficiently by working with a personal tutor. Get 50% off.Affiliate

Wikipedia

Hardy–Weinberg principle

Hardy–Weinberg principle

In population genetics, the Hardy–Weinberg principle, also known as the Hardy–Weinberg equilibrium, model, theorem, or law, states that allele and genotype frequencies in a population will remain constant from generation to generation in the absence of other evolutionary influences. These influences include genetic drift, mate choice, assortative mating, natural selection, sexual selection, mutation, gene flow, meiotic drive, genetic hitchhiking, population bottleneck, founder effect, inbreeding, and outbreeding depression.

Derivation In the simplest case of a single locus with two alleles denoted A and a with frequencies f(A) = p and f(a) = q, respectively, the expected genotype frequencies under random mating are f(AA) = p2 for the AA homozygotes, f(aa) = q2 for the aa homozygotes, and f(Aa) = 2pq for the heterozygotes. In the absence of selection, mutation, genetic drift, or other forces, allele frequencies p and q are constant between generations, so equilibrium is reached. The principle is named after G. H. Hardy and Wilhelm Weinberg, who first demonstrated it mathematically. Hardy's paper was focused on debunking the view that a dominant allele would automatically tend to increase in frequency (a view possibly based on a misinterpreted question at a lecture). Today, tests for Hardy–Weinberg genotype frequencies are used primarily to test for population stratification and other forms of non-random mating. Consider a population of monoecious diploids, where each organism produces male and female gametes at equal frequency, and has two alleles at each gene locus. We assume that the population is so large that it can be treated as infinite. Organisms reproduce by random union of gametes (the "gene pool" population model). A locus in this population has two alleles, A and a, that occur with initial frequencies f0(A) = p and f0(a) = q, respectively. The allele frequencies at each generation are obtained by pooling together the alleles from each genotype of the same generation according to the expected contribution from the homozygote and heterozygote genotypes, which are 1 and 1/2, respectively:

The different ways to form genotypes for the next generation can be shown in a Punnett square, where the proportion of each genotype is equal to the product of the row and column allele frequencies from the current generation.

The sum of the entries is p2 + 2pq + q2 = 1, as the genotype frequencies must sum to one. Note again that as p + q = 1, the binomial expansion of (p + q)2 = p2 + 2pq + q2 = 1 gives the same relationships. Summing the elements of the Punnett square or the binomial expansion, we obtain the expected genotype proportions among the offspring after a single generation:

These frequencies define the Hardy–Weinberg equilibrium. It should be mentioned that the genotype frequencies after the first generation need not equal the genotype frequencies from the initial generation, e.g. f1(AA) ≠ f0(AA). However, the genotype frequencies for all future times will equal the Hardy–Weinberg frequencies, e.g. ft(AA) = f1(AA) for t > 1. This follows since the genotype frequencies of the next generation depend only on the allele frequencies of the current generation which, as calculated by equations (1) and (2), are preserved from the initial generation:

f 1 ( A ) = f 1 ( AA ) + 1 2 f 1 ( Aa ) = p 2 + p q = p ( p + q ) = p = f 0 ( A ) f 1 ( a ) = f 1 ( aa ) + 1 2 f 1 ( Aa ) = q 2 + p q = q ( p + q ) = q = f 0 ( a ) {\displaystyle {\begin{aligned}f_{1}({\text{A}})&=f_{1}({\text{AA}})+{\tfrac {1}{2}}f_{1}({\text{Aa}})=p^{2}+pq=p(p+q)=p=f_{0}({\text{A}})\\f_{1}({\text{a}})&=f_{1}({\text{aa}})+{\tfrac {1}{2}}f_{1}({\text{Aa}})=q^{2}+pq=q(p+q)=q=f_{0}({\text{a}})\end{aligned}}}

For the more general case of dioecious diploids [organisms are either male or female] that reproduce by random mating of individuals, it is necessary to calculate the genotype frequencies from the nine possible matings between each parental genotype (AA, Aa, and aa) in either sex, weighted by the expected genotype contributions of each such mating. Equivalently, one considers the six unique diploid–diploid combinations:

[ ( AA , AA ) , ( AA , Aa ) , ( AA , aa ) , ( Aa , Aa ) , ( Aa , aa ) , ( aa , aa ) ] {\displaystyle \left[({\text{AA}},{\text{AA}}),({\text{AA}},{\text{Aa}}),({\text{AA}},{\text{aa}}),({\text{Aa}},{\text{Aa}}),({\text{Aa}},{\text{aa}}),({\text{aa}},{\text{aa}})\right]}

and constructs a Punnett square for each, so as to calculate its contribution to the next generation's genotypes. These contributions are weighted according to the probability of each diploid–diploid combination, which follows a multinomial distribution with k = 3. For example, the probability of the mating combination (AA,aa) is 2 ft(AA)ft(aa) and it can only result in the Aa genotype: [0,1,0]. Overall, the resulting genotype frequencies are calculated as:

[ f t + 1 ( AA ) , f t + 1 ( Aa ) , f t + 1 ( aa ) ] = = f t ( AA ) f t ( AA ) [ 1 , 0 , 0 ] + 2 f t ( AA ) f t ( Aa ) [ 1 2 , 1 2 , 0 ] + 2 f t ( AA ) f t ( aa ) [ 0 , 1 , 0 ] + f t ( Aa ) f t ( Aa ) [ 1 4 , 1 2 , 1 4 ] + 2 f t ( Aa ) f t ( aa ) [ 0 , 1 2 , 1 2 ] + f t ( aa ) f t ( aa ) [ 0 , 0 , 1 ] = [ ( f t ( AA ) + 1 2 f t ( Aa ) ) 2 , 2 ( f t ( AA ) + 1 2 f t ( Aa ) ) ( f t ( aa ) + 1 2 f t ( Aa ) ) , ( f t ( aa ) + 1 2 f t ( Aa ) ) 2 ] = [ f t ( A ) 2 , 2 f t ( A ) f t ( a ) , f t ( a ) 2 ] {\displaystyle {\begin{aligned}&\left[f_{t+1}({\text{AA}}),f_{t+1}({\text{Aa}}),f_{t+1}({\text{aa}})\right]=\\&\qquad =f_{t}({\text{AA}})f_{t}({\text{AA}})\left[1,0,0\right]+2f_{t}({\text{AA}})f_{t}({\text{Aa}})\left[{\tfrac {1}{2}},{\tfrac {1}{2}},0\right]+2f_{t}({\text{AA}})f_{t}({\text{aa}})\left[0,1,0\right]\\&\qquad \qquad +f_{t}({\text{Aa}})f_{t}({\text{Aa}})\left[{\tfrac {1}{4}},{\tfrac {1}{2}},{\tfrac {1}{4}}\right]+2f_{t}({\text{Aa}})f_{t}({\text{aa}})\left[0,{\tfrac {1}{2}},{\tfrac {1}{2}}\right]+f_{t}({\text{aa}})f_{t}({\text{aa}})\left[0,0,1\right]\\&\qquad =\left[\left(f_{t}({\text{AA}})+{\tfrac {1}{2}}f_{t}({\text{Aa}})\right)^{2},2\left(f_{t}({\text{AA}})+{\tfrac {1}{2}}f_{t}({\text{Aa}})\right)\left(f_{t}({\text{aa}})+{\tfrac {1}{2}}f_{t}({\text{Aa}})\right),\left(f_{t}({\text{aa}})+{\tfrac {1}{2}}f_{t}({\text{Aa}})\right)^{2}\right]\\&\qquad =\left[f_{t}({\text{A}})^{2},2f_{t}({\text{A}})f_{t}({\text{a}}),f_{t}({\text{a}})^{2}\right]\end{aligned}}}

As before, one can show that the allele frequencies at time t + 1 equal those at time t, and so, are constant in time. Similarly, the genotype frequencies depend only on the allele frequencies, and so, after time t = 1 are also constant in time. If in either monoecious or dioecious organisms, either the allele or genotype proportions are initially unequal in either sex, it can be shown that constant proportions are obtained after one generation of random mating. If dioecious organisms are heterogametic and the gene locus is located on the X chromosome, it can be shown that if the allele frequencies are initially unequal in the two sexes [e.g., XX females and XY males, as in humans], f′(a) in the heterogametic sex 'chases' f(a) in the homogametic sex of the previous generation, until an equilibrium is reached at the weighted average of the two initial frequencies.

Deviations from Hardy–Weinberg equilibrium The seven assumptions underlying Hardy–Weinberg equilibrium are as follows:

organisms are diploid only sexual reproduction occurs generations are nonoverlapping mating is random population size is infinitely large allele frequencies are equal in the sexes there is no migration, gene flow, admixture, mutation or selection Violations of the Hardy–Weinberg assumptions can cause deviations from expectation. How this affects the population depends on the assumptions that are violated.

Random mating. The HWP states the population will have the given genotypic frequencies (called Hardy–Weinberg proportions) after a single generation of random mating within the population. When the random mating assumption is violated, the population will not have Hardy–Weinberg proportions. A common cause of non-random mating is inbreeding, which causes an increase in homozygosity for all genes. If a population violates one of the following four assumptions, the population may continue to have Hardy–Weinberg proportions each generation, but the allele frequencies will change over time.

Selection, in general, causes allele frequencies to change, often quite rapidly. While directional selection eventually leads to the loss of all alleles except the favored one (unless one allele is dominant, in which case recessive alleles can survive at low frequencies), some forms of selection, such as balancing selection, lead to equilibrium without loss of alleles. Mutation will have a very subtle effect on allele frequencies through the introduction of new allele into a population. Mutation rates are of the order 10−4 to 10−8, and the change in allele frequency will be, at most, the same order. Recurrent mutation will maintain alleles in the population, even if there is strong selection against them. Migration genetically links two or more populations together. In general, allele frequencies will become more homogeneous among the populations. Some models for migration inherently include nonrandom mating (Wahlund effect, for example). For those models, the Hardy–Weinberg proportions will normally not be valid. Small population size can cause a random change in allele frequencies. This is due to a sampling effect, and is called genetic drift. Sampling effects are most important when the allele is present in a small number of copies. In real world genotype data, deviations from Hardy–Weinberg Equilibrium may be a sign of genotyping error.

Sex linkage Where the A gene is sex linked, the heterogametic sex (e.g., mammalian males; avian females) have only one copy of the gene (and are termed hemizygous), while the homogametic sex (e.g., human females) have two copies. The genotype frequencies at equilibrium are p and q for the heterogametic sex but p2, 2pq and q2 for the homogametic sex. For example, in humans red–green colorblindness is an X-linked recessive trait. In western European males, the trait affects about 1 in 12, (q = 0.083) whereas it affects about 1 in 200 females (0.005, compared to q2 = 0.007), very close to Hardy–Weinberg proportions. If a population is brought together with males and females with a different allele frequency in each subpopulation (males or females), the allele frequency of the male population in the next generation will follow that of the female population because each son receives its X chromosome from its mother. The population converges on equilibrium very quickly.

Generalizations The simple derivation above can be generalized for more than two alleles and polyploidy.

Generalization for more than two alleles

Consider an extra allele frequency, r. The two-allele case is the binomial expansion of (p + q)2, and thus the three-allele case is the trinomial expansion of (p + q + r)2.

( p + q + r ) 2 = p 2 + q 2 + r 2 + 2 p q + 2 p r + 2 q r {\displaystyle (p+q+r)^{2}=p^{2}+q^{2}+r^{2}+2pq+2pr+2qr\,}

More generally, consider the alleles A1, ..., An given by the allele frequencies p1 to pn;

( p 1 + ⋯ + p n ) 2 {\displaystyle (p_{1}+\cdots +p_{n})^{2}\,}

giving for all homozygotes:

f ( A i A i ) = p i 2 {\displaystyle f(A_{i}A_{i})=p_{i}^{2}\,}

and for all heterozygotes:

f ( A i A j ) = 2 p i p j {\displaystyle f(A_{i}A_{j})=2p_{i}p_{j}\,}

Generalization for polyploidy The Hardy–Weinberg principle may also be generalized to polyploid systems, that is, for organisms that have more than two copies of each chromosome. Consider again only two alleles. The diploid case is the binomial expansion of:

( p + q ) 2 {\displaystyle (p+q)^{2}\,}

and therefore the polyploid case is the binomial expansion of:

( p + q ) c {\displaystyle (p+q)^{c}\,}

where c is the ploidy, for example with tetraploid (c = 4):

Whether the organism is a 'true' tetraploid or an amphidiploid will determine how long it will take for the population to reach Hardy–Weinberg equilibrium.

Complete generalization For n {\displaystyle n} distinct alleles in c {\displaystyle c} -ploids, the genotype frequencies in the Hardy–Weinberg equilibrium are given by individual terms in the multinomial expansion of ( p 1 + ⋯ + p n ) c {\displaystyle (p_{1}+\cdots +p_{n})^{c}} :

( p 1 + ⋯ + p n ) c = ∑ k 1 , … , k n ∈ N : k 1 + ⋯ + k n = c ( c k 1 , … , k n ) p 1 k 1 ⋯ p n k n {\displaystyle (p_{1}+\cdots +p_{n})^{c}=\sum _{k_{1},\ldots ,k_{n}\ \in \mathbb {N} :k_{1}+\cdots +k_{n}=c}{c \choose k_{1},\ldots ,k_{n}}p_{1}^{k_{1}}\cdots p_{n}^{k_{n}}}

Significance tests for deviation Testing deviation from the HWP is generally performed using Pearson's chi-squared test, using the observed genotype frequencies obtained from the data and the expected genotype frequencies obtained using the HWP. For systems where there are large numbers of alleles, this may result in data with many empty possible genotypes and low genotype counts, because there are often not enough individuals present in the sample to adequately represent all genotype classes. If this is the case, then the asymptotic assumption of the chi-squared distribution, will no longer hold, and it may be necessary to use a form of Fisher's exact test, which requires a computer to solve. More recently a number of MCMC methods of testing for deviations from HWP have been proposed (Guo & Thompson, 1992; Wigginton et al. 2005)

Example chi-squared test for deviation This data is from E. B. Ford (1971) on the scarlet tiger moth, for which the phenotypes of a sample of the population were recorded. Genotype–phenotype distinction is assumed to be negligibly small. The null hypothesis is that the population is in Hardy–Weinberg proportions, and the alternative hypothesis is that the population is not in Hardy–Weinberg proportions.

From this, allele frequencies can be calculated:

p = 2 × o b s ( AA ) + o b s ( Aa ) 2 × ( o b s ( AA ) + o b s ( Aa ) + o b s ( aa ) ) = 2 × 1469 + 138 2 × ( 1469 + 138 + 5 ) = 3076 3224 = 0.954 {\displaystyle {\begin{aligned}p&={2\times \mathrm {obs} ({\text{AA}})+\mathrm {obs} ({\text{Aa}}) \over 2\times (\mathrm {obs} ({\text{AA}})+\mathrm {obs} ({\text{Aa}})+\mathrm {obs} ({\text{aa}}))}\\\\&={2\times 1469+138 \over 2\times (1469+138+5)}\\\\&={3076 \over 3224}\\\\&=0.954\end{aligned}}}

and

q = 1 − p = 1 − 0.954 = 0.046 {\displaystyle {\begin{aligned}q&=1-p\\&=1-0.954\\&=0.046\end{aligned}}}

So the Hardy–Weinberg expectation is:

E x p ( AA ) = p 2 n = 0.954 2 × 1612 = 1467.4 E x p ( Aa ) = 2 p q n = 2 × 0.954 × 0.046 × 1612 = 141.2 E x p ( aa ) = q 2 n = 0.046 2 × 1612 = 3.4 {\displaystyle {\begin{aligned}\mathrm {Exp} ({\text{AA}})&=p^{2}n=0.954^{2}\times 1612=1467.4\\\mathrm {Exp} ({\text{Aa}})&=2pqn=2\times 0.954\times 0.046\times 1612=141.2\\\mathrm {Exp} ({\text{aa}})&=q^{2}n=0.046^{2}\times 1612=3.4\end{aligned}}}

Pearson's chi-squared test states:

χ 2 = ∑

Tags

  • Classical genetics
  • Population genetics
  • Sexual selection
  • Statistical genetics