WPGMA (Weighted Pair Group Method with Arithmetic Mean) is a simple agglomerative (bottom-up) hierarchical clustering method, generally attributed to Sokal and Michener. The WPGMA method is similar to its unweighted variant, the UPGMA method.
Algorithm The WPGMA algorithm constructs a rooted tree (dendrogram) that reflects the structure present in a pairwise distance matrix (or a similarity matrix). At each step, the nearest two clusters, say i {\displaystyle i} and j {\displaystyle j} , are combined into a higher-level cluster i ∪ j {\displaystyle i\cup j} . Then, its distance to another cluster k {\displaystyle k} is simply the arithmetic mean of the average distances between members of k {\displaystyle k} and i {\displaystyle i} and k {\displaystyle k} and j {\displaystyle j} :
d ( i ∪ j ) , k = d i , k + d j , k 2 {\displaystyle d_{(i\cup j),k}={\frac {d_{i,k}+d_{j,k}}{2}}}
The WPGMA algorithm produces rooted dendrograms and requires a constant-rate assumption: it produces an ultrametric tree in which the distances from the root to every branch tip are equal. This ultrametricity assumption is called the molecular clock when the tips involve DNA, RNA and protein data.
Working example This working example is based on a JC69 genetic distance matrix computed from the 5S ribosomal RNA sequence alignment of five bacteria: Bacillus subtilis ( a {\displaystyle a} ), Bacillus stearothermophilus ( b {\displaystyle b} ), Lactobacillus viridescens ( c {\displaystyle c} ), Acholeplasma modicum ( d {\displaystyle d} ), and Micrococcus luteus ( e {\displaystyle e} ).
First step First clustering Let us assume that we have five elements ( a , b , c , d , e ) {\displaystyle (a,b,c,d,e)} and the following matrix D 1 {\displaystyle D_{1}} of pairwise distances between them :
In this example, D 1 ( a , b ) = 17 {\displaystyle D_{1}(a,b)=17} is the smallest value of D 1 {\displaystyle D_{1}} , so we join elements a {\displaystyle a} and b {\displaystyle b} .
First branch length estimation Let u {\displaystyle u} denote the node to which a {\displaystyle a} and b {\displaystyle b} are now connected. Setting δ ( a , u ) = δ ( b , u ) = D 1 ( a , b ) / 2 {\displaystyle \delta (a,u)=\delta (b,u)=D_{1}(a,b)/2} ensures that elements a {\displaystyle a} and b {\displaystyle b} are equidistant from u {\displaystyle u} . This corresponds to the expectation of the ultrametricity hypothesis. The branches joining a {\displaystyle a} and b {\displaystyle b} to u {\displaystyle u} then have lengths δ ( a , u ) = δ ( b , u ) = 17 / 2 = 8.5 {\displaystyle \delta (a,u)=\delta (b,u)=17/2=8.5} (see the final dendrogram)
First distance matrix update We then proceed to update the initial distance matrix D 1 {\displaystyle D_{1}} into a new distance matrix D 2 {\displaystyle D_{2}} (see below), reduced in size by one row and one column because of the clustering of a {\displaystyle a} with b {\displaystyle b} . Bold values in D 2 {\displaystyle D_{2}} correspond to the new distances, calculated by averaging distances between each element of the first cluster ( a , b ) {\displaystyle (a,b)} and each of the remaining elements:
… excerpt ends here. Continue reading the full article.




