In computer science, Hirschberg's algorithm, named after its inventor, Dan Hirschberg, is a dynamic programming algorithm that finds the optimal sequence alignment between two strings. Optimality is measured with the Levenshtein distance, defined to be the sum of the costs of insertions, replacements, deletions, and null actions needed to change one string into the other. Hirschberg's algorithm is simply described as a more space-efficient version of the Needleman–Wunsch algorithm that uses dynamic programming. Hirschberg's algorithm is commonly used in computational biology to find maximal global alignments of DNA and protein sequences.
Algorithm information Hirschberg's algorithm is a generally applicable algorithm for optimal sequence alignment. BLAST and FASTA are suboptimal heuristics. If X {\displaystyle X} and Y {\displaystyle Y} are strings, where length ( X ) = n {\displaystyle \operatorname {length} (X)=n} and length ( Y ) = m {\displaystyle \operatorname {length} (Y)=m} , the Needleman–Wunsch algorithm finds an optimal alignment in O ( n m ) {\displaystyle O(nm)} time, using O ( n m ) {\displaystyle O(nm)} space. Hirschberg's algorithm is a clever modification of the Needleman–Wunsch Algorithm, which still takes O ( n m ) {\displaystyle O(nm)} time, but needs only O ( min { n , m } ) {\displaystyle O(\min\{n,m\})} space and is much faster in practice. One application of the algorithm is finding sequence alignments of DNA or protein sequences. It is also a space-efficient way to calculate the longest common subsequence between two sets of data such as with the common diff tool. The Hirschberg algorithm can be derived from the Needleman–Wunsch algorithm by observing that:
one can compute the optimal alignment score by only storing the current and previous row of the Needleman–Wunsch score matrix; if ( Z , W ) = NW ( X , Y ) {\displaystyle (Z,W)=\operatorname {NW} (X,Y)} is the optimal alignment of ( X , Y ) {\displaystyle (X,Y)} , and X = X l + X r {\displaystyle X=X^{l}+X^{r}} is an arbitrary partition of X {\displaystyle X} , there exists a partition Y l + Y r {\displaystyle Y^{l}+Y^{r}} of Y {\displaystyle Y} such that NW ( X , Y ) = NW ( X l , Y l ) + NW ( X r , Y r ) {\displaystyle \operatorname {NW} (X,Y)=\operatorname {NW} (X^{l},Y^{l})+\operatorname {NW} (X^{r},Y^{r})} .
Algorithm description
X i {\displaystyle X_{i}} denotes the i-th character of X {\displaystyle X} , where 1 ⩽ i ⩽ length ( X ) {\displaystyle 1\leqslant i\leqslant \operatorname {length} (X)} . X i : j {\displaystyle X_{i:j}} denotes a substring of size j − i + 1 {\displaystyle j-i+1} , ranging from the i-th to the j-th character of X {\displaystyle X} . rev ( X ) {\displaystyle \operatorname {rev} (X)} is the reversed version of X {\displaystyle X} .
… excerpt ends here. Continue reading the full article.
