Preply — Study more efficiently by working with a personal tutor. Get 50% off.Affiliate

Wikipedia

Eigenvalues and eigenvectors

In linear algebra, an eigenvector ( EYE-gən-) or characteristic vector is a (nonzero) vector that has its direction unchanged (or reversed) by a given linear transformation. More precisely, an eigenvector v {\displaystyle \mathbf {v} } of a linear transformation T {\displaystyle T} is scaled by a constant factor λ {\displaystyle \lambda } when the linear transformation is applied to it: ⁠ T v = λ v {\displaystyle T\mathbf {v} =\lambda \mathbf {v} } ⁠. The corresponding eigenvalue, characteristic value, or characteristic root is the multiplying factor λ {\displaystyle \lambda } (possibly a negative or complex number). Geometrically, vectors are multi-dimensional quantities with magnitude and direction, often pictured as arrows. A linear transformation rotates, stretches, or shears the vectors upon which it acts. A linear transformation's eigenvectors are those vectors that are only stretched or shrunk, with neither rotation nor shear. The corresponding eigenvalue is the factor by which an eigenvector is stretched or shrunk. If the eigenvalue is negative, then the eigenvector's direction is reversed. The eigenvectors and eigenvalues of a linear transformation serve to characterize it, and so they play important roles in all areas where linear algebra is applied, from geology to quantum mechanics. In particular, it is often the case that a system is represented by a linear transformation whose outputs are fed as inputs to the same transformation (feedback). In such an application, the largest eigenvalue is of particular importance, because it governs the long-term behavior of the system after many applications of the linear transformation, and the associated eigenvector is the steady state of the system.

Matrices For an n × n {\displaystyle n{\times }n} matrix ⁠ A {\displaystyle A} ⁠ and a nonzero ⁠ n {\displaystyle n} ⁠-vector ⁠ v {\displaystyle \mathbf {v} } ⁠, if multiplying ⁠ A {\displaystyle A} ⁠ by v {\displaystyle \mathbf {v} } (denoted ⁠ A v {\displaystyle A\mathbf {v} } ⁠) simply scales v {\displaystyle \mathbf {v} } by a factor ⁠ λ {\displaystyle \lambda } ⁠, where ⁠ λ {\displaystyle \lambda } ⁠ is a scalar, then v {\displaystyle \mathbf {v} } is called an eigenvector of ⁠ A {\displaystyle A} ⁠, and ⁠ λ {\displaystyle \lambda } ⁠ is the corresponding eigenvalue. This relationship can be expressed as: ⁠ A v = λ v {\displaystyle A\mathbf {v} =\lambda \mathbf {v} } ⁠. Given an ⁠ n {\displaystyle n} ⁠-dimensional vector space and a choice of basis, there is a direct correspondence between linear transformations from the vector space into itself and ⁠ n × n {\displaystyle n\times n} ⁠ square matrices. Hence, in a finite-dimensional vector space, it is equivalent to define eigenvalues and eigenvectors using either the language of linear transformations, or the language of matrices.

Overview Eigenvalues and eigenvectors feature prominently in the analysis of linear transformations. The prefix eigen- is adopted from the German eigen (cognate with the English word own) for 'proper', 'characteristic', 'own'. Originally used to study principal axes of the rotational motion of rigid bodies, eigenvalues and eigenvectors have a wide range of applications, for example in stability analysis, vibration analysis, atomic orbitals, facial recognition, and matrix diagonalization. In essence, an eigenvector v of a linear transformation T is a nonzero vector that, when T is applied to it, does not change direction. Applying T to the eigenvector only scales the eigenvector by the scalar value λ, called an eigenvalue. This condition can be written as the equation

T ( v ) = λ v , {\displaystyle T(\mathbf {v} )=\lambda \mathbf {v} ,}

referred to as the eigenvalue equation or eigenequation. In general, λ may be any scalar. For example, λ may be negative, in which case the eigenvector reverses direction as part of the scaling, or it may be zero, or complex.

The example here, based on the Mona Lisa, provides a simple illustration. Each point on the painting can be represented as a vector pointing from the center of the painting to that point. The linear transformation in this example is called a shear mapping. Points in the top half are moved to the right, and points in the bottom half are moved to the left, proportional to how far they are from the horizontal axis that goes through the middle of the painting. The vectors pointing to each point in the original image are therefore tilted right or left, and made longer or shorter by the transformation. Points along the horizontal axis do not move at all when this transformation is applied. Therefore, any vector that points directly to the right or left with no vertical component is an eigenvector of this transformation, because the mapping does not change its direction. Moreover, these eigenvectors all have an eigenvalue equal to one, because the mapping does not change their length either.

Linear transformations can take many different forms, mapping vectors in a variety of vector spaces, so the eigenvectors can also take many forms. For example, the linear transformation could be a differential operator like ⁠ d d x {\displaystyle {\tfrac {d}{dx}}} ⁠, in which case the eigenvectors are functions called eigenfunctions that are scaled by that differential operator, such as

d d x e λ x = λ e λ x . {\displaystyle {\frac {d}{dx}}e^{\lambda x}=\lambda e^{\lambda x}.}

Alternatively, the linear transformation could take the form of an n × n matrix, in which case the eigenvectors are n × 1 matrices. If the linear transformation is expressed in the form of an n × n matrix A, then the eigenvalue equation for a linear transformation above can be rewritten as the matrix multiplication

A v = λ v , {\displaystyle A\mathbf {v} =\lambda \mathbf {v} ,}

where the eigenvector v is an n × 1 matrix. For a matrix, eigenvalues and eigenvectors can be used to decompose the matrix; for example, by diagonalizing it. Eigenvalues and eigenvectors give rise to many closely related mathematical concepts, and the prefix eigen- is applied liberally when naming them:

The set of all eigenvectors of a linear transformation, each paired with its corresponding eigenvalue, is called the eigensystem of that transformation. The set of all eigenvectors of T corresponding to the same eigenvalue, together with the zero vector, is called an eigenspace, or the characteristic space of T associated with that eigenvalue. If a set of eigenvectors of T forms a basis of the domain of T, then this basis is called an eigenbasis. Additional animations of eigenvectors and eigenvalues in two dimensions, including symmetric transformations and transformations with complex eigenvalues, are available at Wikimedia Commons.

History Eigenvalues are often introduced in the context of linear algebra or matrix theory. Historically, however, they arose in the study of quadratic forms and differential equations. In the 18th century, Leonhard Euler studied the rotational motion of a rigid body, and discovered the importance of the principal axes. Joseph-Louis Lagrange realized that the principal axes are the eigenvectors of the inertia matrix. In the early 19th century, Augustin-Louis Cauchy saw how their work could be used to classify the quadric surfaces, and generalized it to arbitrary dimensions. Cauchy also coined the term racine caractéristique (characteristic root), for what is now called eigenvalue; his term survives in characteristic equation. Later, Joseph Fourier used the work of Lagrange and Pierre-Simon Laplace to solve the heat equation by separation of variables in his 1822 treatise The Analytic Theory of Heat (Théorie analytique de la chaleur). Charles-François Sturm elaborated on Fourier's ideas further, and brought them to the attention of Cauchy, who combined them with his own ideas and arrived at the fact that real symmetric matrices have real eigenvalues. This was extended by Charles Hermite in 1855 to what are now called Hermitian matrices. Around the same time, Francesco Brioschi proved that the eigenvalues of orthogonal matrices lie on the unit circle, and Alfred Clebsch found the corresponding result for skew-symmetric matrices. Finally, Karl Weierstrass clarified an important aspect in the stability theory started by Laplace, by realizing that defective matrices can cause instability. In the meantime, Joseph Liouville studied eigenvalue problems similar to those of Sturm; the discipline that grew out of their work is now called Sturm–Liouville theory. Schwarz studied the first eigenvalue of Laplace's equation on general domains towards the end of the 19th century, while Poincaré studied Poisson's equation a few years later. At the start of the 20th century, David Hilbert studied the eigenvalues of integral operators by viewing the operators as infinite matrices. He was the first to use the German word eigen, which means "own", to denote eigenvalues and eigenvectors in 1904, though he may have been following a related usage by Hermann von Helmholtz. For some time, the standard term in English was "proper value", but the more distinctive term "eigenvalue" is the standard today. The first numerical algorithm for computing eigenvalues and eigenvectors appeared in 1929, when Richard von Mises published the power method. One of the most popular methods today, the QR algorithm, was proposed independently by John G. F. Francis and Vera Kublanovskaya in 1961.

Eigenvalues and eigenvectors of a matrix

Eigenvalues and eigenvectors are often introduced to students in the context of linear algebra courses focused on matrices. Furthermore, linear transformations over a finite-dimensional vector space can be represented using matrices, which is especially common in numerical and computational applications.

Consider two ⁠ n {\displaystyle n} ⁠-dimensional vectors that are formed as a list of ⁠ n {\displaystyle n} ⁠ scalars, such as the three-dimensional vectors

x = [ 1 − 3 4 ] and y = [ − 20 60 − 80 ] . {\displaystyle \mathbf {x} ={\begin{bmatrix}1\\-3\\4\end{bmatrix}}\quad {\mbox{and}}\quad \mathbf {y} ={\begin{bmatrix}-20\\60\\-80\end{bmatrix}}.}

These vectors are said to be scalar multiples of each other, or parallel, or collinear, if there is a scalar ⁠ λ {\displaystyle \lambda } ⁠ such that

y = λ x . {\displaystyle \mathbf {y} =\lambda \mathbf {x} .}

In this example, ⁠ λ = − 20 {\displaystyle \lambda =-20} ⁠. Now consider the linear transformation of ⁠ n {\displaystyle n} ⁠-dimensional vectors defined by an ⁠ n × n {\displaystyle n\times n} ⁠ matrix ⁠ A {\displaystyle A} ⁠:

A v = w , {\displaystyle A\mathbf {v} =\mathbf {w} ,}

or

[ A 11 A 12 ⋯ A 1 n A 21 A 22 ⋯ A 2 n ⋮ ⋮ ⋱ ⋮ A n 1 A n 2 ⋯ A n n ] [ v 1 v 2 ⋮ v n ] = [ w 1 w 2 ⋮ w n ] {\displaystyle {\begin{bmatrix}A_{11}&A_{12}&\cdots &A_{1n}\\A_{21}&A_{22}&\cdots &A_{2n}\\\vdots &\vdots &\ddots &\vdots \\A_{n1}&A_{n2}&\cdots &A_{nn}\\\end{bmatrix}}{\begin{bmatrix}v_{1}\\v_{2}\\\vdots \\v_{n}\end{bmatrix}}={\begin{bmatrix}w_{1}\\w_{2}\\\vdots \\w_{n}\end{bmatrix}}}

where, for each row,

w i = A i 1 v 1 + A i 2 v 2 + ⋯ + A i n v n = ∑ j = 1 n A i j v j . {\displaystyle w_{i}=A_{i1}v_{1}+A_{i2}v_{2}+\cdots +A_{in}v_{n}=\sum _{j=1}^{n}A_{ij}v_{j}.}

If it occurs that ⁠ v {\displaystyle \mathbf {v} } ⁠ and ⁠ w {\displaystyle \mathbf {w} } ⁠ are scalar multiples, that is, if

then ⁠ v {\displaystyle \mathbf {v} } ⁠ is an eigenvector of the linear transformation ⁠ A {\displaystyle A} ⁠ and the scale factor ⁠ λ {\displaystyle \lambda } ⁠ is the eigenvalue corresponding to that eigenvector. Equation (1) is the eigenvalue equation for the matrix ⁠ A {\displaystyle A} ⁠. Equation (1) can be stated equivalently as

where ⁠ I {\displaystyle I} ⁠ is the ⁠ n × n {\displaystyle n\times n} ⁠ identity matrix and ⁠ 0 {\displaystyle \mathbf {0} } ⁠ is the zero vector.

Eigenvalues and characteristic polynomial

Equation (2) has a nonzero solution v if and only if the determinant of the matrix (A − λI) is zero. Therefore, the eigenvalues of A are values of λ that satisfy the equation

Using the Leibniz formula for determinants, the left-hand side of equation (3) is a polynomial function of the variable λ and the degree of this polynomial is n, the order of the matrix A. Its coefficients depend on the entries of A, except that its term of degree n is always (−1)nλn. This polynomial is called the characteristic polynomial of A. Equation (3) is called the characteristic equation or secular equation of A. The characteristic polynomial of an n-by-n matrix A, being a polynomial of degree n, has at most n complex number roots, which can be found by factoring the characteristic polynomial, or numerically by root finding. The characteristic polynomial can be factored into the product of n linear terms:

where the complex numbers λ1, λ2, ..., λn, each of which is an eigenvalue, may repeat. (The number of times an eigenvalue appears in the characteristic polynomial is known as its algebraic multiplicity.) As a brief example, which is described in more detail in the examples section later, consider the matrix

A = [ 2 1 1 2 ] . {\displaystyle A={\begin{bmatrix}2&1\\1&2\end{bmatrix}}.}

Taking the determinant of (A − λI), the characteristic polynomial of A is

det ( A − λ I ) = | 2 − λ 1 1 2 − λ | = 3 − 4 λ + λ 2 . {\displaystyle \det(A-\lambda I)={\begin{vmatrix}2-\lambda &1\\1&2-\lambda \end{vmatrix}}=3-4\lambda +\lambda ^{2}.}

Setting the characteristic polynomial equal to zero, it has roots at λ = 1 and λ = 3, which are the two eigenvalues of A. The eigenvectors corresponding to each eigenvalue λ can be found by solving for the components of v in the equation (A − λI)v = 0. In this example, the eigenvectors are any nonzero scalar multiples of

v λ = 1 = [ 1 − 1 ] , v λ = 3 = [ 1 1 ] . {\displaystyle \mathbf {v} _{\lambda =1}={\begin{bmatrix}1\\-1\end{bmatrix}},\quad \mathbf {v} _{\lambda =3}={\begin{bmatrix}1\\1\end{bmatrix}}.}

If the entries of the matrix A are all real numbers, then the coefficients of the characteristic polynomial will also be real numbers, but the eigenvalues may still have nonzero imaginary parts. The entries of the corresponding eigenvectors therefore may also have nonzero imaginary parts. Similarly, the eigenvalues may be irrational numbers even if all the entries of A are rational numbers or even if they are all integers. However, if the entries of A are all algebraic numbers, which include the rationals, then the eigenvalues must also be algebraic numbers. The non-real roots of a real polynomial with real coefficients can be grouped into pairs of complex conjugates, namely with the two members of each pair having imaginary parts that differ only in sign and the same real part. If the degree is odd, then by the intermediate value theorem at least one of the roots is real. Therefore, any real matrix with odd order has at least one real eigenvalue, whereas a real matrix with even order may not have any real eigenvalues. The eigenvectors associated with these complex eigenvalues are also complex and also appear in complex conjugate pairs.

Spectrum of a matrix The spectrum of a matrix is the list of its eigenvalues, repeated according to their multiplicities; in a shorter notation, the set of its eigenvalues with their multiplicities indicated. An important quantity associated with the spectrum of a matrix is the maximum absolute value of all of its eigenvalues. This is known as the spectral radius of the considered matrix.

Algebraic multiplicity Let λi be an eigenvalue of an n-by-n matrix A. The algebraic multiplicity μA(λi) of the eigenvalue is its multiplicity as a root of the characteristic polynomial, that is, the largest integer k such that (λi − λ)k evenly divides that polynomial. Suppose a matrix A has order n and d ≤ n distinct eigenvalues. Whereas equation (4) factors the characteristic polynomial of A into the product of n linear terms with some terms potentially repeating, the characteristic polynomial can also be written as the product of d terms each corresponding to a distinct eigenvalue and raised to the power of the algebraic multiplicity:

det ( A − λ I ) = ( λ 1 − λ ) μ A ( λ 1 ) ( λ 2 − λ ) μ A ( λ 2 ) ⋯ ( λ d − λ ) μ A ( λ d ) . {\displaystyle \det(A-\lambda I)=(\lambda _{1}-\lambda )^{\mu _{A}(\lambda _{1})}(\lambda _{2}-\lambda )^{\mu _{A}(\lambda _{2})}\cdots (\lambda _{d}-\lambda )^{\mu _{A}(\lambda _{d})}.}

If d = n, then the right-hand side is the product of n linear terms, and this is the same as equation (4). The size of each eigenvalue's algebraic multiplicity is related to the dimension n as

1 ≤ μ A ( λ i ) ≤ n , μ A = ∑ i = 1 d μ A ( λ i ) = n . {\displaystyle {\begin{aligned}1&\leq \mu _{A}(\lambda _{i})\leq n,\\\mu _{A}&=\sum _{i=1}^{d}\mu _{A}\left(\lambda _{i}\right)=n.\end{aligned}}}

If μA(λi) = 1, then λi is said to be a simple eigenvalue. If μA(λi) equals the geometric multiplicity of λi (denoted by γA(λi) and defined in the next section), then λi is said to be a semisimple eigenvalue.

Eigenspaces, geometric multiplicities, and eigenbasis for a matrix Given a particular eigenvalue ⁠ λ {\displaystyle \lambda } ⁠ of the ⁠ n × n {\displaystyle n\times n} ⁠ matrix ⁠ A {\displaystyle A} ⁠, define the set ⁠ E {\displaystyle E} ⁠ to be all vectors ⁠ v {\displaystyle \mathbf {v} } ⁠ that satisfy equation (2):

E = { v : ( A − λ I ) v = 0 } . {\displaystyle E=\left\{\mathbf {v} :\left(A-\lambda I\right)\mathbf {v} =\mathbf {0} \right\}.}

On one hand, ⁠ E {\displaystyle E} ⁠ is precisely the kernel or nullspace of the matrix ⁠ A − λ I {\displaystyle A-\lambda I} ⁠. On the other hand, by definition, any nonzero vector that satisfies this condition is an eigenvector of ⁠ A {\displaystyle A} ⁠ associated with ⁠ λ {\displaystyle \lambda } ⁠; so ⁠ E {\displaystyle E} ⁠ is the union of the zero vector with the set of all eigenvectors of ⁠ A {\displaystyle A} ⁠ associated with ⁠ λ {\displaystyle \lambda } ⁠. The space ⁠ E {\displaystyle E} ⁠ is called the eigenspace or characteristic space of ⁠ A {\displaystyle A} ⁠ associated with ⁠ λ {\displaystyle \lambda } ⁠. In general, ⁠ λ {\displaystyle \lambda } ⁠ is a complex number and the eigenvectors are complex ⁠ n × 1 {\displaystyle n\times 1} ⁠ matrices (column vectors). Because every nullspace is a linear subspace of the domain, ⁠ E {\displaystyle E} ⁠ is a linear subspace of ⁠ C n {\displaystyle \mathbb {C} ^{n}} ⁠. Because the eigenspace ⁠ E {\displaystyle E} ⁠ is a linear subspace, it is closed under addition. That is, if two vectors ⁠ u {\displaystyle \mathbf {u} } ⁠ and ⁠ v {\displaystyle \mathbf {v} } ⁠ belong to the set ⁠ E {\displaystyle E} ⁠, written ⁠ u , v ∈ E {\displaystyle \mathbf {u} ,\mathbf {v} \in E} ⁠, then ⁠ u + v ∈ E {\displaystyle \mathbf {u} +\mathbf {v} \in E} ⁠, or equivalently ⁠ A ( u + v ) = λ ( u + v ) {\displaystyle A(\mathbf {u} +\mathbf {v} )=\lambda (\mathbf {u} +\mathbf {v} )} ⁠. This can be checked using the distributive property of matrix multiplication. Similarly, because ⁠ E {\displaystyle E} ⁠ is a linear subspace, it is closed under scalar multiplication. That is, if ⁠ v ∈ E {\displaystyle \mathbf {v} \in E} ⁠ and ⁠ α ∈ C {\displaystyle \alpha \in \mathbb {C} } ⁠, then ⁠ α v ∈ E {\displaystyle \alpha \mathbf {v} \in E} ⁠, or equivalently ⁠ A ( α v ) = λ ( α v ) {\displaystyle A(\alpha \mathbf {v} )=\lambda (\alpha \mathbf {v} )} ⁠. This can be checked by noting that multiplication of complex matrices by complex numbers is commutative. As long as ⁠ u + v {\displaystyle \mathbf {u} +\mathbf {v} } ⁠ and ⁠ α v {\displaystyle \alpha \mathbf {v} } ⁠ are not zero, they are also eigenvectors of ⁠ A {\displaystyle A} ⁠ associated with ⁠ λ {\displaystyle \lambda } ⁠. The dimension of the eigenspace ⁠ E {\displaystyle E} ⁠ associated with ⁠ λ {\displaystyle \lambda } ⁠, or equivalently the maximum number of linearly independent eigenvectors associated with ⁠ λ {\displaystyle \lambda } ⁠, is referred to as the eigenvalue's geometric multiplicity and denoted by ⁠ γ A ( λ ) {\displaystyle \gamma _{A}(\lambda )} ⁠. Because ⁠ E {\displaystyle E} ⁠ is also the nullspace of ⁠ A − λ I {\displaystyle A-\lambda I} ⁠, the geometric multiplicity of ⁠ λ {\displaystyle \lambda } ⁠ is the dimension of the nullspace of ⁠ A − λ I {\displaystyle A-\lambda I} ⁠, also called the nullity of ⁠ A − λ I {\displaystyle A-\lambda I} ⁠. This quantity is related to the size and rank of ⁠ A − λ I {\displaystyle A-\lambda I} ⁠ by the equation:

γ A ( λ ) = n − rank ⁡ ( A − λ I ) . {\displaystyle \gamma _{A}(\lambda )=n-\operatorname {rank} (A-\lambda I).}

Because of the definition of eigenvalues and eigenvectors, an eigenvalue's geometric multiplicity must be at least one, that is, each eigenvalue has at least one associated eigenvector. Furthermore, an eigenvalue's geometric multiplicity cannot exceed its algebraic multiplicity. Additionally, recall that an eigenvalue's algebraic multiplicity cannot exceed ⁠ n {\displaystyle n} ⁠. In summary,

1 ≤ γ A ( λ ) ≤ μ A ( λ ) ≤ n . {\displaystyle 1\leq \gamma _{A}(\lambda )\leq \mu _{A}(\lambda )\leq n.}

Proof of inequality ⁠ γ A ( λ ) ≤ μ A ( λ ) {\displaystyle \gamma _{A}(\lambda )\leq \mu _{A}(\lambda )} ⁠: Let B = A − λI, where λ is a fixed complex number, and the eigenspace associated with λ is the nullspace of B. Let the dimension of that eigenspace be ⁠ k = γ A ( λ ) {\displaystyle k=\gamma _{A}(\lambda )} ⁠. This means that the last k rows of the echelon form of B are zero. Thus, there is an invertible matrix E coming from Gauss-Jordan reduction, such that

E B = [ ∗ ∗ 0 k × ( n − k ) 0 k × k ] . {\displaystyle EB={\begin{bmatrix}*&*\\\mathbf {0} _{k\times (n-k)}&\mathbf {0} _{k\times k}\end{bmatrix}}.}

Therefore the last k rows of EB − tE are (−t) times the last k rows of E. Therefore the polynomial tk evenly divides the polynomial det(EB − tE), because of basic properties of determinants (homogeneity). On the other hand, det(EB − tE) = det E det(B − tI) = pA(t + λ) det E, so (t − λ)k divides pA(t), and so the algebraic multiplicity of λ is at least ⁠ k {\displaystyle k} ⁠. Q.E.D. Suppose A has d ≤ n distinct eigenvalues λ1, ..., λd, where the geometric multiplicity of λi is γA(λi). The total geometric multiplicity of A,

γ

Tags

  • Abstract algebra
  • Linear algebra
  • Mathematical physics
  • Matrix theory
  • Singular value decomposition