In mathematics, the symmetry of second derivatives (also called the equality of mixed partials) is the fact that exchanging the order of partial derivatives of a multivariate function
f ( x 1 , x 2 , … , x n ) {\displaystyle f\left(x_{1},\,x_{2},\,\ldots ,\,x_{n}\right)}
does not change the result if some continuity conditions are satisfied (see below); that is, the second-order partial derivatives satisfy the identities
∂ ∂ x i ( ∂ f ∂ x j ) = ∂ ∂ x j ( ∂ f ∂ x i ) . {\displaystyle {\frac {\partial }{\partial x_{i}}}\left({\frac {\partial f}{\partial x_{j}}}\right)\ =\ {\frac {\partial }{\partial x_{j}}}\left({\frac {\partial f}{\partial x_{i}}}\right).}
In other words, the matrix of the second-order partial derivatives, known as the Hessian matrix, is a symmetric matrix. Sufficient conditions for the symmetry to hold are given by Schwarz's theorem, also called Clairaut's theorem or Young's theorem. In the context of partial differential equations, it is called the Schwarz integrability condition.
Formal expressions of symmetry In symbols, the symmetry may be expressed as:
∂ ∂ x ( ∂ f ∂ y ) = ∂ ∂ y ( ∂ f ∂ x ) or ∂ 2 f ∂ x ∂ y = ∂ 2 f ∂ y ∂ x . {\displaystyle {\frac {\partial }{\partial x}}\left({\frac {\partial f}{\partial y}}\right)\ =\ {\frac {\partial }{\partial y}}\left({\frac {\partial f}{\partial x}}\right)\qquad {\text{or}}\qquad {\frac {\partial ^{2}\!f}{\partial x\,\partial y}}\ =\ {\frac {\partial ^{2}\!f}{\partial y\,\partial x}}.}
Another notation is:
∂ x ∂ y f = ∂ y ∂ x f or f y x = f x y . {\displaystyle \partial _{x}\partial _{y}f=\partial _{y}\partial _{x}f\qquad {\text{or}}\qquad f_{yx}=f_{xy}.}
In terms of composition of the differential operator Di which takes the partial derivative with respect to xi:
D i ∘ D j = D j ∘ D i . {\displaystyle D_{i}\circ D_{j}=D_{j}\circ D_{i}.}
From this relation it follows that the ring of differential operators with constant coefficients, generated by the Di, is commutative; but this is only true as operators over a domain of sufficiently differentiable functions. It is easy to check the symmetry as applied to monomials, so that one can take polynomials in the xi as a domain. In fact smooth functions are another valid domain. In terms of the total derivative, the second derivative of a twice-differentiable function f : X → Y {\displaystyle f\colon X\to Y} at a point p ∈ X {\displaystyle p\in X} is a linear map
D 2 f p : X → L ( X , Y ) {\displaystyle D^{2}f_{p}\colon X\to {\mathcal {L}}(X,Y)}
Where L ( X , Y ) {\displaystyle {\mathcal {L}}(X,Y)} is the normed vector space of linear maps from X {\displaystyle X} to Y {\displaystyle Y} . By uncurrying, it may be identified with a bilinear form from X × X {\displaystyle X\times X} to Y {\displaystyle Y} . The symmetry of second derivatives then says that this symmetric, ie. D 2 f p ( u ) ( v ) = D 2 f p ( v ) ( u ) {\displaystyle D^{2}f_{p}(u)(v)=D^{2}f_{p}(v)(u)} for all u , v ∈ X {\displaystyle u,v\in X} . Setting u {\displaystyle u} and v {\displaystyle v} to be the standard basis vectors e i {\displaystyle \mathbf {e} _{i}} and e j {\displaystyle \mathbf {e} _{j}} recovers the partial derivatives version. More generally, whenever f {\displaystyle f} is r {\displaystyle r} -times differentiable its derivative at a point p {\displaystyle p} is a symmetric r {\displaystyle r} -linear form D r f p : X ⊗ r → Y {\displaystyle D^{r}f_{p}\colon X^{\otimes r}\to Y} .
History The result on the equality of mixed partial derivatives under certain conditions has a long history. The list of unsuccessful proposed proofs started with Euler's, published in 1740, although already in 1721 Bernoulli had implicitly assumed the result with no formal justification. Clairaut also published a proposed proof in 1740, with no other attempts until the end of the 18th century. Starting then, for a period of 70 years, a number of incomplete proofs were proposed. The proof of Lagrange (1797) was improved by Cauchy (1823), but assumed the existence and continuity of the partial derivatives ∂ 2 f ∂ x 2 {\displaystyle {\tfrac {\partial ^{2}f}{\partial x^{2}}}} and ∂ 2 f ∂ y 2 {\displaystyle {\tfrac {\partial ^{2}f}{\partial y^{2}}}} . Other attempts were made by P. Blanchet (1841), Duhamel (1856), Sturm (1857), Schlömilch (1862), and Bertrand (1864). Finally in 1867 Lindelöf systematically analyzed all the earlier flawed proofs and was able to exhibit a specific counterexample where mixed derivatives failed to be equal. Six years after that, Schwarz succeeded in giving the first rigorous proof. Dini later contributed by finding more general conditions than those of Schwarz. Eventually a clean and more general version was found by Jordan in 1883 that is still the proof found in most textbooks. Minor variants of earlier proofs were published by Laurent (1885), Peano (1889 and 1893), J. Edwards (1892), P. Haag (1893), J. K. Whittemore (1898), Vivanti (1899) and Pierpont (1905). Further progress was made in 1907-1909 when E. W. Hobson and W. H. Young found proofs with weaker conditions than those of Schwarz and Dini. In 1918, Carathéodory gave a different proof based on the Lebesgue integral.
Schwarz's theorem
In mathematical analysis, Schwarz's theorem (or Clairaut's theorem on equality of mixed partials) named after Alexis Clairaut and Hermann Schwarz, states that for a function f : Ω → R {\displaystyle f\colon \Omega \to \mathbb {R} } defined on a set Ω ⊂ R n {\displaystyle \Omega \subset \mathbb {R} ^{n}} , if p ∈ R n {\displaystyle \mathbf {p} \in \mathbb {R} ^{n}} is a point such that some neighborhood of p {\displaystyle \mathbf {p} } is contained in Ω {\displaystyle \Omega } and f {\displaystyle f} has continuous second partial derivatives on that neighborhood of p {\displaystyle \mathbf {p} } , then for all i and j in { 1 , 2 … , n } , {\displaystyle \{1,2\ldots ,\,n\},}
∂ 2 ∂ x i ∂ x j f ( p ) = ∂ 2 ∂ x j ∂ x i f ( p ) . {\displaystyle {\frac {\partial ^{2}}{\partial x_{i}\,\partial x_{j}}}f(\mathbf {p} )={\frac {\partial ^{2}}{\partial x_{j}\,\partial x_{i}}}f(\mathbf {p} ).}
The partial derivatives of this function commute at that point. There exists a version of this theorem where f {\displaystyle f} is only required to be twice differentiable at the point p {\displaystyle \mathbf {p} } . One easy way to establish this theorem (in the case where n = 2 {\displaystyle n=2} , i = 1 {\displaystyle i=1} , and j = 2 {\displaystyle j=2} , which readily entails the result in general) is by applying Green's theorem to the gradient of f . {\displaystyle f.}
An elementary proof for functions on open subsets of the plane is as follows (by a simple reduction, the general case for the theorem of Schwarz easily reduces to the planar case). Let f ( x , y ) {\displaystyle f(x,y)} be a differentiable function on an open rectangle Ω {\displaystyle \Omega } containing a point ( a , b ) {\displaystyle (a,b)} and suppose that d f {\displaystyle df} is continuous with continuous ∂ x ∂ y f {\displaystyle \partial _{x}\partial _{y}f} and ∂ y ∂ x f {\displaystyle \partial _{y}\partial _{x}f} over Ω . {\displaystyle \Omega .} Define
u ( h , k ) = f ( a + h , b + k ) − f ( a + h , b ) , v ( h , k ) = f ( a + h , b + k ) − f ( a , b + k ) , w ( h , k ) = f ( a + h , b + k ) − f ( a + h , b ) − f ( a , b + k ) + f ( a , b ) . {\displaystyle {\begin{aligned}u{\left(h,\,k\right)}&=f\left(a{+}h,\,b{+}k\right)-f\left(a{+}h,\,b\right),\\v{\left(h,\,k\right)}&=f\left(a{+}h,\,b{+}k\right)-f\left(a,\,b{+}k\right),\\w{\left(h,\,k\right)}&=f\left(a{+}h,\,b{+}k\right)-f\left(a{+}h,\,b\right)-f\left(a,\,b{+}k\right)+f\left(a,\,b\right).\end{aligned}}}
These functions are defined for | h | , | k | < ε {\displaystyle \left|h\right|,\left|k\right|<\varepsilon } , where ε > 0 {\displaystyle \varepsilon >0} and [ a − ε , a + ε ] × [ b − ε , b + ε ] {\displaystyle \left[a{-}\varepsilon ,\,a{+}\varepsilon \right]\times \left[b{-}\varepsilon ,\,b{+}\varepsilon \right]} is contained in Ω . {\displaystyle \Omega .}
By the mean value theorem, for fixed h and k non-zero, θ {\displaystyle \theta } , θ ′ {\displaystyle \theta '} , ϕ {\displaystyle \phi } , ϕ ′ {\displaystyle \phi '} can be found in the open interval ( 0 , 1 ) {\displaystyle (0,1)} with
w ( h , k ) = u ( h , k ) − u ( 0 , k ) = h ∂ x u ( θ h , k ) = h [ ∂ x f ( a + θ h , b + k ) − ∂ x f ( a + θ h , b ) ] = h k ∂ y ∂ x f ( a + θ h , b + θ ′ k ) w ( h , k ) = v ( h , k ) − v ( h , 0 ) = k ∂ y v ( h , ϕ k ) = k [ ∂ y f ( a + h , b + ϕ k ) − ∂ y f ( a , b + ϕ k ) ] = h k ∂ x ∂ y f ( a + ϕ ′ h , b + ϕ k ) . {\displaystyle {\begin{aligned}w{\left(h,\,k\right)}&=u{\left(h,\,k\right)}-u{\left(0,\,k\right)}=h\,\partial _{x}u{\left(\theta h,\,k\right)}\\[0.5ex]&=h\left[\partial _{x}f{\left(a{+}\theta h,\,b{+}k\right)}-\partial _{x}f{\left(a{+}\theta h,\,b\right)}\right]\\[0.5ex]&=hk\,\partial _{y}\partial _{x}f{\left(a{+}\theta h,\,b{+}\theta ^{\prime }k\right)}\\[1ex]w{\left(h,\,k\right)}&=v{\left(h,\,k\right)}-v{\left(h,\,0\right)}=k\,\partial _{y}v\left(h,\,\phi k\right)\\[0.5ex]&=k\left[\partial _{y}f{\left(a{+}h,\,b{+}\phi k\right)}-\partial _{y}f{\left(a,\,b{+}\phi k\right)}\right]\\[0.5ex]&=hk\,\partial _{x}\partial _{y}f{\left(a{+}\phi ^{\prime }h,\,b{+}\phi k\right)}.\end{aligned}}}
Since h , k ≠ 0 {\displaystyle h,\,k\neq 0} , the first equality below can be divided by h k {\displaystyle hk} :
h k ∂ y ∂ x f ( a + θ h , b + θ ′ k ) = h k ∂ x ∂ y f ( a + ϕ ′ h , b + ϕ k ) , ∂ y ∂ x f ( a + θ h , b + θ ′ k ) = ∂ x ∂ y f ( a + ϕ ′ h , b + ϕ k ) . {\displaystyle {\begin{aligned}hk\,\partial _{y}\partial _{x}f{\left(a{+}\theta h,\,b{+}\theta ^{\prime }k\right)}&=hk\,\partial _{x}\partial _{y}f{\left(a{+}\phi ^{\prime }h,\,b{+}\phi k\right)},\\\partial _{y}\partial _{x}f{\left(a{+}\theta h,\,b{+}\theta ^{\prime }k\right)}&=\partial _{x}\partial _{y}f{\left(a{+}\phi ^{\prime }h,\,b{+}\phi k\right)}.\end{aligned}}}
Letting h , k {\displaystyle h,\,k} tend to zero in the last equality, the continuity assumptions on ∂ y ∂ x f {\displaystyle \partial _{y}\partial _{x}f} and ∂ x ∂ y f {\displaystyle \partial _{x}\partial _{y}f} now imply that
∂ 2 ∂ x ∂ y f ( a , b ) = ∂ 2 ∂ y ∂ x f ( a , b ) . {\displaystyle {\frac {\partial ^{2}}{\partial x\partial y}}f\left(a,\,b\right)={\frac {\partial ^{2}}{\partial y\partial x}}f\left(a,\,b\right).}
This account is a straightforward classical method found in many text books, for example in Burkill, Apostol and Rudin. Although the derivation above is elementary, the approach can also be viewed from a more conceptual perspective so that the result becomes more apparent. Indeed the difference operators Δ x t , Δ y t {\displaystyle \Delta _{x}^{t},\,\,\Delta _{y}^{t}} commute and Δ x t f , Δ y t f {\displaystyle \Delta _{x}^{t}f,\,\,\Delta _{y}^{t}f} tend to ∂ x f , ∂ y f {\displaystyle \partial _{x}f,\,\,\partial _{y}f} as t {\displaystyle t} tends to 0, with a similar statement for second order operators. Here, for z {\displaystyle z} a vector in the plane and u {\displaystyle u} a directional vector ( 1 0 ) {\displaystyle {\tbinom {1}{0}}} or ( 0 1 ) {\displaystyle {\tbinom {0}{1}}} , the difference operator is defined by
Δ u t f ( z ) = f ( z + t u ) − f ( z ) t . {\displaystyle \Delta _{u}^{t}f(z)={f(z+tu)-f(z) \over t}.}
By the fundamental theorem of calculus for C 1 {\displaystyle C^{1}} functions f
