In mathematics, the Legendre transformation (or Legendre transform), first introduced by Adrien-Marie Legendre in 1787 when studying the minimal surface problem, is an involutive transformation on real-valued functions that are convex on a real variable. Specifically, if a real-valued multivariable function is convex on one of its independent real variables, then the Legendre transform with respect to this variable is applicable to the function. In physical problems, the Legendre transform is used to convert functions of one quantity (such as position, pressure, or temperature) into functions of the conjugate quantity (momentum, volume, and entropy, respectively). In this way, it is commonly used in classical mechanics to derive the Hamiltonian formalism out of the Lagrangian formalism (or vice versa) and in thermodynamics to derive the thermodynamic potentials, as well as in the solution of differential equations of several variables. For sufficiently smooth functions on the real line, the Legendre transform f ∗ {\displaystyle f^{*}} of a function f {\displaystyle f} can be specified, up to an additive constant, by the condition that the functions' first derivatives are inverse functions of each other. This can be expressed as
d f d x = ( d f ∗ d x ) − 1 {\displaystyle {\frac {df}{dx}}=\left({\frac {df^{*}}{dx}}\right)^{-1}~} in Leibniz's notation. The generalization of the Legendre transformation to affine spaces and non-convex functions is known as the convex conjugate (also called the Legendre–Fenchel transformation), which can be used to construct a function's convex hull.
Definition
Definition in one-dimensional real space Let I ⊂ R {\displaystyle I\subset \mathbb {R} } be an interval, and f : I → R {\displaystyle f:I\to \mathbb {R} } a convex function; then the Legendre transform of f {\displaystyle f} is the function f ∗ : I ∗ → R {\displaystyle f^{*}:I^{*}\to \mathbb {R} } defined by
f ∗ ( x ∗ ) = sup x ∈ I ( x ∗ x − f ( x ) ) , I ∗ = { x ∗ ∈ R : sup x ∈ I ( x ∗ x − f ( x ) ) < ∞ } {\displaystyle f^{*}(x^{*})=\sup _{x\in I}(x^{*}x-f(x)),\ \ \ \ I^{*}=\left\{x^{*}\in \mathbb {R} :\sup _{x\in I}(x^{*}x-f(x))<\infty \right\}}
where sup {\textstyle \sup } denotes the supremum over I {\displaystyle I} , e.g., x {\textstyle x} in I {\textstyle I} is chosen such that x ∗ x − f ( x ) {\textstyle x^{*}x-f(x)} is maximized at each x ∗ {\textstyle x^{*}} , or x ∗ {\textstyle x^{*}} is such that x ∗ x − f ( x ) {\displaystyle x^{*}x-f(x)} has a bounded value throughout I {\textstyle I} (e.g., when f ( x ) {\displaystyle f(x)} is a linear function). The function f ∗ {\displaystyle f^{*}} is called the convex conjugate function of f {\displaystyle f} . For historical reasons (rooted in analytic mechanics), the conjugate variable is often denoted p {\displaystyle p} , instead of x ∗ {\displaystyle x^{*}} . If the convex function f {\displaystyle f} is defined on the whole line and is everywhere differentiable, then
f ∗ ( p ) = sup x ∈ I ( p x − f ( x ) ) = ( p x − f ( x ) ) | x = ( f ′ ) − 1 ( p ) {\displaystyle f^{*}(p)=\sup _{x\in I}(px-f(x))=\left(px-f(x)\right)|_{x=(f')^{-1}(p)}}
can be interpreted as the negative of the y {\displaystyle y} -intercept of the tangent line to the graph of f {\displaystyle f} that has slope p {\displaystyle p} .
Definition in n-dimensional real space The generalization to convex functions f : X → R {\displaystyle f:X\to \mathbb {R} } on a convex set X ⊂ R n {\displaystyle X\subset \mathbb {R} ^{n}} is straightforward: f ∗ : X ∗ → R {\displaystyle f^{*}:X^{*}\to \mathbb {R} } has domain
X ∗ = { x ∗ ∈ R n : sup x ∈ X ( ⟨ x ∗ , x ⟩ − f ( x ) ) < ∞ } {\displaystyle X^{*}=\left\{x^{*}\in \mathbb {R} ^{n}:\sup _{x\in X}(\langle x^{*},x\rangle -f(x))<\infty \right\}}
and is defined by
f ∗ ( x ∗ ) = sup x ∈ X ( ⟨ x ∗ , x ⟩ − f ( x ) ) , x ∗ ∈ X ∗ , {\displaystyle f^{*}(x^{*})=\sup _{x\in X}(\langle x^{*},x\rangle -f(x)),\quad x^{*}\in X^{*}~,}
where ⟨ x ∗ , x ⟩ {\displaystyle \langle x^{*},x\rangle } denotes the dot product of x ∗ {\displaystyle x^{*}} and x {\displaystyle x} . The Legendre transformation is an application of the duality relationship between points and lines. The functional relationship specified by f {\displaystyle f} can be represented equally well as a set of ( x , y ) {\displaystyle (x,y)} points, or as a set of tangent lines specified by their slope and intercept values.
Understanding the Legendre transform in terms of derivatives For a differentiable convex function f {\displaystyle f} on the real line with the first derivative f ′ {\displaystyle f'} and its inverse ( f ′ ) − 1 {\displaystyle (f')^{-1}} , the Legendre transform of f {\displaystyle f} , f ∗ {\displaystyle f^{*}} , can be specified, up to an additive constant, by the condition that the functions' first derivatives are inverse functions of each other, i.e., f ′ = ( ( f ∗ ) ′ ) − 1 {\displaystyle f'=((f^{*})')^{-1}} and ( f ∗ ) ′ = ( f ′ ) − 1 {\displaystyle (f^{*})'=(f')^{-1}} . To see this, first note that if f {\displaystyle f} as a convex function on the real line is differentiable and x ¯ {\displaystyle {\overline {x}}} is a critical point of the function of x ↦ p ⋅ x − f ( x ) {\displaystyle x\mapsto p\cdot x-f(x)} , then the supremum is achieved at x ¯ {\textstyle {\overline {x}}} (by convexity, see the first figure in this Wikipedia page). Therefore, the Legendre transform of f {\displaystyle f} is f ∗ ( p ) = p ⋅ x ¯ − f ( x ¯ ) {\displaystyle f^{*}(p)=p\cdot {\overline {x}}-f({\overline {x}})} . Then, suppose that the first derivative f ′ {\displaystyle f'} is invertible and let the inverse be g = ( f ′ ) − 1 {\displaystyle g=(f')^{-1}} . Then for each p {\textstyle p} , the point g ( p ) {\displaystyle g(p)} is the unique critical point x ¯ {\textstyle {\overline {x}}} of the function x ↦ p x − f ( x ) {\displaystyle x\mapsto px-f(x)} (i.e., x ¯ = g ( p ) {\displaystyle {\overline {x}}=g(p)} ) because f ′ ( g ( p ) ) = p {\displaystyle f'(g(p))=p} and the function's first derivative with respect to x {\displaystyle x} at g ( p ) {\displaystyle g(p)} is p − f ′ ( g ( p ) ) = 0 {\displaystyle p-f'(g(p))=0} . Hence we have f ∗ ( p ) = p ⋅ g ( p ) − f ( g ( p ) ) {\displaystyle f^{*}(p)=p\cdot g(p)-f(g(p))} for each p {\textstyle p} . By differentiating with respect to p {\textstyle p} , we find
( f ∗ ) ′ ( p ) = g ( p ) + p ⋅ g ′ ( p ) − f ′ ( g ( p ) ) ⋅ g ′ ( p ) . {\displaystyle (f^{*})'(p)=g(p)+p\cdot g'(p)-f'(g(p))\cdot g'(p).}
Since f ′ ( g ( p ) ) = p {\displaystyle f'(g(p))=p} this simplifies to ( f ∗ ) ′ ( p ) = g ( p ) = ( f ′ ) − 1 ( p ) {\displaystyle (f^{*})'(p)=g(p)=(f')^{-1}(p)} . In other words, ( f ∗ ) ′ {\displaystyle (f^{*})'} and f ′ {\displaystyle f'} are inverses to each other. In general, if h ′ = ( f ′ ) − 1 {\displaystyle h'=(f')^{-1}} as the inverse of f ′ , {\displaystyle f',} then h ′ = ( f ∗ ) ′ {\displaystyle h'=(f^{*})'} so integration gives f ∗ = h + c {\displaystyle f^{*}=h+c} , where c {\displaystyle c} is a constant. In practical terms, given f ( x ) , {\displaystyle f(x),} the parametric plot of x f ′ ( x ) − f ( x ) {\displaystyle xf'(x)-f(x)} versus f ′ ( x ) {\displaystyle f'(x)} amounts to the graph of f ∗ ( p ) {\displaystyle f^{*}(p)} versus p . {\displaystyle p.}
In some cases (e.g., thermodynamic potentials, below), a non-standard requirement is used, amounting to an alternative definition of f * with a minus sign,
f ( x ) − f ∗ ( p ) = x p . {\displaystyle f(x)-f^{*}(p)=xp.}
Definition in physical contexts In analytical mechanics and thermodynamics, the Legendre transformation is usually defined as follows: suppose f {\displaystyle f} is a function of x {\displaystyle x} ; then we have
d f = d f d x d x . {\displaystyle df={\frac {df}{dx}}dx.}
Performing the Legendre transformation on this function means that we take p = d f d x {\displaystyle p={\frac {df}{dx}}} as the independent variable, so that the above expression can be written as
d f = p d x , {\displaystyle df=p\,dx,}
and according to the product rule d ( u v ) = u d v + v d u , {\displaystyle d(uv)=u\,dv+v\,du,} we then have
d ( x p − f ) = x d p + p d x − d f = x d p , {\displaystyle d\left(xp-f\right)=x\,dp+p\,dx-df=x\,dp,}
and taking f ∗ = x p − f , {\displaystyle f^{*}=xp-f,} we have d f ∗ = x d p , {\displaystyle df^{*}=x\,dp,} which means
d f ∗ d p = x . {\displaystyle {\frac {df^{*}}{dp}}=x.}
When f {\displaystyle f} is a function of n {\displaystyle n} variables x 1 , x 2 , ⋯ , x n {\displaystyle x_{1},x_{2},\cdots ,x_{n}} , then we can perform the Legendre transformation on each one or several variables: we have
d f = p 1 d x 1 + p 2 d x 2 + ⋯ + p n d x n , {\displaystyle df=p_{1}\,dx_{1}+p_{2}\,dx_{2}+\cdots +p_{n}\,dx_{n},}
where p i = ∂ f ∂ x i . {\displaystyle p_{i}={\frac {\partial f}{\partial x_{i}}}.} Then if we want to perform the Legendre transformation on, e.g. x 1 {\displaystyle x_{1}} , then we take p 1 {\displaystyle p_{1}} together with x 2 , ⋯ , x n {\displaystyle x_{2},\cdots ,x_{n}} as independent variables, and with Leibniz's rule we have
d ( x 1 p 1 − f ) = x 1 d p 1 − p 2 d x 2 − ⋯ − p n d x n . {\displaystyle d(x_{1}p_{1}-f)=x_{1}\,dp_{1}-p_{2}\,dx_{2}-\cdots -p_{n}\,dx_{n}.}
So for the function φ ( p 1 , x 2 , ⋯ , x n ) = x 1 p 1 − f ( x 1 , x 2 , ⋯ , x n ) , {\displaystyle \varphi (p_{1},x_{2},\cdots ,x_{n})=x_{1}p_{1}-f(x_{1},x_{2},\cdots ,x_{n}),} we have
∂ φ ∂ p 1 = x 1 , ∂ φ ∂ x 2 = − p 2 , ⋯ , ∂ φ ∂ x n = − p n . {\displaystyle {\frac {\partial \varphi }{\partial p_{1}}}=x_{1},\quad {\frac {\partial \varphi }{\partial x_{2}}}=-p_{2},\quad \cdots ,\quad {\frac {\partial \varphi }{\partial x_{n}}}=-p_{n}.}
We can also do this transformation for variables x 2 , ⋯ , x n {\displaystyle x_{2},\cdots ,x_{n}} . If we do it to all the variables, then we have
d φ = x 1 d p 1 + x 2 d p 2 + ⋯ + x n d p n {\displaystyle d\varphi =x_{1}\,dp_{1}+x_{2}\,dp_{2}+\cdots +x_{n}\,dp_{n}} where φ = x 1 p 1 + x 2 p 2 + ⋯ + x n p n − f . {\displaystyle \varphi =x_{1}p_{1}+x_{2}p_{2}+\cdots +x_{n}p_{n}-f.}
In analytical mechanics, people perform this transformation on variables q ˙ 1 , q ˙ 2 , ⋯ , q ˙ n {\displaystyle {\dot {q}}_{1},{\dot {q}}_{2},\cdots ,{\dot {q}}_{n}} of the Lagrangian L ( q 1 , ⋯ , q n , q ˙ 1 , ⋯ , q ˙ n ) {\displaystyle L(q_{1},\cdots ,q_{n},{\dot {q}}_{1},\cdots ,{\dot {q}}_{n})} to get the Hamiltonian:
H ( q 1 , ⋯ , q n , p 1 , ⋯ , p n ) = ∑ i = 1 n p i q ˙ i − L ( q 1 , ⋯ , q n , q ˙ 1 ⋯ , q ˙ n ) . {\displaystyle H(q_{1},\cdots ,q_{n},p_{1},\cdots ,p_{n})=\sum _{i=1}^{n}p_{i}{\dot {q}}_{i}-L(q_{1},\cdots ,q_{n},{\dot {q}}_{1}\cdots ,{\dot {q}}_{n}).}
In thermodynamics, this transformation is applied to variables according to the type of thermodynamic system desired; for example, starting from the energy representation cardinal function of state, the internal energy U ( S , V ) {\displaystyle U(S,V)} , we have
d U = T d S − p d V , {\displaystyle dU=T\,dS-p\,dV,}
so we can perform the Legendre transformation on either or both of S , V {\displaystyle S,V} to yield
d H = d ( U + p V ) = T d S + V d p d F = d ( U − T S ) = − S d T − p d V d G = d ( U − T S + p V ) = − S d T + V d p , {\displaystyle {\begin{aligned}dH&=d(U+pV)\ \ \ \ \ \ \ \ \ \ =\ \ \ \ T\,dS+V\,dp\\dF&=d(U-TS)\ \ \ \ \ \ \ \ \ \ =-S\,dT-p\,dV\\dG&=d(U-TS+pV)=-S\,dT+V\,dp,\end{aligned}}}
and each of these three expressions has a physical meaning. This definition of the Legendre transformation is the one originally introduced by Legendre in his work in 1787, and is still applied by physicists nowadays. Indeed, this definition is mathematically rigorous if we treat all the variables and functions defined above (for example, f , x 1 , ⋯ , x n , p 1 , ⋯ , p n , {\displaystyle f,x_{1},\cdots ,x_{n},p_{1},\cdots ,p_{n},} ) as differentiable functions defined on an open set of R n {\displaystyle \mathbb {R} ^{n}} or on a differentiable manifold, and d f , d x i , d p i {\displaystyle df,dx_{i},dp_{i}} their differentials (which are treated as cotangent vectors in the context of differentiable manifolds). This definition is equivalent to the modern mathematicians' definition as long as f {\displaystyle f} is differentiable and convex for the variables x 1 , x 2 , ⋯ , x n . {\displaystyle x_{1},x_{2},\cdots ,x_{n}.}
Properties The Legendre transform of a convex function, of which double derivative values are all positive, is also a convex function of which double derivative values are all positive.Proof. Let us show this with a doubly differentiable function f ( x ) {\displaystyle f(x)} with all positive double derivative values and with a bijective (invertible) derivative. For a fixed p {\displaystyle p} , let x ¯ {\displaystyle {\bar {x}}} maximize or make the function p x − f ( x ) {\displaystyle px-f(x)} bounded over x {\displaystyle x} . Then the Legendre transformation of f {\displaystyle f} is f ∗ ( p ) = p x ¯ − f ( x ¯ ) {\displaystyle f^{*}(p)=p{\bar {x}}-f({\bar {x}})} , thus, f ′ ( x ¯ ) = p {\displaystyle f'({\bar {x}})=p} by the maximizing or bounding condition d d x ( p x − f ( x ) ) = p − f ′ ( x ) = 0 {\displaystyle {\frac {d}{dx}}(px-f(x))=p-f'(x)=0} . Note that x ¯ {\displaystyle {\bar {x}}} depends on p {\displaystyle p} . (This can be visually shown in the 1st figure of this page above.) Thus x ¯ = g ( p ) {\displaystyle {\bar {x}}=g(p)} where g ≡ ( f ′ ) − 1 {\displaystyle g\equiv (f')^{-1}} , meaning that g {\displaystyle g} is the inverse of f ′ {\displaystyle f'} that is the derivative of f {\displaystyle f} (so f ′ ( g ( p ) ) = p {\displaystyle f'(g(p))
