Introduction
Few developments in the rich history of mathematics may exult in advancing our comprehension and perception of the world as can the advent of calculus. Calculus germinated from the mathematical loam of mensuration, nourished by the burgeoning concerns of physics. The problems of calculus were tackled as early as Eudoxus of Cnidus and Archimedes, who had developed the method of exhaustion to compute areas and volumes by approximating figures with sequences of inscribed and circumscribed polygons, proving results through double reductio ad absurdum arguments. The notion of infinitesimals although spurned for its apparent lack of rigour, nonetheless proved an essential heuristic as evinced by Archimedes’ Method of Mechanical Theorems who envisioned plane figures as comprised of indivisible slices to obtain an ansatz before applying the method of exhaustion. Limit-esque arguments also appeared in fifteenth-century Kerala, where Madhava of Sangamagrama and his successors, obtained infinite series for the sine, cosine, and arctangent functions by subdividing an arc into increasingly many small segments, summing associated geometric quantities, and passing to the limit with infinitely many segments. The critical breakthrough came as the epiphany that the problems of quadrature and tangents were entwined. James Gregory and Isaac Barrow articulated special cases of this reciprocity geometrically. Soon, Newton and Leibniz developed a unification of the concepts of area under the curve and instantaneous rate of change embodied in Theorem 1.
Theorem 1 (Newton-Leibniz) Let \(f\) be an integrable real-valued function and \(g,h\) be differentiable real-valued functions then: \[ \frac{d}{dx}\int_{h(x)}^{g(x)}{f(t) dt} = f(g(x))g'(x) - f(h(x))h'(x) \]
Remark. Seen differently, taking \(h(x) = a, g(x) = x\) and sending \(f \to f'\) we have \(\frac{d}{dx}\bigl(\int_{a}^{x}{f'(t) dt} - f(x)\bigr) = 0\) which is a differential equation with only constant solutions, and since \(\int_{a}^{a}{H(t)dt}=0\) for any \(H\), we have: \(\int_{a}^{x}{f'(t) dt} = f(x) - f(a)\) and evaluating at \(x = b\) we get: \[ \int_{a}^{b}{f'(x) dx} = \int_{a}^{b}{\frac{df}{dx} dx} = \int_{[a,b]}{df} = f|^{b}_{a} \] Setting aside, tentatively, the “cancellation” of \(dx\) as mere notational convenience, latter equality expresses the quantities under concern in a coordinate-free manner on the object integrated over (the interval). Indeed, we may write that Theorem 1 tells us that integrating \(df\) over an elementary set (union of finitely many intervals) in \(\mathbb{R}\) boils down to an alternating sum of evaluations of \(f\) on the boundary points.
Corollary 1 (Multi-variable variant) Let \(f: \mathcal{D} \to \mathbb{R}\) be a scalar field over \(\mathcal{D} \subseteq \mathbb{R}^n\), and be any curve \(C\) with a differentiable parametrization \(\gamma: [0,1] \to \mathbb{R}^n\) then: \[ \int_{C}{\nabla f \cdot d\mathbf{l}} = \int_{0}^{1}{\nabla f(\gamma(t)) \cdot \mathbf{\gamma}'(t) dt} = f \circ \gamma|^{1}_{0} \]
Remark. Follows from \(\frac{d}{dt}(f \circ \gamma)\) and the “law of total derivative”. Alternatively, by first principles, one can re-write the difference quotient involved in the derivative as the sum of difference quotient involving single-variable moves.
Excluding seminal results, the resemblance between the calculus of Newton and the elementary calculus of today is superficial at best, with the former owing much to an intuitive use of infinitesimals which was evidently inconsistent to such an extent that the immaterialist philosopher George Berkeley likened them to religious doctrine in his publication “The Analyst: A Discourse Addressed to an Infidel Mathematician”. Posterity is witness to the salubrity of such works in impelling the “arithmetization” of calculus under Cauchy and Weierstrass, reforming the system under the stalwart concepts of limits, \(\epsilon\delta\)-proofs and real analysis. Despite criticism, the relentless quest of physics to understand the universe had not been delayed for the want of analysis, having already tasted the fruits of this “doctrine” in mechanics, and Newton’s theory of gravitation.
The study of heat conduction, elasticity, gravitation, fluid flow and electromagnetism in the eighteenth and nineteenth centuries required a calculus of vector fields defined over regions of space, and specifically demanded a way to relate quantities inside a region to its behavior on the region’s boundary. Gauss’ laws, which reformulate inverse-square laws for gravitational and electric fields relating flux through a surface directly to an enclosed quantity such as mass or charge, are prime examples. Gauss’ divergence theorem (a weakening of which is Lemma 1), Green’s theorem and Stokes’ theorem grew out of this motivating body of work in a short period of time.
Lemma 1 Let \(R: \mathbb{R}^3 \to \mathbb{R}\) be a differentiable scalar field and \(V = \{(x,y,z) \in \mathbb{R}^3 | (x,y) \in \mathcal{D}, 0 \leq z \leq f(x,y)\}\) be the volume enclosed by the graph of \(f\) over \(\mathcal{D}\), with surface \(\partial V\) then: \[ \int_{V}{\frac{\partial R}{\partial z} dV} = \int_{\partial V}{R \hat{\mathbf{k}} \cdot d\mathbf{S}} \]
Proof. The convention for surface integrals is to choose the outward pointing normal. For the surface defined by \(f\) we get normals: \(\mathbf{n} = -\frac{\partial f}{\partial x}\hat{\mathbf{i}} -\frac{\partial f}{\partial y}\hat{\mathbf{j}} + \hat{\mathbf{k}}\) \[ \iint_{D}{\int_{0}^{f(x,y)}{\frac{\partial R}{\partial z} dz} dy dx} = \iint_{D}{R(x, y, f(x,y)) - R(x,y, 0) dy dx} = \int_{S}{R\hat{\mathbf{k}}\cdot d\mathbf{S}} + \int_{-P}{R\hat{\mathbf{k}}\cdot d\mathbf{S}} \] where \(S\) is the surface defined by \(f\), and \(P\) is the xy-plane restricted to \(\mathcal{D}\) with orientation such that normal is upward. It is clear that vertical faces in \(\partial V\) contribute nothing to the surface integral, so the lemma is established.
Remark. We may conclude that Lemma 1 holds for all unidirectional vector fields since the quantities under consideration are independent of coordinates and we may always choose rectilinear coordinates \((x, y, z)\) such that the field is aligned with \(\hat{\mathbf{k}}\). Additionally, we can extend the result to volumes enclosed by two surfaces (given by \(f_1\), \(f_2\)) glued on their common edge by applying Lemma 1 on \(R, f_1\) and \(-R, -f_2\) and summing: \[ \int_{S_1}{R \hat{\mathbf{k}} \cdot d\mathbf{S}} + \int_{-P}{R \hat{\mathbf{k}} \cdot d\mathbf{S}} + \int_{-S_2}{-R \hat{\mathbf{k}} \cdot d\mathbf{S}} + \int_{-P}{-R \hat{\mathbf{k}} \cdot d\mathbf{S}} =\int_{S_1 + S_2}{R \hat{\mathbf{k}} d\mathbf{S}} \] Indeed, we can take this a step further, noticing that carving out a cuboid in the interior of the volume leaves the flux unchanged, and that the lemma readily applies to cuboids. So, for a closed surface formed by “gluing together” graphs, the theorem applies. A general vector field can be broken down into orthogonal components, and thus Gauss’ theorem for divergence follows by linearity of divergence for such 3D surfaces:
\[ \int_{V}{\nabla \cdot \mathbf{D} dV} = \int_{\partial V}{\mathbf{D} \cdot d\mathbf{S}} \]
Cauchy, three years prior to Green, had begun investigations into Complex Analysis, and discovered the celebrated Cauchy’s Formula for Contour integrals, based on a key result:
Lemma 2 (Cauchy, 1825) Let \(\mathcal{D} \subseteq \mathbb{C}\) be the interior of a contour \(C\), such that \(\forall z \in \mathcal{D}, f \text{ complex-differentiable at } z\) then: \[ \oint_{C}{f(z) dz} = 0 \]
Remark. The superficial similarity between this beautiful result of Cauchy and that of Corollary 1 for a closed loop is tantalizing. This is yet another instance of the behaviour of a function on the interior determining boundary phenomenon.
Gauss’s Disquisitiones generales circa superficies curvas first demonstrated, in 1827, that a surface’s curvature could be characterized entirely by measurements internal to the surface itself, independent of ambient space. Riemann generalized this intrinsic viewpoint to spaces of arbitrary dimension equipped with a metric, supplying the notion of “space” abstracted entirely from any embedding. Hermann Grassmann’s exterior algebra furnished the algebraic apparatus of alternating multilinear forms and wedge products. Élie Cartan synthesized Grassmann’s exterior algebra, Riemannian geometry, and the emerging theory of smooth manifolds, to reformulate the differential geometry of curved spaces in terms of exterior calculus, motivated in no small part by the demands of General Relativity to be rid of the grime of coordinate-based descriptions.
The condensed history of calculus presented here, omits, by necessity of brevity, doubtless multifarious developments such as alternative energy-based methods of Lagrangian mechanics that increasingly drove the need to apply calculus to “higher-order objects” such as functionals, leading to Calculus of Variations or the formalization of infinitesimals in hyperreals under Non-Standard Analysis. The purpose of such a long-winded presentation is to highlight among these gems of calculus, the recurring theme of being able to move from the boundary to interior as the domain of integration while swapping the integrand to a corresponding term with one more “order of derivative”. These gems of calculus, although related by this common theme do not overtly appear to be connected, yet the notion of behaviour on the boundary being determined by variation within the interior prevails intuitively, instilling a desire for their unification under an elegant framework that formalizes this intuition.
This article is unusual in its design, in that it is devoted exclusively to the presentation, as juxtaposed with the development, of the celebrated crown jewel of this body of work due to Cartan, known today as the Generalized Stokes’ Theorem, with the sole purpose of inspiring an intuitive understanding of this beautiful result, solidifying its connections with familiar facts, and obtaining an operational facility in applying it. The diligent reader is referred to the numerous established texts for the purpose of attaining a pragmatic, flexible and deep working knowledge of this fundamental theorem of calculus.
Manifolds
Geometry is, in a word, the study of figures. The properties of figures, being of great importance to us and owing to the infinity of conceivable figures necessitates a taxonomy based on invariants and structure. The topological manifold, is one such label on figures, that articulates among a few topological properties, the geometric idea of small enough patches of the figure being effectively flat.
Definition 1 (Topological Manifold) A second-countable, Hausdorff topological space \((M, \mathcal{T})\) with an open cover \(\{U_{\alpha}\}_{\alpha \in A}\) is an n-dimensional topological manifold, if there exist homeomorphisms \(\psi_{\alpha}: U_{\alpha} \to V_{\alpha}\) where \(V_\alpha\) is open in \(\mathbb{R}^n\).
The ordered pairs \((U_{\alpha}, \psi_{\alpha})\) are termed local charts, and the collection \(\{(U_{\alpha}, \psi_{\alpha})\}_{\alpha \in A}\) is termed an atlas.
Remark. The topological properties are fairly intuitive: a Hausdorff topological space is one in which any two distinct points can be “isolated” in disjoint open sets containing one of them each. A space is second-countable if there is a countable collection of open sets, such that every neighbourhood of a point contains a sub-neighbourhood from that collection. This in turn implies that any open set can be written as a union of countably many open sets from that same collection. This is analogous to how spheres of rational radius on rational lattice points can cover \(\mathbb{R}^n\). These properties disqualify pathological constructions from our consideration. The more geometrically relevant characteristic is the existence of \(\psi_{\alpha}\). A homeomorphism is simply a continuous map with a continuous inverse, but of course the notion of continuity is contingent on the domain, codomain topologies. This additional structure, along with its cover \(U_{\alpha}\), formalizes our intuitive understanding of being able to build surfaces by “gluing” pieces of paper. Indeed, in 3D, a homeomorphism of \(\mathbb{R}^2\) is equivalent to bending, folding, or stretching a piece of paper, and its mapping into the cover \(U_{\alpha}\) is the aspect of pasting these pieces of paper together to form the surface. Indeed, second-countability may be taken as considering only higher-dimensional analogoues of this process where only countably-many pieces of paper are used. It is noteworthy, that in essence the definition presents effecitvely an opaque type \(M\), which can only be inspected or accessed via \(\psi_{\alpha}\) which are inherently myopic.
Each local chart can be envisaged as local maps (in the cartographic sense) with custom coordinate systems. In our requirements in Definition 1 we neglected to consider \(U_{\alpha}\) disjoint, consequently we may have topological manifolds where \(x \in U_{\alpha} \cap U_{\beta}\). We will then have to investigate the transition map: \[ f_{\alpha \beta}: \psi_{\alpha}(U_\alpha \cap U_\beta) \to \psi_{\beta}(U_\alpha \cap U_\beta), x \mapsto (\psi_{\beta} \circ \psi_{\alpha}^{-1})(x) \] It follows from Definition 1 that \(f_{\alpha \beta}\) is a homeomorphism. Evidently, if we wish to make the observation of properties such as continuity, differentiability on a function \(f: M \to \mathbb{R}\) consistent with multiple viable local charts, it is sufficient to imbue the transition maps \(f_{\alpha \beta}\) with the corresponding properties.
Definition 2 (\(C^k\) with respect to an atlas) A function \(f: M \to \mathbb{R}\) is \(C^k\) with respect to a chart \((U_\alpha, \psi_\alpha)\) of an n-dimensional topological manifold if \(f \circ \psi_{\alpha}^{-1} \in C^k(\mathbb{R}^n)\). For an atlas \(\mathcal{A} = \{U_\alpha, \psi_\alpha\}_{\alpha \in A}\) with all transition maps \(f_{\alpha\beta} \in C^k(\mathbb{R}^n)\) \(f\) is \(C^k\) with respect to \(\mathcal{A}\), iff \(f\) is \(C^k\) with respect to every local chart.
In this manner, the choice of atlas \(\mathcal{A}\) seems to control “smoothness” on the manifold. We may define an equivalence relation between \(C^k\)-atlases \(\mathcal{A}, \mathcal{B}\) if \(\mathcal{A} \cup \mathcal{B}\) is a \(C^k\)-atlas. The union operator here is tantamount to requiring inter-atlas transitions being \(C^k\) themselves. We can consider the equivalence class \([\mathcal{A}]\) generated by \(\mathcal{A}\). This is closed under union, consequently, we can identify a maximal \(C^k\)-atlas \(\mathcal{A}^{*}\)
Definition 3 (\(C^k\)-structure) An n dimensional topological manifold \(M\) is said to have \(C^k\)-structure generated by an atlas \(\mathcal{A}\) when it is equipped with the maximal atlas \(\mathcal{A}^{*} \in [\mathcal{A}]\)
Remark. When talking about the \(C^\infty\) case, this is commonly referred to as smooth structure We can consider the category of all \(C^k\)-manifolds \((M, \mathcal{A}^{*})\), where morphisms are given by \(C^k\)-maps. This gives us a notion of isomorphic smooth structures. In particular, for two \(C^k\)-structures \([A], [B]\) we have \([A] \cong [B]\) if there exists an autohomeomorphism \(h: M \to M\) such that for all \((U_a, \psi_a) \in A, (U_b, \psi_b) \in B\), the following maps, when defined, are \(C^k\): \[ \begin{align} \psi_a \circ h \circ \psi_b^{-1}: \psi_b(U_b \cap h^{-1}(U_a)) \to \psi_a(U_a) \\ \psi_b \circ h^{-1} \circ \psi_a^{-1}: \psi_a(U_a \cap h(U_b)) \to \psi_b(U_b) \end{align} \] To round out this discussion of smooth structures, let us examine two incompatible atlases that nonetheless generate isomorphic structures. Consider the real line \(\mathbb{R}\) as the manifold of choice under the identity map. Take the first atlas as \(\mathcal{A} = \{(\mathbb{R}, \text{id})\}\). For the second one, let us choose a homeomorphism of \(\mathbb{R}\) that is not a diffeomorphism. For instance, \(g(x) = x^3, g^{-1}(x) = x^{\frac{1}{3}}, (g^{-1})' = \frac{1}{3}x^{\frac{-2}{3}}\) which is not differentiable at 0. Thus the second chart is \(\mathcal{B} = \{(\mathbb{R}, g)\}\). \(f_{AB} = g(x)\) checks out, but \(f_{BA} = g^{-1}\) fails to be \(C^1\). Let us take the identity function on \(\mathbb{R}\). This is a smooth function in the conventional sense on our manifold \(\mathbb{R}\), yet, it is not smooth on \(\mathcal{B}\). Nonetheless, taking \(h = g\), shows us that the two are isomorphic smooth structures.
Example 1 (Orthographic Projection and Hemisphere charts are compatible.) Consider the sphere in n-dimensions \(\mathbb{S}^{n-1}\). This is the solution set of the equation \(\sum_{i=1}^{n}{x_i^2} = 1\). One way to cover it is via \(2n\) charts, which are graphs of \(x_i = \pm\sqrt{1-\sum_{j \neq i}{x_j^2}}\) on the interior of their well-defined domain. This can be seen by induction. Choosing \(x_n\) first, the leftover points of \(\mathbb{S}^{n-1}\) are themselves a \(\mathbb{S}^{n-2}\). Another approach would be to select a pair of antipodal points as poles, and then draw secants from the pole to every other point. Such a secant will also intersect their bisecting plane, and thus define a mapping from a point on the sphere to a point on the plane. In this way, all but the pole is mapped into the chart. We map this pole, by using the corresponding chart from the other pole. As \(\mathbb{S}^{n-1}\) is Hausdorff, any set formed by removing exactly one point is open. Formally, the map under question is: \[ \psi: \mathbb{S}^{n-1} \setminus \{(\mathbf{0}, 1)\} \to \mathbb{R}^{n-1}, (x_1, ..., x_n) \to \bigg(\frac{x_1}{1-x_n}, ..., \frac{x_{n-1}}{1-x_n}\bigg) \] Which readily inverts as: \[ \psi^{-1}: \mathbb{R}^{n-1} \to \mathbb{S}^{n-1} \setminus \{(\mathbf{0}, 1)\}, (x_1, ..., x_{n-1}) \to \bigg(\frac{2x_1}{\sum{x_i^2}+1}, ..., \frac{\sum{x_i^2}-1}{\sum{x_i^2}+1}\bigg) \] Clearly, transitions from this chart to the earlier graph-like atlas are smooth since it is effectively composition with a projection. Likewise, we need only check the smoothness of functions: \[ \frac{\sqrt{1-\sum_{j\neq i}{x_j^2}}}{1\pm x_n} \text{ and } \frac{x_i}{1\pm\sqrt{1-\sum{x_j^2}}} \] which have all derivatives when \(x_n \neq \pm 1\) (resp).
Remark. Smooth atlases make precise the intuitive requirement for a set of maps (cartographic) to constitute an atlas useful for navigation. Adding a map in which tracking a moving point or the current location is so contrived that a small displacement easily tracked in one corresponds to jumps in position on the other maps in the atlas is seldom of any value.
All of this is to say that, by choosing a particular kind of smooth structure from an “obvious” atlas, functions we define in terms of coordinates on the manifold will remain smooth when switching to any other atlas from the same smooth structure. So far, the discussion has revolved around the properties of diffeomorphism and \(C^k\) functions. However, \(C^1\) maps do not always admit a non-trivial explicit form for their inverse. For the purposes of Definition 1, local regularity is sufficient. One convenient method, supported by Theorem 2, is to check the invertibility of Jacobian on the desired domain which boils down to examining the roots of the determinant.
Theorem 2 (Euclidean Inverse Function Theorem) Let \(f: A \to \mathbb{R}^n\) be a \(C^1\)-map on \(A \subseteq \mathbb{R}^n\), and \(\mathbf{x_0} \in A^\circ\) have \(\text{det }\mathbf{J}f(\mathbf{x}_0) \neq 0\), then there exists an open set \(U \subseteq A\) such that \(f|_{U}\) is a \(C^1\)-diffeomorphism. In particular for \(g = (f|_U)^{-1}, \mathbf{y} = f(\mathbf{x})\), we have \(\mathbf{J}g(\mathbf{y}) = (\mathbf{J}f)^{-1}(\mathbf{x})\)
Proof. It is sufficient to prove the former claim when \(\mathbf{J}f\) is invertible in \(U\) as \(\mathbf{J}(g \circ f)(\mathbf{x}) = \mathbf{J}g(f(\mathbf{x}))\mathbf{J}f(\mathbf{x}) = I\) by the chain rule applied on Jacobian, which implies \(\mathbf{J}g(\mathbf{y}) = (\mathbf{J}f)^{-1}(\mathbf{x})\). The properties of the claim are invariant under translation and invertible linear transformation, send \(f(\mathbf{x}) \leftarrow (\mathbf{J}f(\mathbf{x}_0))^{-1}(f(\mathbf{x} + \mathbf{x}_0)-f(\mathbf{x}_0))\) so that \(\mathbf{x_0} \leftarrow \mathbf{0}, \mathbf{J}f(\mathbf{0}) = I\) without loss of generality. Set \(r(\mathbf{x}) = f(\mathbf{x}) - \mathbf{x}, T_{\mathbf{y}}(\mathbf{x}) = \mathbf{y} -r(\mathbf{x})\) to attempt a fixed-point iteration method. For an \(\varepsilon \in (0, 1)\) we may find open sets such that: \(\text{det }\mathbf{J}f \in (1-\varepsilon, 1+\varepsilon)\), and \(||\mathbf{J}r||_{F}^2 < \frac{\varepsilon^2}{n}\) since continuity is closed under ring operations. A non-empty ball \(B(0, \rho)\) satisfies all conditions since the corresponding open sets \(U_1, U_2\) have \(\mathbf{0} \in U_1 \cap U_2 \in \mathcal{T}(\mathbb{R}^n) \implies B(0, r) \subseteq U_1 \cap U_2\). The latter condition extends to gradient of each component \(r_i = \pi_i \circ r\) by non-negativity of norm so \(\forall i \in \mathbb{N}_{\leq n}, ||\nabla r_i|| < \frac{\varepsilon}{\sqrt{n}}\). Applying Corollary 1 and Cauchy-Schwarz, we get \(r_i(\mathbf{x}) < \frac{\varepsilon}{\sqrt{n}}||x||\). This implies, \(||r(\mathbf{x})||^2 = \sum{||r_i(\mathbf{x})||^2} < \varepsilon^2 ||x||^2\). So we have \(||r(\mathbf{x})|| < \varepsilon ||\mathbf{x}||\). By the same argument, choosing any other \(\mathbf{x'} \in B(0, \rho)\), we can obtain \(||r(\mathbf{x})-r(\mathbf{x'})|| < \varepsilon ||\mathbf{x} - \mathbf{x'}||\).
Now suppose we chose \(\mathbf{y} \in B(0, (1-\varepsilon)\rho)\), then \(||T_{\mathbf{y}}(\mathbf{x})|| \leq ||\mathbf{y}|| + ||r(\mathbf{x})|| < \rho\). Thus, \(T_{\mathbf{y}}|_{\overline{B}(\mathbf{0}, \rho)}\) is a contraction, which admits a unique fixed point \(\mathbf{x}^{*} = T_y(\mathbf{x}^{*})\) by the Banach Fixed-point Theorem on complete metric spaces. By construction, the fixed point \(\mathbf{x}^{*}\) belongs to the preimage of \(\mathbf{y}\) under \(f\). Since \(f\) is continuous, the preimage \(D = f^{-1}(B(\mathbf{0}, (1-\varepsilon)\rho))\) is open, with \(\mathbf{0} \in D \cap B(0, \rho) = U\). We have found an open set \(U\) on which \(f|_U\) is invertible. We must now argue that the inverse is \(C^{1}\). By applying both the triangle inequality and reverse-triangle inequality on the expression: \(||(r(\mathbf{x}) - r(\mathbf{x}')) - (\mathbf{x}'-\mathbf{x})||\) we can obtain:
\[ \begin{align} (1-\varepsilon)||\mathbf{x}-\mathbf{x'}|| < ||f(\mathbf{x}) - f(\mathbf{x'})|| < (1+\varepsilon)||\mathbf{x}-\mathbf{x'}|| \\ (1-\varepsilon)||g(\mathbf{y})-g(\mathbf{y'})|| < ||\mathbf{y} - \mathbf{y'}|| < (1+\varepsilon)||g(\mathbf{y})-g(\mathbf{y'})|| \end{align} \] In particular, \(g = (f|_U)^{-1}\) is Lipschitz with constant \(\frac{1}{1-\varepsilon}\). Now move to another pair \(\mathbf{x}, \mathbf{y} = f(\mathbf{x})\) in the domain. We repeat the aforementioned argument for \(r(\mathbf{x})\) on \(R(\mathbf{x'}) = f(\mathbf{x'}) - f(\mathbf{x}) - \mathbf{J}f(\mathbf{x})(\mathbf{x'} -\mathbf{x})\) so for any \(\eta \in (0,1), U' \subseteq U\), we have \(||R(\mathbf{x'}) - R(\mathbf{x})|| < \eta ||\mathbf{x'} - \mathbf{x}||\) Now, let \(\mathbf{y}_t = \mathbf{y} + t\mathbf{e}_i\), then: \[ \mathbf{y_t} - \mathbf{y} = \mathbf{J}f(\mathbf{x})[g(\mathbf{y_t}) - g(\mathbf{y})] + R(g(\mathbf{y_t})) \] Which implies: \[ [\mathbf{J}f(x)]^{-1}\mathbf{e}_i - \frac{1}{t}g(\mathbf{y_t}) - g(\mathbf{y}) = [\mathbf{J}f(\mathbf{x})]^{-1}\frac{||g(\mathbf{y_t}) -g(\mathbf{y})||}{t}\frac{R(g(\mathbf{y_t}))}{||g(\mathbf{y_t})-g(\mathbf{y})||} \] We can now apply the Lipschitz property of \(g\), limit-like property of \(R(\mathbf{x'})\) and the following observation: \[ ||\mathbf{J}f(\mathbf{x}) \mathbf{v}|| = ||\mathbf{J}r(\mathbf{x})\mathbf{v} + \mathbf{v}|| > \bigg(1 - \frac{\varepsilon}{\sqrt{n}}\bigg)||\mathbf{v}|| \] Send \(\mathbf{v} \to \mathbf{J}f(\mathbf{x})\mathbf{v}\) so that: \(||[\mathbf{J}f(\mathbf{x})]^{-1}\mathbf{v}|| < (1 - \frac{\varepsilon}{\sqrt{n}})^{-1}||\mathbf{v}||\) Finally, \[ ||[\mathbf{J}f(x)]^{-1}\mathbf{e}_i - \frac{1}{t}g(\mathbf{y_t}) - g(\mathbf{y})|| < \bigg(1 - \frac{\varepsilon}{\sqrt{n}}\bigg)^{-1}(1-\varepsilon)^{-1}\eta \] As before, the inequality extends to the components, and we have that: \[ \lim_{t \to 0}{\frac{g_j(\mathbf{y} + t\mathbf{e}_i) - g_j(\mathbf{y})}{t}} \text{ exists} \] Thus \(\mathbf{J}g\) exists, and by the argument at the start of the proof, \(g \in C^1\)
Remark. The proof presented here is rather tedious, with the geometric intuition not immediate from this rather dry presentation. The diligent reader is suggested to seek out better proofs. Interestingly, we can bootstrap this result to show that the regularity of \(f\) is preserved.
Corollary 2 (\(C^k\) Euclidean Inverse Function) Let \(f: A \to \mathbb{R}^n\) be a \(C^k\)-map on \(A \subseteq \mathbb{R}^n\), and \(\mathbf{x_0} \in A^\circ\) have \(\text{det }\mathbf{J}f(\mathbf{x}_0) \neq 0\), then there exists an open set \(U \subseteq A\) such that \(f|_{U}\) is a \(C^k\)-diffeomorphism. In particular for \(g = (f|_U)^{-1}, \mathbf{y} = f(\mathbf{x})\), we have \(\mathbf{J}g(\mathbf{y}) = (\mathbf{J}f)^{-1}(\mathbf{x})\)
Proof. Suppose \(g \in C^r\), for some \(r < k\), the Jacobian \(\mathbf{J}g\) is a composition of \((\mathbf{J}f)^{-1}\) and \(g\) which are both \(C^r\), so consequently, \(g \in C^{r+1}\). By Theorem 2, we have the base case \(r=1\), and the claim follows by induction.
Corollary 3 (Euclidean Implicit Function Theorem) Let \(F: A \to \mathbb{R}^m\) be a \(C^1\)-map on \(A \subseteq \mathbb{R}^{n-m}\times \mathbb{R}^m\) notated as \(F(\mathbf{x}, \mathbf{y})\), with \(F(\mathbf{x}_0, \mathbf{y}_0) = \mathbf{c}_0\) for some \((\mathbf{x}_0, \mathbf{y}_0) \in A^\circ\) such that \(\operatorname{det}\mathbf{J}F_{\mathbf{x}_0}(\mathbf{y}_0) \neq 0\) then there exist open sets \(U \subseteq \mathbb{R}^{n-m}, V \subseteq \mathbb{R}^m\) with \((\mathbf{x}_0,\mathbf{y}_0) \in U \times V\) and a \(C^1\)-map \(g: U \to V\) such that \(\forall \mathbf{x} \in U, F(\mathbf{x}, g(\mathbf{x})) = \mathbf{c_0}\). In other words, \(F^{-1}(\mathbf{c_0}) \cap U \times V\) is the graph of \(g\) on \(U\).
Proof. Construct the map \(\phi(\mathbf{x}, \mathbf{y}) = (x, F(\mathbf{x}, \mathbf{y}))\). We know that the following determinant is non-zero since it is lower triangular: \[ \operatorname{det}\mathbf{J}\phi(\mathbf{x_0, y_0}) = \operatorname{det} \begin{bmatrix} I & \mathbf{0} \\ \mathbf{J}F_y(\mathbf{x_0}) & \mathbf{J}F_x(\mathbf{y_0}) \end{bmatrix} \neq 0 \] We can then setup the inverse \(\phi^{-1}(\mathbf{x}, \mathbf{c})\) by Theorem 2 over some open domain \(U \times V \in \mathcal{T}(\mathbb{R}^{n-m}\times\mathbb{R}^{m})\). By construction, \(\pi_\mathbf{x} \circ \phi^{-1} = \text{id}_{\mathbf{x}}\). Fix \(\mathbf{c} \leftarrow \mathbf{c_0}\), so \(g = \pi_{\mathbf{y}} \circ \phi^{-1}|_{\mathbf{c_0}}\) is a function of \(\mathbf{x} \in U\) with outputs \(\mathbf{y}\) in the preimage of \(U\) which is open.
Remark. The regularity is preserved due to Corollary 2, and the fact that all compositions are with trivially smooth functions. The striking resemblance between Corollary 3 and Theorem 2 is due to their logical equivalence, which follows from the equivalence of the problems \(x = f^{-1}(y)\) and \(f(x) - y = 0\).
The graph of a \(C^k\) real-valued function on \(U \subseteq \mathbb{R}^n\) is a \(C^k\)-submanifold of \(\mathbb{R}^{n+1}\). More precisely, \(\Gamma(f) = \{(\mathbf{x}, f(\mathbf{x})) \in \mathbb{R}^{n+1} | \mathbf{x} \in U\}\) equppied with subspace topology and the atlas \(\{(\Gamma(f), (\mathbf{x}, f) \to \mathbf{x})\}\) comprises a \(C^k\)-submanifold, inherting the structure from \(\mathbb{R}^{n+1}\), since the coordinate map retains regularity. We may now propose a sufficient condition for the solution set of a system of equations to constitute a smooth manifold, following on the intuitive conclusions that can be drawn from Corollary 3.
Corollary 4 (Smooth Manifold from System of Equations) Let \(f: U \to \mathbb{R}^m\) be a \(C^{\infty}\)-map corresponding to a system of \(m\) equations on \(n\) variables drawn from \(U \in \mathcal{T}(\mathbb{R}^n)\) represented as: \[ \begin{align*} f_1(x_1, x_2, ..., x_n) = c_1 \\ f_2(x_1, x_2, ..., x_n) = c_2 \\ \vdots \\ f_m(x_1, x_2, ..., x_n) = c_m \end{align*} \] such that \(f(\mathbf{x}) = \mathbf{c} \implies \operatorname{rank}\mathbf{J}f(\mathbf{x}) = m\), then \(f^{-1}(\mathbf{c})\) is a smooth manifold of dimension \((n-m)\).
Example 2 (Quadric surfaces) Let \(A \in \mathrm{Sym}_n(\mathbb{R})\), \(\mathbf{x} = [x_1, x_2, ..., x_n]\) be the cartesian coordinates, and \(c \in \mathbb{R} \setminus \{0\}\), then the solution set of the following equation is a smooth manifold or empty: \[ \mathbf{x}^TA\mathbf{x} = c \] since a zero gradient, i.e \(2A\mathbf{x} = \mathbf{0}\), would only occur on points outside of the level set. This readily supplies us with plenty of familiar surfaces are smooth manifolds. In particular, the n-dimensional (referring here to ambient space) sphere \(\mathbb{S}^{n-1}\) is a smooth manifold, as are ellipsoids, and hyperboloids and their open submanifolds like hemispheres. But not cones, since \(c = 0\).
Now let us indulge in some idle curiosity. By the Spectral Theorem, \(A\) admits an eigenvalue decomposition \(A = V \Sigma V^T, \Sigma = \text{diag}(\lambda_i), V \text{ orthogonal}\). If we wish to identify surfaces described by the aforementioned equation upto scaling, we can fix \(c = 1\), and identify them via \(\mathbf{\hat{\lambda}} = \frac{\mathbf{\lambda}}{||\mathbf{\lambda}||} \in \mathbb{S}^{n-1}\). It is clear the decomposition is equivariant, consequently, we frame the quadratic surface by a choice of an order for \(\lambda_i\) and the eigenvectors (the frame-basis). This is analogous to the situation in classical coordinate geometry, where one has to find a quantity intrinsic to the geometry, say for instance the area, of a tilted ellipse by proposing a rotation of the axes that reduces it to a “standard” form, whereby the original configuration from the problem is made immaterial.
Example 3 The set of all (nonzero level-set, origin-centered, framed) quadratic \(n\)-dimensional (intrinsic) surfaces, upto scaling and choice of frame-basis, is itself a smooth manifold of \(n\) dimensions isomorphic (diffeomorphic) to \(\mathbb{S}^{n}\), and upto rotation, isomoprhic to the orbifold quotient \(\mathbb{S}^{n}/S_{n+1}\)
Remark. This is a rather pleasing and poetic conclusion. Just as a transient point may move smoothly on set of points (surface), so can a transient surface move smoothly in the set of all such surfaces. It is revealing that the smooth structure cannot throw away frame information. This is largely due to the fact that we yet distinguish two surfaces with the same set of eigenvalues if they are permuted. If we fix an order, say, descending, we effectively have a “canonical labelling” of the “axes” of the surfaces under consideration. But this cannot be meaningfully defined intrinsically, especially due to possible symmetries of the object under consideration. Symmetry here is essentially an inability to distinguish. This matter can be evinced clearly by considering ellipsoids. If the ellipsoid is “fully asymmetric”, then the surface itself, intrinsically, suggests a labelling of axes as principal, secondary, etc. However, suppose two lengths are equal, then a cross section of the ellipsoid is a circle. What axes shall we choose here, and how shall we distinguish them by labels? An entire family of rotations is available here. In the event that eigenvalues \(\lambda_i\) are distinct, we can choose any point \(\lambda^{'}\) in the open box \(|{\lambda}_i^{'} - {\lambda}_i| < \frac{1}{2}\min{|\lambda_i -{\lambda}_j|}\), retain the same labelling of axes, and move to it smoothly. However, if \(\lambda_i = \lambda_j\) for some \(i \neq j\), then every open box would contain a point such that our labels would swap, even if the surfaces themselves could smoothly morph. Therefore, it is the inability to maintain this canonical naming of axes, that makes the quotient object not smooth. This leads us directly to three fixes: remove the troublesome symmetric surfaces, or distinguish frames or embrace another kind of structure.
Example 4 (Matrix groups that are Smooth Manifolds) Consider \(\mathrm{GL}_n(\mathbb{R})\) which is the group of all invertible \(n\times n\) matrices. The set of all \(n \times n\) matrices is isomorphic to \(\mathbb{R}^{n^2}\), so it inherits smoothness naturally. The general linear group can be considered a preimage of the continuous function \(\operatorname{det}\), i.e, \(\mathrm{GL}_n(\mathbb{R}) = \operatorname{det}^{-1}(\mathbb{R} \setminus \{0\})\) which must be open, and therefore \(\mathrm{GL}_n(\mathbb{R})\) is a smooth open submanifold. Moreover, the set of orthogonal \(n \times n\) matrices, denoted \(\mathrm{O}(n)\) is also a smooth manifold. We shall evince this by the characterization that every member of \(\mathrm{O}(n)\) is an array of independent orthonormal row-vectors spanning \(\mathbb{R}^n\), which gives \(n + \binom{n}{2} = \frac{n(n+1)}{2}\) constraints. Suppose \([\mathbf{r}^T_1 ... \mathbf{r}^T_n]^T = A \in \mathrm{O}(n)\), we shall reason that the constraint gradients \(G^{ij} = \frac{\partial(\mathbf{r}_i \cdot \mathbf{r}_j)}{\partial A}\) are linearly independent. The \(k\)-th row of \(G^{ij}\), i.e, \(G^{ij}_{i}\) satisfies: \(G^{ij}_k = \delta_{ik}\mathbf{r_j} + \delta_{jk}\mathbf{r_i}\). Thus, any non-trivial linear combination of \(G^{ij}\) yields a linear combination of \(\mathbf{r}_i\). Consequently, by contrapositive, no such linear combination may exist. With this, the Jacobian for this system of equations has full rank, and \(\mathrm{O}(n)\) is, by Corollary 4, a smooth manifold of \(\frac{n(n-1)}{2}\) dimensions. The normal subgroup \(SO(n)\), is the special orthogonal group of orthogonal matrices with determinant 1. This is an open, path-connected submanifold of \(O(n)\), with an abelian quotient group \(O(n)/SO(n)\) of order 2.
The smooth structure on manifolds, allows us to speak of smooth operations and functions on elements of the manifold. This is particularly relevant when the manifold doubles as a group. In such cases, the smooth manifold is a Lie group when the group and inverse operations are both smooth.
Definition 4 (Lie Group) If a group \((G, *)\) is also a \(C^{\infty}\) manifold, and both the binary operation \(*: G \times G \to G\) of the group and the map \(g \mapsto g^{-1}\) of taking the inverse are \(C^{\infty}\), then G is called a Lie group.
Lemma 3 (Compact Exhaustion of Manifold) Let \(M\) be a topological manifold, then there exists a sequence of compact sets \((K_i)_{i \in \mathbb{N}}\) such that \(K_i \subseteq K_{i+1}^{\circ}\) and \(\cup{K_i} = M\).
Proof. Since \(M\) is second-countable, there exists a countable basis \(\mathcal{O} =\{O_i\}_{i \in \mathbb{N}}\). Fix a point \(p \in U \subseteq M\) in a chart \((U, \psi)\). There exists a neighbourhood \(N\) of \(\psi(p)\) such that \(p \in O_p \subset \psi^{-1}(\overline{N}) \subset U\) for some \(O_p \in \mathcal{O}\). Clearly \(\overline{O_{p}} \subseteq \psi^{-1}(\overline{N})\) is compact as it is a closed subset of the image of a compact set under a continuous map. We can now replace \(\mathcal{O}\) by the countable basis \(\mathcal{O}'\) found in this manner across all points of \(M\). Now set \(K_i = \overline{\cup_{j = 1}^{i}{O'_j}}\). For compact \(M\), the sequence has a tail of \(K_i = M\). The desired sequence is a subsequence of \(K_i\) so defined, skipping terms until \(K_{i_k} \subseteq K_{i_{k+1}}^{\circ}\), which must always be possible since \(K_i\) can be covered by finitely many \(O'_i\) and we may simply include open sets until the largest index.
A an open cover is said to be locally finite, if for every point there exists a neighbourhood which intersects only finitely many sets from the cover.
Proposition 1 (Locally finite open cover refinement on Manifolds) Let \(M\) be a topological manifold and \(\mathcal{U}\) be a countable open cover of \(M\). Then there exist countable open covers \(\mathcal{V}=\{V_i\}_{i\in I}\) and \(\mathcal{W}\) of \(M\) such that:
\[ \mathcal{U} \;\xleftarrow{\;\text{ref}\;}\; \begin{array}{c} \mathcal{V}\\[2pt] \text{locally finite}\\[4pt] \overline{V_i} \ \text{compact} \end{array} \;\underset{\text{ref}}{\stackrel{\exists\,\psi_i\ (\cong)}{\rightleftarrows}}\; \mathcal{W}=\{\psi_i^{-1}(B(\mathbf{0},1))\} \] where “ref.” indicates cover refinement, and \(\psi_i : V_i \to B(\mathbf{0}, 3)\) are homeomorphisms.
Proof. By Lemma 3, we have a sequence of compact sets \(K_{-1} = K_{0} = \emptyset, E_i = K_i \setminus K_{i-1}^{\circ}\) whose union gives \(M\). Since \(E_i\) is compact, only finitely many \(U_j \in \mathcal{U}\) are needed to cover it. Choosing small neighbourhoods \(p \in V_p \subseteq U_j \cap (K_{i+1}^{\circ} \setminus K_{i-2})\) (note \(K_{i+1}^{\circ} \setminus K_{i-2} \supseteq E_i\)) for each point in each \(U_j \cap E_i\). We can further restrict \(V_p\) to be the preimage of a ball with compact closure (similar to Lemma 3), using the corresponding coordinate map for the topological manifold, and select homeomorphisms \(\psi_p\) trivially. Now, constructing \(W_p\) and applying compactness we get finitely many \(W_j, V_j\) for each \(i\). Taking the unions of these families iteratively over all \(i \in \mathbb{N}\), we get \(\mathcal{V}\) and \(\mathcal{W}\). Evidently, \(\mathcal{V}\) is locally finite, since for any \(x \in K_i^{\circ}\), \(K_i^{\circ}\) intersects atmost finitely many members of \(\mathcal{V}\) from only the first \(i + 1\) iterations. Thus the diagram from the claim holds.
While studying manifolds, we would need to construct various objects, specifying them on local charts and “glue” them together. One simple method to reconcile these objects is to assign an affine combination to each point of the manifold. This idea is made precise by the parition of unity.
Definition 5 (Partition of Unity) Let \(M\) be a \(C^k\) manifold. A family of countably many \(C^k\)-functions \(\{f_i\}\) is called a partition of unity if:
- \(\forall i \in \mathbb{N}, f_i(p) \geq 0\).
- \(\{\operatorname{supp}{f_i}\}\) is locally finite, where \(\operatorname{supp}{g} = \overline{\{x \in X | f(x) \neq 0\}}\) is defined for any map \(g: X \to \mathbb{R}\) from a topological space.
- \(\forall p \in M, \sum_{i}{f_i(p)} = 1\).
If \(\{\operatorname{supp}{f_i}\}\) is the refinement of an open cover \(\mathcal{U}\), then the partition of unity \(\{f_i\}\) is said to be subordinate to \(\mathcal{U}\).
The existence of partitions of unity can be inferred directly from Proposition 1.
Proposition 2 For every open cover on a smooth manifold, there exists a partition of unity with compact support subordinate to it.
Proof. Applying Proposition 1, choosing \(g_i\) such that \(\operatorname{supp}{g_i} = \overline{V_i}\), and normalizing \(f_i = \frac{g_i}{\sum_{j}{g_j}}\) yields the partition of unity. It is easy to verify that this is indeed a valid construction. This follows via local finiteness of \(\mathcal{V}\), the fact that there exists a smooth function \(b: \mathbb{R}^n \to \mathbb{R}_{0}^{+}\) that satisfies \(b(B(\mathbf{0}, 1)) = \{1\}, b(\mathbb{R}^n \setminus B(\mathbf{0}, 3)) = \{0\}\), and the refinement \(\mathcal{W}\) via coordinate maps \(\psi_i\). For instance, such a function can be constructed by the “flat” function \(u(x) = \exp{(-x^{-1})}\mathbb{I}_{x > 0}\). Details are omitted here for brevity.
Remark. The proof carries over just as easily to the \(C^k\) case.