Every Formula Is a First Term
E = mc² is the first term of a series that never ends. So is F = ma, so is Hooke’s law, so is the ideal gas. That is not because the true law is infinitely long. It is because every smooth structure looks like its tangent from wherever you happen to be standing — and an expansion point is a frame. Where the series stops working, something is waiting on the axis you cannot see.
Contents
01 The most famous formula is a first term
Einstein never wrote E = mc² in the paper that is supposed to contain it. The September 1905 note in the Annalen der Physik is three pages long, and its conclusion is a sentence: if a body gives off energy L in the form of radiation, its mass diminishes by L/V², where V is the speed of light. No E, no c, no equals sign in the famous shape. The letter E does appear, but for the body’s energy in a chosen frame, not as the left side of a slogan.
I start there for the same reason Part Seven started with the fact that F = ma is not Newton’s. The famous formulas are not the things physicists discovered. They are the parts of what physicists discovered that fit on a T-shirt. And the way they got to fit is the subject of this lesson: each of them is the first term of something longer, cut off where the rest stopped mattering.
Here is the full expression Einstein’s note is actually about. The energy of a body of mass m moving at speed v is
The Greek letter γ (gamma) is just a name for that square root; Part Five showed it is the hyperbolic cosine of the rapidity, but you do not need that here. Now do what a calculus student does with any function near a point: expand it about v = 0. The binomial theorem gives
Look at what fell out. The first term is mc² — the rest energy, exact, the thing on the T-shirt. The second term is ½mv², which is Newton’s kinetic energy, the formula every schoolchild learns two years before they hear the word relativity. Newton’s mechanics is not a rival theory that Einstein overthrew. It is the second term of Einstein’s expression, and Einstein said so himself in that same September note, keeping the v² term and “neglecting magnitudes of fourth and higher orders.” Everything after the second term is what you pay for going fast: at a tenth of light speed the third term is three-quarters of one percent of the kinetic energy, at half of light speed it is nineteen percent, and near c the series needs every one of its infinitely many terms to reach the answer at all.
The most famous formula, and the series it begins
- truncated series
- exact
The same thing happens to F = ma. The exact law is F = dp/dt — force is the rate of change of momentum — and Part Four established that this equation is not approximate. What is approximate is the momentum inside it. Newton uses p = mv; the exact expression is p = γmv. Push a body along the line it is already moving on and the derivative works out to
So the brief this lesson began from — the great formulas are the beginning of an infinitely long one — is literally true for the two most famous formulas in physics. I am going to spend the rest of the lesson making it precise, because the precise version says something the slogan does not: the infinitely long formula is not the true one either.
02 Why the first term is always simple
The theorem underneath all of this is Taylor’s, and for this audience it is worth stating in full, because every symbol in it is going to be doing philosophical work later.
Take any function f that is smooth — meaning you can differentiate it as many times as you like — and any point a where you happen to know its value. Write f′(a) for its slope at that point, f″(a) for the slope of the slope, and so on. Then for x near a,
Read it as a sentence. Whatever the structure is, near where you stand it looks like its value, plus its tangent, plus a small correction for curvature, plus smaller corrections forever. Brook Taylor published it in 1715 without a remainder and without asking whether the sum converges; Lagrange supplied the remainder in 1797; James Gregory had the idea in private around 1671; and the Kerala school under Madhava had the series for sine and cosine by about 1400, more than two centuries before anyone in Europe. None of that history matters to the argument. What matters is the shape of the statement: it applies to every smooth function, and it says the first term after the constant is always linear.
That is why the great formulas are simple. Not because nature is generous, and not because early physicists were lucky, but because linear is what everything looks like up close. Hooke’s law — ut tensio, sic vis, as the extension so the force, first published in 1676 as an anagram so no one could steal it and decoded in 1678 — is not a discovery about springs. Take any potential energy V(x) with a stable resting point at x₀. At a minimum the slope is zero by definition, so the linear term of the Taylor series vanishes, and the first thing left standing is the quadratic:
Differentiate and the force is −k(x − x₀): Hooke. This is the reason the harmonic oscillator is in every chapter of every physics book. It is not one system among many; it is the second term of all of them. And the third term is not decoration. A crystal whose atoms sat in perfectly quadratic wells would not expand when heated — the well would be symmetric and the average position would never shift. Thermal expansion, the reason bridges have gaps in them, is the cubic term of the Taylor series of the interatomic potential.
The cleanest place to watch a first term fail is the pendulum, because you can measure it with string and a stopwatch. The textbook period T₀ = 2π√(L/g) comes from replacing sin θ by θ — keeping the first term of sine’s own Taylor series. The exact period is an elliptic integral, and its expansion in the amplitude θ₀ is
Where the textbook pendulum fails, and why the series fails at 180 degrees
- truncated series
- exact
Pull the bob back 23 degrees and the formula in the textbook is one percent slow. Pull it back to horizontal and it is eighteen percent short of the true period. Pull it all the way to the top and the series stops converging altogether — and it stops for a physical reason, which is the subject of the next section.
03 An expansion point is a frame
Here is the correction I want to make to the brief, and it is the point of the lesson. The slogan says the true formula is infinitely long. It is not. E = γmc² is six characters long. The infinite series is what that short expression looks like from the point v = 0. Choose a different point — expand about some cruising speed v₀ instead of rest — and you get a different infinite series for the same energy, with every single coefficient changed. The series is not the structure. The series is the structure’s coordinates in a basis of powers of (x − a), and the choice of a is the choice of basis.
Part Five put it this way for spacetime: a frame is a choice of basis, changing frames is a rotation, and the things that survive every rotation are the invariants. The same three sentences hold here word for word. An expansion point is a choice of basis. Moving the expansion point recomputes every coefficient. What survives is the function. So “E = mc² + ½mv²” is not a law and not an approximation to a law. It is the energy written in the coordinates of the rest frame — and in that case the metaphor is not even a metaphor, because v = 0 is the rest frame. For the pendulum, θ₀ = 0 is not a frame in the physicist’s sense, and I am using the word by analogy. I flag it once and will not apologise for it again: the mathematical claim (a Taylor series is a coordinate description relative to a point) is exact; the relativity vocabulary is borrowed.
The infinitely long formula is what a short structure looks like from where you are standing. The structure is the invariant. The series is the view.
Now for the part that connects to the series’ oldest vocabulary. Every Taylor series has a radius of convergence — a distance from the expansion point beyond which the sum stops meaning anything. And the theorem that fixes that radius is one of the strangest facts a student meets: the radius of convergence equals the distance to the nearest point where the function misbehaves — measured in the complex plane, even when the function, the point, and everything you care about are real.
Take f(x) = 1/(1 + x²). On the real line it is a gentle bump, smooth everywhere, never larger than one, nothing remotely wrong with it at x = 1. Expand it about zero and the series is 1 − x² + x⁴ − x⁶ + …, which converges for |x| < 1 and dies at |x| = 1, for no reason you can see on the real line. The reason is at x = ±i, where 1 + x² = 0 and the function goes to infinity. Those two poles sit on the imaginary axis, exactly one unit from the origin, and it is their distance that kills the real series at |x| = 1. Move the expansion point to x = 1 and the nearest pole is now √2 away, so the new series converges on a bigger disc. The function never changed. The view did.
The reason the real series dies is not on the real line
Readers of Part Three will recognise the shape of this. The thing that limits what you can see from the real line is not on the real line. You feel only its consequence — the series failing — and the cause is perpendicular to everything you can measure. The pendulum series dies at 180 degrees for the same kind of reason: a pendulum balanced exactly upside down takes forever to fall, so the period is infinite there, and that singularity sets the radius of the whole series even for a student who only ever swings the bob a few degrees. The failure of the first term far from home is caused by something at the edge of the world.
04 The catalogue
Once you know what to look for, the pattern is everywhere, and in every case the next term has been measured. That is the part that makes this a lesson in physics rather than a lesson in rhetoric. Here is the catalogue, with the expansion point named in each case, because naming the point is naming the frame.
Six first terms, their next terms, and the experiment that saw the next term
- F = maabout v = 0 · the rest frame × [1 + (3/2)v²/c² + …] Every particle accelerator ever built. The bracket is why the Large Hadron Collider cannot make a proton reach c however hard it pushes.
- Newtonian gravityabout weak field · GM/rc² → 0 first post-Newtonian correction Mercury’s perihelion: 43 arcseconds per century that Newton’s first term could not account for. Einstein, November 1915. Computed here from the formula: 42.99″.
- Hooke’s lawabout zero strain · the resting length + cubic term of the potential Thermal expansion. A perfectly quadratic well would not expand when heated (Grüneisen, 1912). Bridges have gaps because the third term exists.
- Ideal gas, PV = nRTabout zero density · V → ∞ PV/nRT = 1 + B(T)/V + C(T)/V² + … The virial expansion. Kamerlingh Onnes named the coefficients in 1901; for a van der Waals gas the second one is B = b − a/RT, and it has been tabulated for a century.
- Maxwell’s equationsabout zero field strength · E « m²c³/eħ + (2α²/45m⁴)[(E²−B²)² + 7(E·B)²] Light scattering off light — forbidden by Maxwell, predicted by the next term (Euler and Kockel, 1935), observed by ATLAS at CERN in 2019 with lead nuclei grazing past each other.
- Small-angle pendulumabout θ₀ = 0 · hanging still × [1 + θ₀²/16 + …] Any stopwatch. At 23 degrees the first term is already one percent out; see Figure 02.
Two honesty clauses before the catalogue turns into a slogan. First, not everything is a first term. E² = (pc)² + (mc²)² is not a truncation of anything; it is the Minkowski norm from Part Five, exact, and it is the thing that generates the series in section 01. Conservation laws, which are consequences of symmetry, are exact. Identities are exact. Part Seven’s claim that the Schrödinger equation is linear is a postulate, not an approximation. The pattern is: constitutive relations and dynamical laws expanded about a comfortable point are first terms; structural identities are not. Second, Hooke and Ohm are a slightly different case from gravity and F = ma. Their “more complete relation” is a material’s response function, not a deeper law of nature. They are first terms of something, but the something is a property of copper, not of the universe.
05 Where the series is all we have
So far the structure behind each series has been known. We have γmc²; we have the elliptic integral; we have general relativity sitting behind Newton. The series was a view of a thing we could also write down whole. Now the harder case, and the one where the brief is most exactly right.
Since Kenneth Wilson’s work on the renormalisation group in 1971, the working picture of fundamental physics has been that every theory we have is an effective theory — valid below some energy scale Λ (lambda, the cutoff), and written as a series in powers of E/Λ, energy over cutoff. Steven Weinberg’s 1979 statement of the method is worth quoting in paraphrase because it sounds like a recipe and is actually a metaphysics: write down every term the symmetries allow, infinitely many, ordered by how fast they shrink as the energy drops; keep as many as your experiment can resolve. The Standard Model of particle physics, on this view, is the leading terms of such a series. The catalogue of what comes next has been written out — in 2010 four physicists in Warsaw counted 59 independent next-order terms — and the experiments now underway at CERN are, in the most literal sense, measurements of Taylor coefficients.
Even Einstein’s equations of gravity have been put in this form. In 1994 John Donoghue wrote the gravitational action as Einstein’s term plus terms in the square of the curvature and showed how to compute what they do. Einstein’s equation, the most beautiful thing in physics, is on this reading the first term of a series whose later terms are suppressed — we assume — by the Planck scale.
And here the radius of convergence comes back with a new face. What sets the radius of an effective theory’s series? The nearest singularity — and in a quantum field theory the nearest singularity is the mass of the lightest particle you left out. Fermi’s theory of the weak force is a series that stops converging at the mass of the W boson, because the W is the pole at ±i: invisible from below, and the reason the low-energy series dies where it does. That is the picture from Figure 03, with a particle standing where the pole was. It is not an analogy. It is the same theorem.
Now the part that has to be said carefully, because two different series are in play and popular accounts run them together. The series in E/Λ converges below the cutoff. But there is a second series inside every quantum field theory — the expansion in the strength of the interaction, the coupling constant, which for electromagnetism is α ≈ 1/137 — and in 1952 Freeman Dyson gave an argument, which he himself called tentative, that this second series has a radius of convergence of zero. The argument is one paragraph long. If the series converged for small positive α, it would converge for small negative α too. But negative α means like charges attract, and in such a world the vacuum is not stable: it can lower its energy without limit by producing pairs. So the function has no Taylor series at α = 0 at all. The infinitely long formula, in this case, does not add up.
A philosopher’s first reaction is that this must make the theory unreliable. It does not, and the reason is a fact about divergent series that every student should meet once. A series can diverge and still be the most accurate arithmetic ever performed: the terms shrink for a while, reach a smallest term, and then grow — and if you stop at the smallest term, the error is about the size of that term. Quantum electrodynamics, expanded in a series that does not converge, predicts the electron’s magnetic moment to about twelve significant figures. Dyson’s own estimate was that the terms would start growing after roughly 137 of them, and physicists have computed five.
A series that does not converge and works anyway
Divergent does not mean wrong. It means the series is a ladder you can climb ten rungs of and then must step off, because the eleventh rung goes down.
06 What no series can see
One more boundary, and it is the one that stops the slogan cold. Cauchy noticed in 1823 that the function e−1/x², with the value 0 filled in at x = 0, is perfectly smooth — infinitely differentiable — and yet every one of its derivatives at zero is zero. Its Taylor series about the origin is 0 + 0x + 0x² + …, which converges beautifully, everywhere, to the wrong function. The series is not wrong about any coefficient. It is blind. The function is flat at the origin to every order and non-zero a hair’s breadth away, and no amount of differentiating at zero will ever detect it.
Now recall the shape of the tunnelling exponent from Part Ten. The chance of crossing a wall the classical world says is impassable goes as e−c/ħ for some positive constant c set by the wall. As a function of Planck’s constant that is Cauchy’s function, with ħ in place of x². Expand it about ħ = 0 — which is exactly what “classical physics plus quantum corrections” means — and every coefficient is zero. Tunnelling is invisible to every order of the series. It is not a small correction to the classical answer; it is a term the expansion about the classical point cannot express at all. Physicists call such effects non-perturbative, and the Sun, which Part Ten showed runs on tunnelling, runs on one.
So the honest map has three regions. There are structures we can write down whole, whose series are views of them and converge out to the nearest singularity. There are structures we cannot write down, where the series is the only thing we have and, expanded in the coupling, does not converge at all — and works anyway. And there are features of the world that no series about the easy point can see, flat to every order, present nonetheless. The great formulas are first terms. The rest of the world is not always a second term.
07 On “math is why”
The spine of this series is Tegmark’s: the world does not obey a mathematical structure, it is one, and the direction of explanation runs from the structure to the phenomenon. Every part so far has performed that inversion once. Here it is for this one.
The ordinary story says: physicists observed that laws are simple — linear, quadratic, inverse-square — and marvelled at the simplicity. Eugene Wigner’s 1960 essay on the unreasonable effectiveness of mathematics is the classic statement of the marvel. The inversion says: the simplicity of the first term is Taylor’s theorem. Any smooth structure whatsoever is linear to first order about any point. It could not fail to be. The linearity of Hooke’s law is not information about springs; it is information about what differentiable means. To the extent the great formulas are simple because they are first terms, their simplicity is explained by the mathematics of smoothness and needs no further miracle.
I have to be exact about how far that goes, because this audience will push, and should. Wigner’s actual puzzle was not that laws are locally simple. It was that concepts invented for their own beauty — complex numbers, Hilbert space — turn out to apply; and that a law fitted to crude data over a narrow range keeps holding when extrapolated far beyond it. Taylor’s theorem answers neither of those directly. It explains why the first term is linear near a point. It does not explain why the same first term, with the same constants, holds from a laboratory bench to the orbit of a galaxy. What explains that is a physical fact, not a theorem: the world is arranged so that the small parameter — v/c, GM/rc², E/Λ — is tiny across enormous ranges. Gravity at the Sun’s surface is a GM/rc² of two parts in a million. That separation of scales is why first terms travel so far, and calculus does not supply it. So the honest statement is this: Taylor’s theorem explains why laws look simple nearby. Something about the world, not about calculus, explains why “nearby” is so large. I take the first half to dissolve part of Wigner’s wonder and the second half to relocate the rest. What remains to wonder at is not why the tangent line is straight but why we live so far from every singularity.
There is one more consequence, and it is the deepest thing the series has to say about why first terms are trustworthy at all. A function that equals its Taylor series — an analytic function, which is what everything in section 01 through 04 was — has a property philosophers ought to find remarkable: knowing it perfectly at one point fixes it everywhere the series reaches. The infinitely many derivatives at v = 0 are the energy at v = 0.9c. That is why Newton’s gravity, fitted in the solar system, could be trusted at all in a galaxy: the structure is rigid, and rigidity is a mathematical property of the structure, not a kindness of the world. Where it fails — at the pole, at the W boson, at Cauchy’s flat function — it fails for a mathematical reason too.
Now the required hedge, stated once. Physics is compatible with and suggestive of reading the analytic structure as the reality and the series as its shadow. It does not prove it. A realist about objects can accept every word above and hold that the function summarises dispositions of things, and that the things come first. And there is a serious philosophical literature that reads the truncation the other way: Nancy Cartwright argued in 1983 that fundamental laws, read literally, are false of real situations and hold only ceteris paribus. I would put her point in this lesson’s vocabulary and then disagree with the moral. The laws do not lie. They truncate, and a truncation carries its own error term — Lagrange’s remainder is precisely a bound on the ceteris that are not paribus. And Robert Batterman’s 2002 argument that the interesting physics lives in singular limits, where the expansion breaks, is exactly right and is section 06 of this lesson: the flat function is his point in my notation.
Teach the mathematics as established. Teach the reading of it as an argument you find persuasive. The direction of explanation is the part I would defend to the end: the circle does not explain eit, and the simplicity of Hooke’s law does not explain Taylor’s theorem. It is the other way round.
08 The ledger
- What is load-bearing
- E = γmc² expands to mc² + ½mv² + …, and F = dp/dt with p = γmv gives γ³ma along the line of motion. Taylor’s theorem makes every smooth function linear to first order about any point, and quadratic about any minimum. The radius of convergence equals the distance to the nearest singularity in the complex plane. e−1/x² is smooth with an identically zero Taylor series. Every number in the figures was computed, not sketched; the captions say how.
- What is convention
- Calling mc² “the first term” and ½mv² “the second” is counting from zero; some books count the other way, and on their count Newton’s kinetic energy is the first term and the rest energy is a constant that does not move. Nothing hangs on it. The choice to expand about v = 0 rather than some other speed is likewise a choice — that is the whole point of section 03.
- Where the shorthand breaks
- “Every formula is a first term” is false for identities, symmetries and the conservation laws they entail: E² = (pc)² + (mc²)² is not the beginning of anything, it is the norm that generates the series. “The true formula is infinitely long” is false whenever the closed form is known — and is the wrong description even when it is not, because the series depends on the point and the structure does not. “An expansion point is a frame” is literal for v = 0 and analogy for θ₀ = 0. And “the series” is two different series in field theory: the one in E/Λ converges below the cutoff; the one in the coupling does not converge at all, and no theorem yet proves that for electrodynamics — Dyson’s argument is physics, not mathematics, though Barry Simon proved the analogous statement for the anharmonic oscillator in 1970.
- Where I would push back on myself
- Section 07 leans on Taylor’s theorem to deflate Wigner, and Taylor’s theorem is local. The thing that actually needs explaining — why the first term holds across twenty orders of magnitude of scale, with the same constants — is separation of scales, and that is a contingent fact about this universe that no theorem about smooth functions delivers. A critic can grant every equation in this lesson and say I have relocated the mystery rather than dissolved it. I think that is right, and I think the relocation is progress, because “why are we so far from the nearest singularity” is a sharper question than “why is mathematics effective.” But sharper is not answered.
09 Exercises
- Count the terms Using the expansion in section 01, find the speed at which the third term, (3/8)mv⁴/c², equals one percent of the second term, ½mv². Express it as a fraction of c, and then in kilometres per second. Is anything you have ever ridden in close?
- Move the frame Expand 1/(1 + x²) about x = 2 instead of x = 0 or x = 1. Without computing a single coefficient, state the radius of convergence and say why. Then say which of the following changed when you moved the point: the function, the series, the poles, the radius.
- Find the singularity with a stopwatch The pendulum series in section 02 has radius π in θ₀. Explain in one paragraph, with no equations, why a student who never swings the bob past ten degrees is nevertheless bound by what happens at 180. Then say whether the same argument applies to Fermi’s theory and the W boson.
- Climb the divergent ladder Compute the partial sums of Σ (−1)ⁿ n! xⁿ at x = 0.05 by hand or by machine, find the smallest term, and estimate where the best stopping point is. Compare with Figure 05 at x = 0.1. State the rule that relates the best stopping point to x, and then say what it predicts for α = 1/137.
- Argue the other side Defend the claim that this lesson gets the direction of explanation backwards: that the great formulas are exact statements about a world of objects and their powers, that Taylor expansions are bookkeeping we impose, and that “the structure is the invariant” confuses a description with what it describes. Use Cartwright. Make the case at full strength. Then say what it costs — in particular, what the object-realist has to say about why the radius of convergence of a real series is set by a pole at ±i that no object ever occupies.
Sources
- A. Einstein, “Ist die Trägheit eines Körpers von seinem Energieinhalt abhängig?” Annalen der Physik 18, 639 (1905). The L/V² paper; the small-v expansion is here, not in the June paper.
- A. Einstein, “Zur Elektrodynamik bewegter Körper,” Annalen der Physik 17, 891 (1905), §10 for W = μV²[(1−v²/V²)−1/2 − 1].
- A. Einstein, “Erklärung der Perihelbewegung des Merkur aus der allgemeinen Relativitätstheorie,” Sitzungsber. Preuss. Akad. Wiss. 1915, 831 (18 November 1915).
- B. Taylor, Methodus Incrementorum Directa et Inversa (London, 1715), Prop. VII, Thm. 3, Cor. 2. No remainder, no convergence.
- J.-L. Lagrange, Théorie des fonctions analytiques (Paris, 1797) — the remainder term.
- A.-L. Cauchy, Résumé des leçons données à l’École Royale Polytechnique sur le calcul infinitésimal (1823), 38th lesson — e−1/x².
- R. Hooke, Lectures de Potentia Restitutiva, or of Spring (London, 1678); the anagram appeared in A Description of Helioscopes (1676).
- E. Grüneisen, “Theorie des festen Zustandes einatomiger Elemente,” Annalen der Physik 39, 257 (1912) — anharmonicity and thermal expansion.
- H. Kamerlingh Onnes, Comm. Phys. Lab. Leiden 71 (1901) — the virial coefficients named; the series form is Thiesen (1885).
- H. Euler and B. Kockel, Naturwissenschaften 23, 246 (1935) — the quartic term; W. Heisenberg and H. Euler, Z. Phys. 98, 714 (1936) — the full one-loop Lagrangian.
- ATLAS Collaboration, “Observation of light-by-light scattering in ultraperipheral Pb+Pb collisions,” Phys. Rev. Lett. 123, 052001 (2019); evidence in Nature Physics 13, 852 (2017).
- K. G. Wilson, Phys. Rev. B 4, 3174 and 3184 (1971).
- S. Weinberg, “Phenomenological Lagrangians,” Physica A 96, 327 (1979).
- B. Grzadkowski, M. Iskrzyński, M. Misiak and J. Rosiek, “Dimension-six terms in the Standard Model Lagrangian,” JHEP 10 (2010) 085 — the 59 operators.
- J. F. Donoghue, “General relativity as an effective field theory: the leading quantum corrections,” Phys. Rev. D 50, 3874 (1994).
- F. J. Dyson, “Divergence of perturbation theory in quantum electrodynamics,” Phys. Rev. 85, 631 (1952).
- C. M. Bender and T. T. Wu, Phys. Rev. 184, 1231 (1969); B. Simon, Ann. Phys. 58, 76 (1970) — the anharmonic oscillator’s divergent series, found and then proved.
- E. P. Wigner, “The Unreasonable Effectiveness of Mathematics in the Natural Sciences,” Comm. Pure Appl. Math. 13, 1 (1960).
- N. Cartwright, How the Laws of Physics Lie (Oxford, 1983).
- R. W. Batterman, The Devil in the Details: Asymptotic Reasoning in Explanation, Reduction, and Emergence (Oxford, 2002).
- J. Worrall, “Structural Realism: The Best of Both Worlds?” Dialectica 43, 99 (1989) — epistemic; J. Ladyman, Stud. Hist. Phil. Sci. 29, 409 (1998) — ontic.
- M. Tegmark, “The Mathematical Universe,” Found. Phys. 38, 101 (2008).
2 thoughts on “Every Formula Is a First Term”