An advanced companion to the second edition
The Reality EquationMathematical Companion
The quotient is one line. Understanding what it preserves, what it loses, and what follows from it takes a book.
A chapter at a time
Open any chapter heading to read its derivation, then reveal the worked solution after trying the problem. Close a chapter to return to the compact outline. Five parts organize the twenty chapters.
How to use this companion
Every claim must carry its premises.
The Reality Equation distinguishes an event from the relation through which an entity receives it. Its second edition builds a program around complex expectation, logarithmic surprise, predictive geometry, and decomposition. This companion works through that program in mathematical detail.
Read Chapters 1–8 for the algebra and geometry, 9–12 for inference and dynamics, and 13–20 for the connections to probability, coherence, and intelligence. Algebra is assumed. Complex numbers, derivatives, entropy, and information bounds are developed where they are needed.
The second edition is the main text. The later foundation paper and mathematical essays sharpen its conventions. Their differences are identified in the source ledger below. This companion derives consequences, supplies missing assumptions, and offers clearly labeled refinements where a formula cannot support its intended claim.
A proof establishes what follows from a model. Measurement establishes whether a model earns its interpretation. Both are necessary. Neither can substitute for the other.
Chapter 01Before the divisionThe observable is part of the equation
Begin with a record and a question. A package arrived in six days. That sentence supplies many possible observables: elapsed time, damage, delivery cost, completion of a promise. An equation about elapsed time is not an equation about care. Before calculating, declare the entity X, the observable O, the encounter time t, and a positive reference scale x₀.
Write the recorded magnitude as a, the predictive magnitude as p, and their dimensionless counterparts as A = a/x₀ and P = p/x₀. The book calls these normalized quantities à and P̃. Here the tildes are suppressed after normalization. B denotes the dimensionless ideational magnitude. The ordinary domain is A > 0, P > 0, B ≥ 0.
These are definitions within the model. That experience is represented by this quotient is a modeling claim. That reported valence tracks s is a further empirical claim. Neither is established by dividing two numbers correctly. The arithmetic can be exact while an interpretation remains untested.
The projection A = ΠX,O,t(world)/x₀ expresses observable-relativity. One world can supply different projections. If two witnesses track different observables, canceling their numerators is invalid. If a record is incomplete, uncertainty belongs to our estimate of A; it does not imply the historical occurrence was itself revised.
Positivity also carries a cost. A quantity with a meaningful zero and positive ratios fits directly. Signed profit, a rating scale with an arbitrary origin, or Celsius temperature does not enter unchanged. Adding a convenient constant to make every value positive changes the ratios. A transformation must be declared and justified before the outcome is known.
A longer delivery time gives a larger time-ratio, yet may be worse for the recipient. Thus positive s is not automatically good. The assignment of valence depends on what the observable measures and how it is oriented. A reciprocal speed-like observable reverses a time-ratio; it creates a different declared coordinate, not a moral sign hidden in arithmetic.
This companion develops the second edition’s mathematical program using the later essays’ distinctions. It uses definition for a chosen quantity, proposition for a consequence with stated premises, and postulate for a bridge requiring evidence. When a definition needs refinement, the refinement is identified.
Chapter 02Perform the divisionKeep the complex quotient intact
Complex division is ordinary algebra with one additional rule: i² = −1. Multiply the numerator and denominator by the conjugate P − iB. The cross terms cancel in the denominator, leaving a positive real number.
Define D = P² + B² and r = |E| = √D. If θ = atan2(B,P), then E = r eiθ. Division has two separate effects: divide length by r and reverse the angle.
With A = 5, P = 3, B = 4, the denominator is 3 + 4i and has length 5. Its quotient is 0.6 − 0.8i. The modulus is 1, but the quotient is not 1. This is the first distinction a student must be able to maintain without hesitation.
Proposition. Full neutrality R = 1 holds exactly when P = A and B = 0. Magnitude neutrality |R| = 1 holds when P² + B² = A². Proof: R = 1 means E = A, a positive real number. Magnitude neutrality merely equates the two lengths. One condition is a point in the expectation plane; the other is a circular arc in the ordinary domain.
In the book’s nonnegative-B convention, E lies in the first quadrant and R in the fourth, including the positive real boundary. The angle φ lies in (−π/2, 0]. A signed-coordinate extension may allow B ∈ ℝ, giving −π/2 < φ < π/2. That extension is useful, but it must not quietly replace a magnitude with a signed coordinate.
The distinction between Re R and |R| matters just as much. Re R = A P/D, whereas |R| = A/√D. These coincide only when B = 0 in the ordinary domain. Taking the real part before taking a logarithm would produce a different scalar: ln(Re R) = s + ln(cos θ). It is not the scalar the book defines.
Chapter 03What the imaginary magnitude retainsA resultant is a compression
To construct ideation, represent conditioned idea-directions as phasors zj = ajeiθⱼ, with aj ≥ 0. The aggregate is Z = Σzj. The book’s magnitude convention sets B = |Z|. Expanding the squared modulus gives the exact contribution of every pair:
Equal directions reinforce. Opposite directions subtract. Orthogonal directions contribute no cross term. For two phasors of lengths 3 and 4, B is 7 when aligned, 1 when opposed, and 5 when perpendicular. “How many ideas?” cannot answer “How large is the resultant?”
The triangle inequality gives B ≤ W, where W = Σaj. For W > 0, the ratio B/W measures net alignment of these particular phasors. It does not measure how many distinctions the host retains. Two equally weighted opposite phasors produce B = 0 despite both remaining represented.
There are two angles here. The direction ψ = arg Z tells us which direction the phasor sum points within idea-space. The expectation angle θ = atan2(B,P) tells us the relative contribution of ideation and prediction. They are not interchangeable. If the denominator receives only |Z|, it discards ψ. Rotating every idea-phasor by the same angle changes ψ while preserving B, E, and R.
That loss of information is a fact about the chosen map. Recovering the identity of the orienting idea requires retaining Z, or its phase, elsewhere in the entity’s state. The basic quotient cannot tell us whether two equally large resultants point toward different ideals.
An alternative signed model could set B = Re(e−iψ₀Z) along a declared ideational axis ψ₀. It would preserve one projection and discard the perpendicular one. Substituting the entire complex Z into E = P + iZ is a third model: its real part becomes P − Im Z. That generally breaks the identification of Re E with prediction. A mathematical companion must say which compression is being used before assigning meaning to its output.
Chapter 04Normalize without erasing the eventInvariance is a precise permission
The early hyperbola y = 1/x survives inside the complex model as a normalized special case. If A > 0, divide the entire denominator by A. Define Ê = E/A. Then R = 1/Ê. Actual has become 1 by a change of scale; its former value survives in the normalized denominator.
Proposition. The quotient and its logarithm are invariant under a common positive rescaling. The proof is cancellation of c. This is the exact content of “units cancel.” Rescaling A and P while leaving B fixed does not preserve the quotient; all components must participate in the same coordinate change.
In a dimensional formulation, one can first specify an ideational equivalent b in the same units as a and p, then use B = b/x₀. Alternatively, B can be calibrated directly as dimensionless relative to a declared x₀. Either method needs a calibration rule. Labeling an orientation dimensionless does not make its numerical scale uniquely determined.
A common additive shift is different. Five divided by eight is 0.625. Adding ten to each gives fifteen divided by eighteen, about 0.833333. Consequently the ratio is appropriate to a ratio scale with a meaningful origin, not automatically to every scale on which a questionnaire prints numbers.
Normalization by each event’s A is valid for that event. It does not license forgetting that the scale varies across events. Derivatives and comparisons must either carry the changing scale or be computed from the original ratios. The statement “A is always 1” is a coordinate convention, not a proof that all arrivals have identical content.
The older scalar picture also has a simple geometric limit. On B = 0 with A = 1, R = 1/P. The curve never reaches either axis at finite positive P. Its point nearest the origin is (1,1), because P² + P−2 ≥ 2, with equality only at P = 1. This proves the distance √2. It does not, by itself, prove that a physical standing wave occupies that segment. Geometry and physical realization require separate premises.
Chapter 05Why the logarithm appearsA uniqueness theorem, with its premises
Suppose a real readout f assigns a number to each positive ratio. Require it to add when ratios multiply: f(xy) = f(x) + f(y). Also require continuity at 1. These assumptions are sufficient to determine its form up to one constant.
Proposition. Under those assumptions, f(x) = c ln x for every x > 0. To prove it, define g(u) = f(eu). Then g(u + v) = g(u) + g(v). Taking u = v = 0 gives g(0) = 0; taking v = −u gives g(−u) = −g(u). Repeated addition gives g(nu) = ng(u) for integers n.
For a positive integer n, g(1) = ng(1/n), so g(m/n) = (m/n)g(1). Thus g agrees with a straight line on every rational number. For any real u, choose rationals qn → u. Continuity extends the equality to g(u) = ug(1). Return through x = eu: f(x) = g(1)ln x. That is the proof.
The normalization f(e) = 1 fixes c = 1. An increasing, nontrivial readout requires c > 0. Without this normalization, mathematics permits any real c, including the zero function. Continuity can be weakened, but some regularity is essential; unrestricted additive functions need not be straight lines.
Reciprocal symmetry and neutrality are consequences of additivity. They need not be three separate miracles. The theorem tells us what a continuous additive readout must be. It does not prove that a body uses an additive readout, that an idea has a measured numerical magnitude, or that experience is a ratio in the first place.
The unit conversion is equally exact. A real log-ratio in bits is s₂ = s/ln 2. A doubling yields 1 bit or ln 2 nats; a halving yields their negatives. If the entire complex logarithm is divided by ln 2, its imaginary coordinate is rescaled too. To retain φ in radians while reporting s in bits, report the ordered pair (s₂, φ) rather than pretending it is the unmodified complex logarithm.
Natural logarithms are especially convenient for calculus because the derivative of ln x is 1/x. That convenience fixes a unit, not a psychological law. The later logarithm essay correctly makes the empirical premise visible; the proof here supplies the steps between that premise and the result.
Chapter 06The complex logarithmBranches, winding, and the limits of addition
For a nonzero complex number z = reiθ, every number ln r + i(θ + 2πk), k ∈ ℤ, exponentiates to z. The complex logarithm is therefore multivalued until a branch is chosen. The principal value selects −π < arg z ≤ π; as an analytic function it excludes the negative real axis and zero.
Because A and P are positive, R stays in the open right half-plane. The ordinary model never meets the principal branch cut. This restricted domain makes its logarithm particularly well behaved. A full circuit around zero belongs to an enlarged model, not to an ordinary trajectory with P > 0.
For arbitrary nonzero complex factors, the principal logarithm obeys Log(z₁z₂) = Log z₁ + Log z₂ only up to an integer multiple of 2πi. For example, let z₁ = z₂ = e3πi/4. Their individual principal angles sum to 3π/2; their product has principal angle −π/2. The two logarithmic expressions differ by 2πi.
The real parts still add exactly: ln|z₁z₂| = ln|z₁| + ln|z₂|. For angles accumulated along a continuous path, use an unwrapped phase. If a closed path winds once around zero, its unwrapped logarithm gains 2πi although its endpoint complex number returns to its starting value. The line integral ∮dz/z records that winding.
There is another subtlety. Chapter 5’s real uniqueness theorem does not by itself force the coefficient of the imaginary angle. A continuous additive readout on radius and an unwrapped phase could weight ln r and θ independently. Choosing the analytic complex logarithm couples them. Analytic continuation of the real logarithm on a connected branch domain fixes the complex function. That analyticity is an additional mathematical structure we have adopted.
Nor does the symbol i prove an experiential channel is unfelt. “The body reads only s” is a proposed observation map, h(S) = Re S. It may be tested against alternatives h(s,φ). Mechanics does not establish that map for psychological experience merely because it also uses complex notation. The quotient and its derivatives survive either outcome.
Branch conventions follow NIST DLMF §4.2. The phenomenological interpretation is the Reality Equation’s separate modeling claim.
Chapter 07The geometry of inversionCircles, rays, and a neutral arc
Hold A fixed and regard E as the variable. The map E ↦ A/E is a complex reciprocal followed by scaling. Along a ray E = reiθ, it reverses the angle and replaces r with A/r. Along a circle centered at zero, it preserves circularity and replaces the radius with its reciprocal times A.
The differential is dR/dE = −A/E². It is nonzero wherever E ≠ 0, so the map is locally conformal: it preserves angles between intersecting smooth curves. Although the position angle changes sign, the full complex map is holomorphic, not an orientation-reversing reflection. Radial inversion and angle reversal act together. This is an instance where a correct drawing can invite an incorrect verbal conclusion.
Now hold P = p > 0 and vary B. Let R = u + iv. Since E = A/R, the real part of E is A u/(u² + v²). Setting it equal to p and completing the square gives:
The vertical prediction line maps to a circle through the origin, with center A/(2p) on the real axis. Under B ≥ 0, the trajectory traces its lower semicircle from A/p toward zero, never reaching zero for finite B. The approach to zero is a limit, not an allowed state with A > 0 and finite E.
In the expectation plane, constant real surprise s = c gives a circle of radius A e−c. Constant angular surprise φ gives a ray at angle −φ. Thus a magnitude reading supplies a circle; an angle supplies a ray; together, with A known, they locate E. This geometry is the inverse problem before any statistical noise is added.
Neutrality gives a particularly useful demonstration. P = A, B = 0 produces S = 0. Moving along the circle P² + B² = A² preserves s = 0 while changing φ. Moving along the line P = A while increasing B makes s negative. The phrase “add ideas without changing prediction” describes the line, not the circle. Confusing those paths produces apparently contradictory claims that are simply different experiments.
Chapter 08Sensitivity and local errorWhat changes, by how much, under which constraint
Let D = P² + B². Differentiate the two coordinates directly. The resulting equations are exact differentials, not approximations:
The first separates fractional arrival change from radial expectation change. The second measures rotation. A perturbation parallel to (P,B) changes s but not φ. A perpendicular perturbation leaves s unchanged to first order while changing φ.
At fixed A and P, ∂s/∂B = −B/D. For B ≥ 0 the signed reading decreases as ideational magnitude grows. That does not imply the absolute reading decreases. With A = P = 5, B = 0 gives s = 0; B = 5 gives s = −½ln 2 ≈ −0.346574. The signed value fell and |s| grew. “Dampens the modulus of Reality” is true; “cannot intensify a negative deviation” is false if intensity means |s|.
When P and B are positive, the logarithmic sensitivities are −P²/D and −B²/D. They sum to −1. A simultaneous one-percent increase in both components therefore reduces s by approximately 0.01; exactly, scaling both by 1.01 subtracts ln 1.01.
For a local expansion near full neutrality, write A/P = 1 + ε and b = B/P. Then:
Small ideation affects angle at first order and real surprise at second order around B = 0. A scalar sensor can therefore be insensitive to a small angular change. This is a mathematical reason to distinguish channels even before interpreting them.
Uncertainty propagates through the same derivatives. For an estimate vector (Â,P̂,B̂) with small error covariance Σ, let g = (1/A, −P/D, −B/D). The delta-method approximation is Var(ŝ) ≈ gᵀΣg. Correlated estimation errors matter; adding three independent error bars without their covariance can misstate the uncertainty. Close to A = 0 or E = 0 the derivatives become large and a linear approximation can fail.
A partial derivative is not an instruction to the witness. “Holding P fixed while changing B” specifies a comparison of modeled states. It supplies no mechanism for consciously setting either component.
Chapter 09One reading, several unknownsIdentifiability before interpretation
A reading is not a diagnosis. In log coordinates u = ln A and v = ln|E|, the real observation is s = u − v. Its Jacobian is the one-row matrix [1, −1]. Its null direction is (1,1): moving both coordinates equally changes neither the reading nor anything inferable from that reading alone.
If s = ln(5/8), then (A,|E|) can be (5,8), (10,16), (8,12.8), or any positive common rescaling. The observation defines a line in log coordinates and a ray in ordinary coordinates. More accurate measurement of the same s narrows uncertainty around that line; it does not select a point on it.
Knowing A adds an independent constraint and recovers |E| = Ae−s. But P and B remain on an arc P² + B² = A²e−2s. A written point prediction estimates P; it does not automatically measure the full denominator. If A, s, and P are known exactly, B² follows, but a signed B retains a sign ambiguity. A negative inferred B² is a model-or-measurement inconsistency, not an imaginary ideational magnitude to be ignored.
Proposition. The full complex reading plus a known positive A uniquely determines E within the declared branch. Proof: exponentiate −S and multiply by A. Without known A, even full complex surprise retains a common-scale ambiguity. An extra channel does not abolish every unknown.
With independent Gaussian observation noise of variance τ² on s, the Fisher information for (u,v) from one reading is τ−2[[1,−1],[−1,1]]. Its determinant is zero. Repeating that same measurement improves precision in the difference direction but never identifies the shared-scale direction. A prior can choose a preferred point; it cannot retroactively turn that choice into information supplied by the observation.
For two entities sharing exactly the same A, s₁ − s₂ = ln(|E₂|/|E₁|). If their Actual projections differ, an additional ln(A₁/A₂) remains. Repeated outcomes with an unchanged denominator give the mirror identity. Both experiments work only to the extent their held-fixed assumptions hold.
Even an identified parameter is not automatically a cause. Establishing that P changed does not establish why it changed. That requires a temporal model, relevant observations, and an intervention or other defensible causal design.
Chapter 10The equation does not contain a learning lawDynamics must be added openly
The static quotient tells us the relation among an arrival and a standing expectation. It does not tell us how tomorrow’s prediction is formed. To study learning, add a state-update rule. The choice is substantive and testable.
A useful elementary model works in log coordinates. Set ut = ln At and zt = ln Pt. With B = 0, let a prediction update after the arrival according to:
For a constant arrival A, define the log error et = ln A − zt. Substitution gives et+1 = (1 − η)et, hence et = (1 − η)te₀. The prediction converges when 0 < η < 2. It converges monotonically when 0 < η ≤ 1 and alternates around the target when 1 < η < 2. At η = 2 the error alternates without decaying; larger gains are unstable.
This is a theorem about the specified recursion. It does not establish a biological learning rate. For η = 0.2, the error retains 80 percent of its preceding value per encounter. Its continuous-valued half-life is ln(1/2)/ln(0.8), about 3.106 encounters.
Noise produces a trade-off. Suppose ut = μ + εt, with independent zero-mean errors of variance v. In stationarity, the variance V of z satisfies V = (1 − η)²V + η²v. Therefore V = ηv/(2 − η). Larger gains track a changing target faster but transmit more observational variation into the prediction. This conclusion depends on the independence and stationarity assumptions.
Now restore a fixed B > 0. If this learning rule makes P → A, the limiting real surprise is −½ln(1 + (B/A)²), not zero. Learning the predictive magnitude alone does not produce full neutrality. Adding a rule for B is a separate modeling step.
The later foundation paper places agency after the received quotient: actions generate artifacts and subsequent encounters can reshape prediction. A dynamical representation can therefore use a hidden state z, an action policy u = π(R,z), and a future state update driven by arriving records. It must not write a conscious command directly into E merely because the analyst can vary E on paper.
Chapter 11Accumulation is an operatorA signed balance and an attention load are different
The attention essay defines a normalized accumulation of complex surprise. In discrete time, choose nonnegative weights wj summing to one and form U = ΣwjSj. In continuous time, choose a nonnegative kernel K with integral one over the attention window and integrate K S. The kernel sets a temporal weighting; its normalization prevents arbitrary growth merely from increasing the sample rate.
The last inequality is the triangle inequality. It separates net accumulated surprise from gross surprise load. Two equally weighted real episodes +1 and −1 give U = 0 and L = 1. Calling both quantities “attention” would erase precisely the distinction the model needs to explain.
The published demonstration maps U to U/(1 + |U|), optionally after a threshold. That map preserves direction, has magnitude below one, and compresses large inputs. Those are proven properties of the map. That the nervous system applies it is an open empirical proposal. Nor does the formal complex direction identify a semantic object: an additional addressing rule is needed to name what attention is about.
Thresholding is nonlinear and order matters. Let Tδ(x) keep a real x only when |x| ≥ δ. Two opposite sub-events can be detected individually and then cancel in a sum. Conversely, accumulating many small same-sign unnormalized increments before applying a threshold can produce a detected total even when no increment was detected. A normalized mean behaves differently again. Specify whether the mechanism sums, averages, integrates evidence, or forgets.
A related exact identity underlies the conservation essays. For positive intermediate expectations E₀,…,Eₙ and a final A:
The intermediate terms telescope. But arbitrary episode surprises Σln(Ak/Ek) do not telescope unless the terms actually form such a chain. Discounting, thresholds, saturation, and taking absolute values also destroy the simple endpoint identity.
For a decline from 100 to 60, the total log change is ln 0.6 ≈ −0.510826. In ten equal multiplicative steps, each change is about −0.051083. A toy hard threshold at 0.06 misses each increment and detects the whole change. This is an illustrative detector, not a measured threshold for human expectation. The mathematics distinguishes the telescoping ledger from the instrument reading it.
Chapter 12Many entities, one aggregateGeometric means and correlated shocks
A collective score requires an aggregation rule. Let fixed weights αj ≥ 0 sum to one. For positive quotient moduli rj, averaging real log readings yields the logarithm of their weighted geometric mean:
This is not generally the log of the arithmetic mean. Concavity of ln gives Σαjln rj ≤ ln(Σαjrj). For equally weighted moduli 2 and 1/2, the average log is zero, while the average modulus is 1.25. “The average Reality is neutral” depends on which average was declared.
If all entities share A, their aggregate real surprise equals ln A minus the weighted mean of ln|Ej|. The effective denominator for that statistical aggregate is a geometric mean. It is not automatically the denominator of a family, firm, or other composite entity. A composite needs its own observable, predictive mechanism, and ideational state. Averaging its members is one possible model, not a consequence of being a composite.
The population essay on correlated surprise motivates a second calculation. If s is a random vector with covariance matrix Σ, then Var(s̄) = αᵀΣα. With n equal-weight readings, each of variance v and each pair sharing correlation ρ, this becomes:
For independent readings, the variance is v/n. For perfectly correlated readings, it is v. When ρ is positive and fixed, increasing the population leaves a variance floor ρv. The formula is exact for the stipulated equal-correlation model; admissible ρ must satisfy −1/(n−1) ≤ ρ ≤ 1.
Shared denominators are not the only source of correlated surprise. A common arrival can correlate readings even when denominator variations are independent. In log coordinates sj = u − vj; its cross-covariance includes Var(u), Cov(vj,vk), and their cross terms. Identifying a common information diet as the cause therefore requires measurements beyond correlated outcomes.
One can define an effective independent population neff = n/[1 + (n−1)ρ] in the positive-correlation case. It measures variance reduction relative to independent equally variable readings. It does not count autonomous minds or quantify their moral worth.
Chapter 13Four quantities that share a nameSigned surprise, surprisal, divergence, and error
Let q assign probabilities to discrete outcomes. The surprisal of an observed outcome x is ℓ(x) = −ln q(x). It is nonnegative because q(x) ≤ 1. The Reality Equation’s real magnitude reading s = ln(A/|E|) can be negative. They coincide under a particular embedding: define A = 1, P = q(x), B = 0. Then s = −ln q(x).
This embedding is a declared event-indicator comparison. It does not require the mistaken claim that observing an outcome retroactively changes its pre-outcome probability to one. The realized indicator is one; the forecast probability remains the historical forecast. For continuous variables, probabilities of exact points are generally zero and densities depend on units. Use declared finite bins or a likelihood ratio with a shared reference measure.
Let the true discrete distribution be μ, with q positive wherever μ is positive. Expanding the logarithm gives the cross-entropy identity:
Nonnegativity of the divergence follows from Jensen’s inequality applied to −ln(q/μ), or the log-sum inequality. Equality requires q = μ on the relevant support. Improving a model can reduce excess expected log loss, but cannot reduce the source entropy by forecasting better. This theorem concerns probability loss, not automatically the signed magnitude quotient.
Bayesian surprise is another object: a divergence between posterior and prior beliefs about a hidden parameter. Prediction error is another: a residual such as A − P. A log-ratio approximates relative residual error near agreement, ln(A/P) ≈ (A − P)/P, but approximation is not identity. A Gaussian residual in log space yields squared log-ratio loss after a specified noise model; it does not make an arbitrary raw residual equal to a logarithm.
A revealing calculation uses a positive lognormal arrival. Suppose ln A has mean μ and variance v, and use the arithmetic-mean predictor P = E[A] = exp(μ + v/2), with B = 0. Then E[s] = μ − ln P = −v/2. A correct arithmetic mean does not produce mean-zero signed log surprise. The geometric predictor P = exp(E[ln A]) does. With deterministic B added, the expected real reading becomes μ − ½ln(P² + B²).
This is Jensen’s gap, not evidence of pessimism. “Calibrated” must identify a target: mean, median, quantile, probability distribution, or expected log outcome. Different loss functions legitimately select different forecasts.
For the information-theoretic quantities, see Shannon’s original paper. The signed-ratio comparisons and examples above are derived here.
Chapter 14The cloud is not the guessA scalar quotient cannot recover a distribution
The second edition separates a predictive distribution p(x) from the scalar it sends to the denominator, P = G[p]. The resolving functional G may be a mean, mode, quantile, or sample. Specifying it is indispensable.
Two clouds can have the same mean and radically different shapes. Consider probability one-half at each of 1 and 3, versus a point mass at 2. Both have mean 2; the first has variance 1 and the second variance 0. With the same A and B, they produce identical R and S. The quotient cannot identify spread, coherence, or modality from the resolved mean.
Now consider the book’s counterfactual response S(a) = Log(a/E) with a positive and E fixed. Its real derivative with respect to a is 1/a and its imaginary part is constant. It is strictly increasing on every continuous positive interval. It has no interior local minima, no multiple basins, and no curvature encoding the modes of p. Restricting it to disconnected support does not make its formula encode probability weights.
This matters for Chapter 16 of the second edition: the displayed fixed-E log-ratio cannot alone be the multibasin landscape described in its prose. A companion should expose the missing structure instead of repeating the conclusion.
A probability surprise landscape Lp(a) = −ln p(a) is a different object. For a smooth positive density, its minima occur at density maxima. A mixture with separated modes can create separated basins. Its absolute height depends on coordinate units; basin locations survive a common linear rescaling after transforming the density correctly. Finite-bin probabilities make the measurement protocol explicit.
There are therefore two legitimate experiments. Vary a while holding E fixed to measure the quotient’s response. Separately estimate p and examine its log-loss landscape to study predictive geometry. If one wants a counterfactual E(a), that function must be specified; it is additional machinery.
A further point concerns latent uncertainty. Variation in reported guesses can reflect sampling from a fixed cloud, changing clouds across contexts, observation noise, or a deterministic reporting policy. Those mechanisms cannot be separated by calling all variation “spread.” Repeated controlled contexts and a model of the reporting channel are needed.
The conceptual distinction survives this repair: the machine can carry rich structure even when its realized quotient is near neutrality. What changes is the proposed instrument for measuring that structure. We need the cloud or a valid proxy for it; the resolved scalar does not contain it.
Chapter 15Coherence needs a defensible measureBoundedness is something to prove
The second edition proposes an entropy-based coherence κ = 1 − h(p)/hmax(v), where h is differential entropy and hmax(v) = ½ln(2πev). The Gaussian maximizes differential entropy at fixed positive variance. That upper bound does not supply a lower bound on h(p).
Even at variance one, h(p) can become arbitrarily negative: concentrate probability into very narrow separated peaks while preserving their overall variance. Consequently the proposed ratio can exceed one even after variance normalization. If hmax is zero it is undefined; if negative, its inequality behavior changes. Selecting convenient units cannot provide a universal [0,1] guarantee.
A mathematically safe refinement starts with a density p having finite entropy and finite nonzero variance. Let g be the Gaussian with the same mean and variance. Then:
To prove the equality, expand −ln g(x); it is a constant plus a quadratic. Its expectation under p equals its expectation under g because their first two moments match. Therefore −∫p ln g = h(g), and the identity follows. Under a common invertible affine change of variable, the entropy shifts cancel and J remains unchanged.
This κJ is a proposed bounded refinement, not a retroactive quotation from the book. It measures non-Gaussianity. Non-Gaussianity is not automatically useful organization: skewness, heavy tails, or narrow arbitrary peaks can raise it. A behavioral interpretation still needs independent validation. For a discrete or singular cloud, use a declared discrete reference distribution or a smoothing protocol; the density derivation cannot simply be reused.
The spectral candidate is also precise when its inputs are declared. With nonnegative weights summing to a positive value, κ₁ = |Σwjeiφⱼ|/Σwj lies between zero and one by the triangle inequality. But it measures alignment, not every form of phase organization. Two perfectly stable equal-weight phases at 0 and π give κ₁ = 0. Their second harmonic κ₂ = |Σwje2iφⱼ|/Σwj equals 1.
Thus a zero first-order parameter can coexist with perfect anti-phase organization. Nor are entropy deficit and phase alignment generally equivalent: a static amplitude distribution and a collection of dynamic phases are different data. Calling both “coherence” does not create a theorem connecting them.
Report which κ was measured, its weighting, reference family, sampling resolution, and uncertainty. A product score is only as meaningful as those ingredients. The bridge from any such structural score to felt experience remains the correspondence postulate.
Chapter 16The peak has conditionsDerive the capacity curve completely
The core structural score is C = σκ: spread times a chosen coherence measure. The bare product does not imply a peak. If κ remains a positive constant, C rises linearly with σ. An inverted-U needs a specified decline of coherence.
Adopt the second edition’s family, with σc > 0, m > 0, and raw normalized dispersion σ ≥ 0:
For m > 1 the numerator of C′ decreases from 1 to negative values and crosses zero once. Thus the unique maximum occurs at σ* = σc(m−1)−1/m. Evaluating C there gives:
For m = 2, σ* = σc and Cmax = σc/2. For m = 4, σ* ≈ 0.759836σc and Cmax ≈ 0.569877σc. The peak equals the named capacity only in the m = 2 special case. “At the edge” is an interpretation of this parameterized relation, not a universal numerical identity.
For m = 1 the score rises toward σc without an interior maximum. For 0 < m < 1 it grows without bound like σcmσ1−m. Finite named capacity alone does not force superlinear decay; the assumption m > 1 does.
The second edition also allows squashing spread to σ/(1+σ). That changes the score’s functional form and generally moves its maximum. The derivative above applies to raw normalized dispersion. If dispersion is restricted to a bounded admissible interval, the observed maximum can be at its boundary instead of at σ*. A theorem must retain its domain.
One can fit σc and m to independently measured coherence versus dispersion, then test the predicted structural peak on held-out observations. Defining κ by the desired curve and displaying a matching C peak is a demonstration of the definition, not evidence that the curve describes a real system.
Finally, none of this calculus identifies a conscious system. The product is a structural index. The book’s correspondence postulate relates it to experiential richness where experience exists; it does not turn C > 0 into a proof of sentience.
Chapter 17When a component pays for itselfCompression, interactions, and honest accounting
Let y be an arrival signal, V a declared vocabulary, and M a model composed of candidate components. A linear decomposition writes y = Σcj + rres. The remainder is system-relative: it is what this decomposition leaves unresolved. The identity alone does not tell us which components deserve credit.
The book’s five tests supply that discipline: distinguishability, stability under perturbation, internal coherence, net compression gain, and validity on held-out signals. These tests concern different failure modes. A named component may be stable but redundant, compressive but fragile, or attractive in training data but useless elsewhere.
For precise accounting, fix a code family and define the total description length L(M,y) = L(M) + L(y|M). Charge a component’s description inside L(M) or separately, exactly once. Its gain must compare complete ledgers.
The final identity telescopes when the models form one nested chain. This is a safe way to make additive component credit. In contrast, summing “remove each component from the full model” gains can double-count interacting structure. Those comparisons do not share canceling middle terms.
A small parity example exposes the issue. Suppose a target bit is the XOR of two independent fair input bits. Either input alone tells us nothing about the target; the pair determines it. Over 100 examples, imagine a baseline code of 100 bits, a one-component total code of 101 bits, and a joint total code of 2 bits because the rule plus residual is exceptionally short. These are stipulated illustrative code costs.
Along one nested path the gains are −1 and 99, summing to the true joint gain 98. Removing either component from the full model gives 99, whose sum 198 exceeds the total. Conversely, rejecting every component with negative individual forward gain would reject a useful pair. The admissible unit may need to be a block. A group-aware selection rule is an extension, not a result of naive counting.
Out-of-sample testing must preserve the ledger. Fit the vocabulary and model-selection choices on training data, then assess a locked code on unseen signals. If held-out data repeatedly guide selection, they become training data and another test set is needed. Report a distribution of gains, not merely whether one finite sample happened to be positive.
Compression gain is measured in bits when the code uses log₂. It is not automatically conserved attention, utility, or truth. A model can efficiently describe a biased record. Testing fidelity to the intended observable is a distinct task.
The code-length approach follows the minimum-description-length tradition; see Grünwald’s tutorial. The nested-gain ledger here makes the component accounting explicit.
Chapter 18What finite capacity actually boundsA finite message and a lifetime of compression
The second edition sketches Iact ≤ Ipot, with potential capacity defined by a finite number of distinguishable states. To make a theorem, specify what is registered, how long the encounter lasts, and which information quantity is bounded.
Let Y be an arrival and Z a representation taking at most M distinct values. For discrete variables:
The first inequality follows because conditional entropy is nonnegative. The second follows because the uniform distribution maximizes entropy on M outcomes. This proves a capacity bound on the information about Y registered in one such representation. It also covers a randomized encoder, provided the output alphabet remains the declared M states.
It does not prove that every compression score over an arbitrarily long signal is bounded by log₂ M. A small stored rule can compress repetitions indefinitely. A device that recognizes a long all-zero block can describe its length and the rule in roughly logarithmic space while saving almost one bit per symbol relative to a naive baseline. Those growing savings are not newly registered distinctions in a single fixed-size state.
Likewise a one-bit register reused across T rounds can transmit up to T bits through its sequence of states. Its instantaneous capacity is one bit; its transcript has up to 2T possibilities. Memory capacity, throughput, model knowledge, and compression benefit are different quantities with different units or time dependence.
Thus the original broad bound needs an explicit bridge from the chosen decomposition gain to registered information under a fixed channel and encounter window. This chapter supplies the valid finite-alphabet bound without pretending it resolves every scoring interpretation.
The angular special case is straightforward. If a circle is partitioned into M distinguishable cells, capacity is log₂ M bits. Equal one-degree cells give M = 360 and about 8.491853 bits. Two independent such coordinates give 360² joint states and about 16.983706 bits. Correlated or constrained coordinates have fewer accessible joint states.
Radians and degrees produce the same M when both circumference and resolution use the same unit. The ratio 2π/δθ was already dimensionless; the logarithm’s benefit is additive accounting for independent joint configurations, not repairing inconsistent angular units. When the ratio is not an integer, declare an actual partition or a packing criterion instead of counting fractional distinguishable states.
A capacity number does not say how quickly useful states are found or whether their use serves an appropriate goal. Search cost, action coupling, and outcome quality remain separate axes.
Chapter 19The remainder and the recordFidelity needs its own experiment
The decomposition remainder and the surprise reading meet through a model, not through their names. A large residual energy can be expected noise; a small residual can be highly improbable under a narrow forecast. To infer surprise from a remainder, specify its probability model or an explicit map back to the positive observable.
Under log loss, improving q toward μ lowers the divergence term in expected surprisal. It need not drive magnitude surprise to zero. Consider a positive arrival that is 1 or 3 with equal probability. A perfect distributional model knows both possibilities and their probabilities. Its mean is 2. With B = 0, the realized signed readings are ln(1/2) and ln(3/2), neither zero. Perfect probabilistic calibration leaves irreducible realized variation.
This supplies a counterexample to any unconditional bridge claiming that perfect model fit or complete distributional knowledge makes |ln(A/P)| vanish. Such a result would require additional restrictions, for example a deterministic arrival class and a matching predictor. The book marks the remainder-to-ratio bridge as open; the companion preserves that status.
Trace fidelity introduces another independent uncertainty. Let a latent event project to A and a record deliver an estimate Â. If the denominator is fixed, the error in the real reading is:
A record overstating A by 10 percent shifts s upward by ln 1.1 ≈ 0.095310. A decomposer may process that record flawlessly and remain wrong about the world it represents. Rechecking the same corrupted record does not create independent evidence.
The book’s factorization Ieff = IactΦnum is a proposed index. It is not a theorem without an operational definition of fidelity and evidence for a multiplicative response. A scalar fidelity factor can hide which features were corrupted. A one-percent transcription error in a decision-critical field may matter more than a large error in irrelevant formatting.
A stronger evaluation therefore reports a profile: decomposition gain under the declared code; record error by relevant feature; out-of-sample performance; cost of obtaining independent ground truth; and the sensitivity of downstream decisions to those errors. A combined scalar is permissible once its weights and purpose are published.
Estimation should also distinguish observational uncertainty from the ontology of the Past. Present artifacts may conflict, disappear, or be corrected. That is a fact about the evidence channel. A cryptographic hash establishes the consistency of bytes with a commitment, not the truth of the event those bytes describe.
The mathematical discipline is consistent across both sides of the equation: identify what the instrument observes, identify what was inferred, and do not let a precise calculation conceal an unmeasured input.
Chapter 20A complete encounterCompute, identify, and leave the open question open
Bring the pieces together in a fully specified mathematical example. The numbers are constructed, not measured psychological data. Let A = 5, P = 6, B = √28, with the same entity, observable, time convention, and reference scale throughout. We have D = 36 + 28 = 64 and |E| = 8.
The quotient is approximately 0.468750 − 0.413399i. Its real part is not its modulus. Its negative real surprise does not, without an observation model, diagnose a feeling or assign blame to an arrival.
Suppose the only available observation is s. The pair (A,|E|) remains unidentified. Even if a reliable record fixes A = 5, the split between P and B remains unidentified. If a second independent observation supplies φ, the inverse formula recovers the full denominator. If φ is not observable, it must not be invented from the story.
Next compare a modeled state with the same A and P but twice B. The new squared denominator is 36 + 4·28 = 148. Its real surprise is ln(5/√148) ≈ −0.889168, and the signed change is −½ln(148/64) ≈ −0.419165. This is a state comparison. It says nothing about whether the person can produce that change at will.
Alternatively scale all three components by ten. Every quotient coordinate remains unchanged. Alternatively keep |E| = 8 but set B = 0 and P = 8. The real reading remains −0.470004 and the phase becomes zero. These are three different operations: ideational growth at fixed prediction, a units change, and a rotation at fixed length. They must never be described as one operation called “changing expectation.”
For the predictive cloud, infinitely many distributions can resolve to P = 6. Its width and organization require additional observations. For decomposition intelligence, the same arrival may admit many signal representations; code family, component vocabulary, and held-out tests must be fixed. For attention, one encounter does not determine the weighting kernel, threshold, or semantic address.
A serious test of the framework would record forecasts before outcomes; declare a positive ratio-scale observable; compare the log-ratio model against plausible alternatives; specify any angular measurement independently; and evaluate predictions on held-out encounters. A test of coherence would measure spread and organization independently rather than choosing values to match reported experience. A test of decomposition would retain model costs and null-signal controls.
Mathematics makes these obligations visible. The companion’s central result is not a new slogan. It is a chain of computations whose inputs, invariances, missing measurements, and boundary conditions can be inspected. The quotient is compact. The discipline needed to use it is substantial.
Appendix A
One notation, throughout
- X, O
- Entity and declared observable.
- A
- Positive Actual magnitude after normalization.
- P
- Positive resolved prediction after normalization.
- B
- Nonnegative ideational resultant magnitude in the ordinary convention.
- E
- Complex Expectation, P + iB.
- R
- Complex Reality, A/E.
- S
- Principal complex log reading, Log R.
- s, φ
- Real log reading and angular coordinate: S = s + iφ.
- p, q, μ
- Predictive cloud, probability model, and reference arrival distribution; their roles are stated locally.
- σ, κ
- Dispersion and a declared coherence measure.
- U
- Weighted accumulation of complex surprise.
- C
- Structural coherent-spread score, not a certification of experience.
- L(M,y)
- Total code cost of a model and signal.
- r_res
- System-relative decomposition remainder.
- I(Y;Z)
- Mutual information, distinct from the decomposition-intelligence notation in the book.
Appendix B
What this companion clarifies
- A = 1 is normalization. It does not make all observables or events identical. All denominator coordinates must transform with the reference scale.
- The complex and scalar readings have separate names. S denotes Log R; s denotes its real part. Zero s need not mean zero S.
- The idea-resultant angle and the expectation angle differ. Taking B = |Z| discards the former. A later theory of ideation must retain or reconstruct it explicitly.
- Magnitude damping is not a theorem about absolute distress. At fixed A and P, larger B decreases s; when s is negative, its absolute magnitude can increase. A psychological “jolt” needs a specified observation map.
- The logarithm’s uniqueness is conditional. Ratio composition, additivity, regularity, and a unit select the real logarithm. The analytic complex extension adds structure; global principal-branch additivity requires care.
- The equation is not a controller. Partial derivatives compare states. Learning and action need a separate causal model consistent with the later foundation paper’s received denominator.
- The cloud needs an independent instrument. A fixed-E log-ratio is monotone in positive A. A multibasin probability landscape belongs to −ln p, a different function.
- The entropy ratio needs refinement. Variance normalization alone does not bound 1 − h/h_max. Chapter 15 offers a bounded non-Gaussianity score and states what it does not establish.
- A coherence peak requires a decay law. The specific inverted-U follows for m > 1 in the specified family, not from a finite-capacity label alone.
- Component gains need a common ledger. Conditional nested gains telescope; arbitrary deletion gains do not. Interactions can require grouped components.
- The capacity theorem has a precise target. Registered information in an M-state representation is bounded by log₂ M. Lifetime compression gain is a different quantity.
- The correspondence postulate remains a postulate. No calculation here proves that a system has felt experience or that a structural score is consciousness itself.
Appendix C · Source map
The book and the working archive
The source review combined all 694 published entries in the Reality Equation category with paginated site searches for Reality Equation, Immutable Past, and Math Lab. That produced 1,353 distinct related posts as of October 2, 2026. Their body text was indexed and screened; 243 mathematics or foundation titles were flagged for closer review. The full second edition and 24-page foundation paper were read alongside the central mathematical essays. The search inventory is broader than a count of articles devoted solely to the equation.
This is a mathematical synthesis, not an assertion that every historical formulation is mutually consistent. The original sixteen-chapter serial, second edition, August foundation paper, and September essays record an evolving argument. The links below identify the principal sources used to resolve its notation and extend its derivations.
- The Reality Equation — Second Edition · 2026-07-27
- The Reality Equation – Chapter 1 · 2026-04-10
- The Reality Equation – Chapter 3 · 2026-04-10
- The Reality Equation – Chapter 4 · 2026-04-10
- The Reality Equation – Chapter 6 · 2026-04-10
- The Reality Equation – Chapter 10 · 2026-04-10
- The Reality Equation – Chapter 11 · 2026-04-10
- The Reality Equation – Chapter 15 · 2026-04-10
- The Immutable Past · 2026-08-05
- Complex Reality (2D) · 2025-08-22
- Complex Reality (2D) — Advanced Notes (α & γ, v2) · 2025-08-22
- SPEC Formalism: Entity, Reduced-State Specification, and the Reality Quotient · 2025-09-15
- Attention Is Normalized Accumulated Surprise · 2026-06-28
- The Shape of the Prediction Machine · 2026-07-01
- Intelligence as Decomposition · 2026-07-03
- What Is a Logarithm · 2026-08-05
- John Rector’s Math Lab · 2026-08-27
- Surprise Has an Imaginary Part · 2026-09-03
- The Conservation of Surprise · 2026-09-04
- Surprise Has a Noise Floor · 2026-09-14
- One Reading, Two Unknowns · 2026-09-20
- Correlated Surprise · 2026-09-24
- The Reality Equation, Second Edition — complete book · 132 PDF pages. Principal crosswalk: Chapters 3–12 for the quotient; 13–20 for the cloud; 21–27 for decomposition; 29 and 34–38 for attention, boundaries, and open problems.
- The Immutable Past — foundation paper · 24 PDF pages. Consult especially Sections 2, 7, and 12 for records, the received quotient, and open questions.
- NIST Digital Library of Mathematical Functions, §4.2 · complex logarithms and branch conventions.
- Claude E. Shannon, A Mathematical Theory of Communication · the information-theoretic foundation.
- Peter Grünwald, A Tutorial Introduction to the Minimum Description Length Principle · explicit coding and model costs.
The worked numerical cases and curves in this companion are constructed calculations. They are not measurements of people, organizations, or artificial systems. The proposed refinements are identified in Chapters 14, 15, 17, and 18; their interpretations remain open to criticism and testing.