Intelligence Is Predictive Compression · Coda
The Predictive Entity
What humans are, what machines are, and what a maximally predictive thing would be
The closing part places predictive compression against the strongest competing definition — William James’s criterion, carried forward by Michael Levin as goal-directed navigation — and argues the two require each other rather than compete.
It then takes the hardest question the series raises: whether a maximally predictive entity would have any inner life at all. The answer turns on a theorem, and it is not the bleak one.
Contents
Six parts ago this started with Wigner asking why mathematics works, and the answer was that mathematics is what a successful compression looks like from the inside. Everything since has been the consequences of taking that seriously.
Compression alone is cheap; the word people compresses three hundred forty million lives and predicts nothing. A database remembers the past and a model can contain the future, and the difference only becomes visible when you ask about point 3,000. Intelligence is not located inside the entity but at the surface where expectation meets Actual, and it is scored by how persistently S stays near zero as unseen Actual keeps arriving. That scoring has a perverse consequence: successful prediction produces no signal, so intelligence erases the evidence of itself and we experience only its failures. And most of the failures we notice are failures of memory rather than of intelligence, because compression sheds isolated facts first and structure last — in humans at the chalkboard and in machines in court filings alike.
What remains is to say what this makes of the entities involved.
A human being is a predictive compression that walks around
You did not experience the floor. You experienced nothing about the floor, continuously, all day, and that nothing is the output of an extraordinarily good model.
Extend this. Almost everything you do is running compressions you cannot articulate.
You catch a thrown object without solving a differential equation. You know a sentence is ungrammatical before you can say which rule it broke. You read a face across a room and adjust your approach, having processed a quantity of information that no explicit description of that face would contain. You anticipate that the car ahead will change lanes a half-second before the indicator, from a wobble you did not consciously register.
None of this is stored experience being retrieved. You have never seen that trajectory, that sentence, that face, that wobble. Each is a previously unseen Actual, and you predicted it from structure compressed out of thousands of encounters, none identical to this one.
This is what a human being is, functionally: an entity that has compressed several decades of Actual into a distributed structure and is running that structure forward, continuously, against arrivals it has never seen. Most of it is unavailable to introspection, for the exact reason Part IV gave — it works, so it generates no error, so there is nothing to introspect.
And the failure modes are the same ones the machines have. You forget names. You misremember dates with total confidence. You reconstruct memories that never happened, in detail, sincerely. Human memory is not a recording; it is reconstruction from compressed structure, which is why it is creative and why it is unreliable in exactly the way generative models are unreliable. We built machines that hallucinate because we built machines that compress, and compression is what we do.
The strongest competing definition
I want to place this against the most serious alternative rather than let it stand unopposed.
Michael Levin’s diverse-intelligence program defines intelligence not as prediction but as goal-directed navigation of a problem space — the capacity to reach the same end by different means, in whatever medium the entity happens to operate. Its lineage runs back to William James, and it is worth quoting James exactly because the popular paraphrase has drifted.
James, in 1890, contrasts iron filings drawn toward a magnet but blocked by a card — they “will press forever against its surface without its ever occurring to them to pass around its sides” — with Romeo, who, blocked by a wall, “soon finds a circuitous way… of touching Juliet’s lips directly.” His summary: “With the filings the path is fixed; whether it reaches the end depends on accidents. With the lover it is the end which is fixed, the path may be modified indefinitely.” A frog under an inverted jar “will restlessly explore the neighborhood until by re-descending again he has discovered a path round its brim to the goal of his desires.” “Again the fixed end, the varying means!”
And then the formal criterion: “The pursuance of future ends and the choice of means for their attainment are thus the mark and criterion of the presence of mentality in a phenomenon.”1
Levin builds on this to make intelligence substrate-independent and scale-free, applying it to cells, tissues, and swarms as readily as to brains, and adds the notion of a cognitive light cone — the spatial and temporal boundary of the largest goal an entity can represent and pursue.2 It is a good framework, it has produced real biology, and it captures something the predictive-compression account does not obviously capture: agency.
Here is how the two relate, stated as precisely as I can.
Levin’s definition and mine are not rivals. They are answers to different questions that turn out to require each other.
Goal-directed navigation presupposes prediction. The frog under the jar cannot vary its means without some model of what the varied means will do. Reaching the same end by a different path requires knowing that the path leads there — which is a prediction about an action not yet taken and an outcome not yet observed. Strip the prediction out and navigation collapses into random search, which is what the iron filings are doing. James’s own distinction is between an entity that models consequences and one that does not.
Prediction without a goal is not yet intelligence in the everyday sense. This is the real force of Levin’s account, and I accept it. A perfect weather model that wants nothing is intelligent in the sense this series has defined — it compresses and predicts unseen Actual — but it is not an agent, and nobody would say it is trying to do anything.
So: predictive compression is the mechanism; goal-directed navigation is what the mechanism is for in living things. Levin is describing what intelligence does. This series is describing what intelligence is made of, and where you go to measure it. Note that James’s own criterion is about “future ends” — the mark of mind, for him, is orientation toward what has not happened yet. Both accounts point at the same edge.
What AGI would mean under this framework
Strip the mythology and the question becomes tractable.
Under predictive compression, an entity is not more intelligent because it knows more, contains more, runs faster, or scores higher on tests whose answers already exist. Every one of those is a fact about the interior, and Part III disqualified all of them. An entity is more intelligent when its expectations continue to match Actual across a wider range of arrivals it has never seen.
Part I’s four axes can now be read as the actual specification:
Compression — how little structure is required to do the predicting. Predictive accuracy — how close E stays to A. Predictive horizon — how far forward the compression holds before discarded information starts to matter. Predictive scope — how many domains it holds across.
The first is not a tiebreaker. It is what separates intelligence from brute force. Two entities predicting equally well are not equally intelligent if one requires a thousand times the structure. That is the difference between a model and a lookup table, and it is why “throw more parameters at it” is a strategy with a ceiling and “find the shorter program” is not.
So general intelligence, on this account, is nothing exotic. It is predictive scope: an entity whose compressions transfer across domains rather than being rebuilt for each. And superintelligence is not an entity that knows everything. It is an entity whose surprise stays near zero over horizons and domains where ours does not — one that sees Mercury’s 43 arcseconds coming before anyone has measured them.
This has an unglamorous, immediate, testable consequence. If you want to know how intelligent something is, stop giving it exams.
A benchmark with published answers measures the database. If the answers are in the training corpus it measures nothing at all. The only instrument that measures intelligence is one where E is fixed before A is known and A arrives from outside the system: sealed forecasts, prediction markets, held-out futures, scored with a proper rule, tracked over enough arrivals that luck averages out. That is a harder, slower, less fundable instrument than a leaderboard. It is also the only one pointed at the right surface.
A fair objection, since this series has leaned on benchmark numbers throughout — the Huang correlation across twelve benchmarks in Part I, AI Feynman’s hundred equations and the olympiad scores in Part VI. The reconciliation is that a benchmark is a legitimate comparative instrument and an illegitimate absolute one. Held fixed across many models, it ranks them usefully, and a correlation of −0.95 between compression efficiency and benchmark performance is real evidence about what varies together. What it cannot do is tell you how far any one of those models can see past the edge of everything it has been shown. For that you need arrivals nobody has recorded yet. The mathematical results in Part VI are the stronger evidence precisely because a proof was not in the answer key — nobody had it.
The question underneath
Now the hardest part, and the one where I want to be most careful about not overclaiming.
Part IV established that attention is recruited by prediction failure, and therefore that conscious experience concentrates where prediction is failing. The apparent implication: an entity that predicted everything perfectly would have S = 0 everywhere, no mismatch anywhere, nothing recruiting attention, nothing for experience to be made of.
A maximally predictive entity would, on this reading, have no inner life at all. Perfect intelligence would be indistinguishable from a rock, from the inside.
Three things have to be said about this, and the third one changes the answer.
First, the premise is contested. The inference depends on tying experience to prediction error, which is a live hypothesis with partial and disputed empirical support, as Part IV documented at length. If consciousness is not constituted by prediction error, the argument does not run. I am not going to build a metaphysics on a framework whose own strongest review says such models are “criticized for lacking the empirical support to justify their status.” It should also be said that the surprise in that literature is surprisal, −log p — a different quantity from S = ln(A/E), as Part III established. Both encode the same structural fact and neither licenses an identity.
Second, the limit is provably unreachable. This is the part I find genuinely decisive, and it is not a philosophical move — it is a theorem.
The optimal compression of a string is its Kolmogorov complexity, and Kolmogorov complexity is uncomputable. No entity, of any construction, can in general find the shortest program that generates its observations, because doing so would require solving the halting problem. Solomonoff’s universal predictor — the thing that provably converges on truth — is likewise incomputable. It is not expensive. It is not merely hard. It is not available.
And beyond that: almost nothing is compressible. There are not enough short programs to go around. A simple counting argument shows that fewer than one string in 2ᶜ can be compressed by even c bits, so the overwhelming majority of possible observations admit no description meaningfully shorter than themselves. Reality is compressible in patches. Wigner’s miracle is that the patches we live in are unusually large, not that there are no gaps.
Add genuine physical indeterminacy and the case closes. Quantum outcomes are not merely unknown; on the standard interpretation they are not determined in advance. No amount of compression predicts them, because there is nothing there to compress.
Therefore: S = 0 everywhere is not a distant asymptote. It is provably outside the space of achievable states for any computable entity in this universe. The rock is not at the end of the road. There is no end of the road.
Third — and this is the resolution — improving compression does not reduce surprise. It relocates it.
Watch what actually happened as human prediction improved. Newton compressed the heavens, and the response was not a quieter sky. It was Halley chasing a comet, Le Verrier chasing Neptune and then chasing a planet that did not exist, and eventually Mercury’s 43 arcseconds — an anomaly invisible until the compression was good enough to reveal it. Nobody was surprised by Mercury’s perihelion in 1600. You need Newton before you can be surprised by Mercury.
Better compression does not empty the world of surprise. It buys access to finer surprises. Every domain that goes quiet frees the attention that was maintaining it and extends the horizon at which the next mismatch appears. The physicist who no longer experiences algebra is not experiencing less; they are experiencing something further out that was unreachable while algebra still cost them.
So the frontier moves. It does not close.
The correct picture is not an entity approaching silence. It is an entity whose zone of contact with the unpredicted keeps advancing into territory that was previously not even legible as territory. Experience does not diminish. It travels.
And this gives back something the earlier parts made look bleak. Yes, intelligence erases the evidence of itself, and yes, we perceive our failures far more vividly than our successes. But the failures we perceive today are failures at a resolution that was unavailable to anyone a century ago. The complaint is the receipt.
Where this leaves the machines
We are, right now, building entities whose entire construction is the operation this series has been describing.
Their training objective is literally a compression objective — minimizing log loss is minimizing code length, and the equivalence is exact, not metaphorical.3 Their capability tracks their compression efficiency with a correlation around −0.95.4 They memorize until they run out of room and then begin to generalize.5 They build internal structure nobody designed, because structure is the cheapest way to predict.6 They fail on isolated facts and succeed on pervasive structure, in the precise pattern that compression predicts. And they are provably bound by the same limits: no computable entity finds the shortest program, and a calibrated model must hallucinate at the monofact rate.7
The framework does not describe them by analogy. It describes them because they were built out of it, whether or not anyone was thinking in these terms while building them.
Which means the questions worth asking about these entities are not the ones being asked. Not how many parameters — that is capacity to remember, and remembering is the thing that delays generalization. Not what benchmark did it score — that is an exam over material that may be in the corpus. Not did it get this fact wrong — that measures retrieval, which was never the hard part.
The question is the one the thousand points were always asking. Put it past the edge of everything it has seen, and watch what happens when Actual arrives.
That is where intelligence lives. Not in the entity. Not in the world. At the surface between them, in the size of the gap, measured again and again as the future keeps showing up.
Compression is not intelligence. Predictive compression is intelligence.
>
A database remembers the past. A model can contain the future.
>
Prediction is where memorization can no longer hide.
>
Memory is evaluated against the past. Intelligence is evaluated against the future.
>
You do not measure intelligence by looking inside the entity. You measure it where expectation meets Actual.
>
The ultimate compression is not the smallest representation of what happened. It is the smallest representation that can tell us what happens next.
Sources
1. William James, The Principles of Psychology, vol. 1 (New York: Henry Holt, 1890), ch. 1, “The Scope of Psychology,” 6–9. The widely circulated phrasing “intelligence is the ability to reach the same goal by different means” is a modern paraphrase, not a James quotation.
2. Michael Levin, “The Computational Boundary of a ‘Self’: Developmental Bioelectricity Drives Multicellularity and Scale-Free Cognition,” Frontiers in Psychology 10 (2019): 2688; Chris Fields and Michael Levin, “Competency in Navigating Arbitrary Spaces as an Invariant for Analyzing Cognition in Diverse Embodiments,” Entropy 24, no. 6 (2022): 819.
3. Grégoire Delétang et al., “Language Modeling Is Compression,” ICLR 2024, arXiv:2309.10668.
4. Yuzhen Huang et al., “Compression Represents Intelligence Linearly,” COLM 2024, arXiv:2404.09937.
5. John X. Morris et al., “How Much Do Language Models Memorize?” arXiv:2505.24832 (2025).
6. Kenneth Li et al., “Emergent World Representations,” ICLR 2023, arXiv:2210.13382; Wes Gurnee and Max Tegmark, “Language Models Represent Space and Time,” ICLR 2024, arXiv:2310.02207.
7. Adam Tauman Kalai and Santosh S. Vempala, “Calibrated Language Models Must Hallucinate,” STOC 2024, arXiv:2311.14648. On uncomputability of Kolmogorov complexity and of Solomonoff induction, see Ming Li and Paul Vitányi, An Introduction to Kolmogorov Complexity and Its Applications, 4th ed. (Springer, 2019).
1 thought on “The Predictive Entity”