Intelligence Is Predictive Compression · Part IV of VII
The Vanishing Act
Intelligence, when it works, produces nothing to notice
Zero surprise is not a quiet event. It is the absence of an event. A prediction that succeeds produces no signal, because a signal is a difference and there is no difference.
This part follows that consequence through habituation, mismatch negativity, inattentional blindness and the predictive-processing literature — reported honestly, including its substantial critics — to a conclusion about why we experience the failures of intelligence far more vividly than its successes.
Contents
You are not currently thinking about the floor.
It is holding you up. It has been holding you up continuously, and at no point in the last hour did you allocate a single unit of conscious attention to the question of whether it would continue. You did not verify it. You did not check. Your expectation of the floor is so completely satisfied, so continuously, that the floor has no experiential presence at all.
Now consider what happens the instant it gives way.
The floor becomes the only thing in the universe. Every other cognitive process stops. Attention arrives with a violence that has no equivalent in ordinary experience, and it arrives late — after the failure, not before it.
This is not a fact about floors. It is the general structure of conscious experience, and it follows directly from the framework of the last three parts.
If intelligence is predictive compression, and surprise is S = ln(A/E), then a well-compressed domain is one where A keeps matching E and S keeps landing near zero. And zero surprise is not a quiet event. It is the absence of an event. There is nothing there to attend to. A prediction that succeeds produces no signal, because a signal is a difference and there is no difference.
A note on vocabulary before going further, because this chapter draws on a literature that uses the same word for a different quantity. When neuroscientists say surprise they usually mean surprisal, −log p, the negative log probability of the sensory data under a generative model. Part III established that S = ln(A/E) is not that quantity and this series will not pretend otherwise. What carries across between them is not the mathematics. It is the structural fact both encode: a satisfied expectation generates nothing. That observation is what the rest of this chapter rests on, and it survives regardless of which formalism you prefer.
Which means the better an entity gets at a region of reality, the more completely that region vanishes from its experience.
The economy of attention
Attention is scarce. This is not a metaphor about busy lifestyles; it is a hard constraint. Sensory systems deliver vastly more information per second than can be processed at the level where conscious deliberation happens, and the excess is not queued. It is discarded.
Given scarcity, the allocation question is unavoidable: what gets processed?
The answer, empirically, is that attention is recruited by prediction failure. Not by importance. Not by magnitude. By mismatch.
The simplest demonstration is the oldest form of learning we know. Present an organism with a repeated, inconsequential stimulus and its response declines — habituation, characterized systematically by Thompson and Spencer in 1966 and observable in essentially everything with a nervous system.1 Nothing about the stimulus changed. What changed is that it became predictable, and predictable stimuli lose their claim on the organism.
The neural version is repetition suppression: repeat a stimulus and the neural response to it shrinks. The interesting question is whether that shrinkage is fatigue or expectation, and the evidence points at least partly toward expectation — manipulate what a subject expects to be repeated, holding the actual repetition constant, and the suppression tracks the expectation.2 The brain is not tired of the stimulus. The brain already had it.
The electrophysiology gives it a signature. Play a sequence of identical tones and then a deviant one, and roughly 100–250 milliseconds later there is a negative deflection in the difference wave — the mismatch negativity. It appears without attention, without task relevance, without the subject noticing. The brain has registered a violated regularity before anyone decided to care.3
And the cost of this allocation policy is measurable. In the best-known demonstration in cognitive psychology, subjects counting basketball passes failed — 46 percent of them, across all conditions — to notice a person in a gorilla suit walking through the middle of the scene, stopping, and beating their chest.4 The gorilla was large, central, and visible. It was not attended to, because attention was elsewhere and the gorilla was not a violation of what the visual system had been asked to model.
The predictive processing account
There is a research program built entirely on this, and it deserves to be stated carefully — both because it is the most developed formal account of the phenomenon and because it is more contested than popular treatments admit.
Rao and Ballard’s 1999 model of visual cortex proposed that feedback connections from higher to lower cortical areas carry predictions of lower-level activity, while feedforward connections carry the residual errors between those predictions and what actually arrived.5 Trained on natural images, such a network develops receptive fields resembling those of real simple cells, and the error-carrying units reproduce extra-classical effects like end-stopping. The architecture is inverted from the textbook picture: what flows up is not the signal but the part of the signal that was not anticipated.
Karl Friston generalized this into the free-energy principle. In his own summary: “Adaptive agents must occupy a limited repertoire of states and therefore minimize the long-term average of surprise associated with sensory exchanges with the world.” And: “Surprise rests on predictions about sensations, which depend on an internal generative model of the world. Although surprise cannot be measured directly, a free-energy bound on surprise can be, suggesting that agents minimize free energy by changing their predictions (perception) or by changing the predicted sensory inputs (action).”6
That second clause is the elegant part. Two ways to reduce mismatch: update the model, or change the world until it matches the model. Perception and action as a single operation viewed from two sides.
Within this framework, attention has a formal definition. It is the precision assigned to prediction error — the estimated reliability of an error signal, implemented as gain on the units carrying it.7 Attending to something is not shining a light on it. It is deciding that mismatch from that source is trustworthy enough to act on. Errors you judge unreliable get down-weighted into irrelevance; errors you judge precise get amplified into consciousness.
Andy Clark’s framing, which pushed this into philosophy of mind, is that brains are essentially prediction machines, continuously attempting to match incoming sensory input against top-down expectation.8 Jakob Hohwy extended the same principle across perception, attention, action, delusion, and the self.9
Where honesty requires slowing down
Now the part that most treatments skip.
Predictive processing is a serious, productive, and genuinely contested framework. It is not established fact, and this series will not lean on it as though it were.
The most careful recent assessment of the neurophysiological evidence puts it plainly: these models “have become increasingly influential in cognitive neuroscience,” but “are also criticized for lacking the empirical support to justify their status.”10 The same review notes that many of the phenomena routinely cited as evidence for predictive processing “can also be accommodated within traditional models by invoking additional mechanisms” — repetition suppression, for instance, is explicable by plain neural adaptation without any predictive machinery. Studies testing the framework’s most distinctive claims have produced directly conflicting results.
The honest verdict, as of the most recent thorough review: real support for the general claim that top-down expectation shapes early sensory processing; weak and contested support for the specific architectural claims — segregated prediction and error populations, particular laminar and oscillatory signatures, hierarchical message-passing.
There are structural objections too. The “dark room problem” asks why, if organisms minimize surprise, they do not seek out the most predictable possible environment and remain there. The standard reply is that organisms carry evolutionarily-shaped priors entailing exploration and homeostasis, so a dark room is in fact highly surprising for a creature like us.11 That reply is coherent, and it also illustrates the deeper worry: a framework that can absorb any counterexample by positing whatever priors are needed to fit the observed behavior is difficult to falsify. Philosophers have documented this explicitly — the free-energy principle “has been called a postulate, an unfalsifiable principle, a natural law, and an imperative,” and “its epistemic status is unclear.”12 A sharper critique argues that the framework illegitimately slides from Markov blankets as a statistical modeling device to Markov blankets as the literal physical boundary of an organism.13 Even the attention-as-precision reduction has drawn direct dispute.
I am reporting all of this because the argument of this series does not depend on predictive processing being correct. The argument depends on something much weaker and much better established: attention is recruited disproportionately by prediction failure. That is supported by habituation, by mismatch negativity, by inattentional blindness, and by the plain phenomenology of the floor. Predictive processing is the most developed theory of why. If it is wrong in its architectural specifics, the observation survives.
One more distinction, because it is constantly blurred. Midbrain dopamine neurons encode a reward prediction error — the gap between received and expected value.14 That is a different signal, in a different system, in a different currency, from the perceptual prediction error of Rao and Ballard. Both are “actual minus expected,” and the shared mathematical form is genuinely interesting, but they are not the same thing and treating them as one is a popularization error. Actual arrives in more than one register.
What this does to conscious experience
Now put the pieces together, because the consequence is strange and I think underappreciated.
Attention goes where prediction is failing. Prediction fails where compression is poor. Therefore conscious experience is systematically biased toward the regions where an entity’s intelligence is worst.
Not where it is best. Where it is worst.
Every domain you have genuinely mastered has become invisible to you. The native speaker does not experience grammar; the grammar of a language you are learning is nothing but experience. The experienced driver does not experience the clutch; the learner experiences almost nothing else. The physicist does not experience the algebra; they experience only the part that will not come out. Mastery is precisely the process of a domain going quiet.
Which produces the observational effect at the center of this piece:
We experience the failures of intelligence far more vividly than its successes. The successes disappear into normality.
This is not a psychological quirk. It is forced by the structure. A success is S near zero, and S near zero is nothing happening. There is no way to build a system that both predicts well and continuously experiences its own predicting, because the experience is made of the residue and good prediction leaves none.
The implications compound in unpleasant directions.
Skill is systematically under-credited. The competent person, doing the thing they are best at, has the least to report about it. Ask an expert how they did it and you will usually get a bad answer — not because they are withholding, but because the process that produced the answer generated no prediction error and therefore left no experiential trace. Expertise is largely unavailable to introspection by construction.
Infrastructure is invisible until it fails. Water systems, power grids, supply chains, immune responses, institutional norms. Every one is a compression of enormous accumulated learning about what reality does, and every one is experienced as nothing right up until the moment it stops working — at which point it is experienced as a catastrophe and, frequently, as evidence that it was never any good.
Progress feels like decline. As predictive systems improve, the set of things that surprise us shrinks toward the residual — the hard cases, the anomalies, the failures. Attention, being drawn to mismatch, concentrates entirely there. The result is a population that experiences a steadily improving system as a steadily worsening one, because the improving part goes silent and the failing part gets all the light. This is not irrationality. It is attention working exactly as designed, applied to a moving baseline.
And this applies with unusual force to artificial intelligence. A model that produces ten thousand competent responses and one fabricated citation will be discussed almost exclusively in terms of the fabricated citation. The ten thousand generated no surprise, so they generated no experience, so they generated no discourse. The one failure generated all three.
I want to be careful here, because this is where the argument could be misused. The failures are real and they matter. A fabricated citation in a legal brief is a serious professional harm regardless of how it was produced, and the fact that attention is drawn to failures does not mean attention is wrong to be drawn to them. It is the correct allocation policy: mismatch is where information is, and where the work remains.
The claim is narrower and it is about calibration. When you form a judgment about the intelligence of any entity — an institution, a person, a machine, yourself — you are working from a sample that is structurally biased toward that entity’s worst-compressed domains. The successes are not merely underweighted in your assessment. They were never in your assessment, because they were never in your experience.
The paradox, stated cleanly
An entity that predicted its entire environment perfectly would have S = 0 everywhere, always. Nothing would ever mismatch. Nothing would ever recruit attention. There would be no prediction error anywhere for consciousness to be made of.
Whatever such an entity is, its inner life would not resemble ours — and by every account of consciousness that ties experience to prediction error, it might not have one at all.
That is the shape of the thing:
Intelligence, at the limit of its success, erases the evidence of itself.
We only ever see it working badly. When it works well, there is nothing to see.
Which sets up the last confusion this series has to clear away — the most common one, and the one that has done the most damage to how people evaluate both machines and each other.
If we only notice the failures, we will draw our conclusions from them. And the failures we notice most are not failures of intelligence at all.
They are failures of memory.
Sources
1. Richard F. Thompson and W. A. Spencer, “Habituation: A Model Phenomenon for the Study of Neuronal Substrates of Behavior,” Psychological Review 73, no. 1 (1966): 16–43. Updated criteria in Rankin et al., “Habituation Revisited,” Neurobiology of Learning and Memory 92, no. 2 (2009): 135–138.
2. Kalanit Grill-Spector, Richard Henson, and Alex Martin, “Repetition and the Brain: Neural Models of Stimulus-Specific Effects,” Trends in Cognitive Sciences 10, no. 1 (2006): 14–23; Christopher Summerfield et al., “Neural Repetition Suppression Reflects Fulfilled Perceptual Expectations,” Nature Neuroscience 11, no. 9 (2008): 1004–1006; Ana Todorovic and Floris P. de Lange, “Repetition Suppression and Expectation Suppression Are Dissociable in Time in Early Auditory Evoked Fields,” Journal of Neuroscience 32, no. 39 (2012): 13389–13395. The expectation-based interpretation remains contested — see note 10.
3. Marta I. Garrido, James M. Kilner, Klaas E. Stephan, and Karl J. Friston, “The Mismatch Negativity: A Review of Underlying Mechanisms,” Clinical Neurophysiology 120, no. 3 (2009): 453–463.
4. Daniel J. Simons and Christopher F. Chabris, “Gorillas in Our Midst: Sustained Inattentional Blindness for Dynamic Events,” Perception 28, no. 9 (1999): 1059–1074.
5. Rajesh P. N. Rao and Dana H. Ballard, “Predictive Coding in the Visual Cortex: A Functional Interpretation of Some Extra-Classical Receptive-Field Effects,” Nature Neuroscience 2, no. 1 (1999): 79–87.
6. Karl Friston, “The Free-Energy Principle: A Unified Brain Theory?” Nature Reviews Neuroscience 11, no. 2 (2010): 127–138.
7. Harriet Feldman and Karl Friston, “Attention, Uncertainty, and Free-Energy,” Frontiers in Human Neuroscience 4 (2010): 215.
8. Andy Clark, “Whatever Next? Predictive Brains, Situated Agents, and the Future of Cognitive Science,” Behavioral and Brain Sciences 36, no. 3 (2013): 181–204.
9. Jakob Hohwy, The Predictive Mind (Oxford: Oxford University Press, 2013).
10. Kevin S. Walsh, David P. McGovern, Andy Clark, and Redmond G. O’Connell, “Evaluating the Neurophysiological Evidence for Predictive Processing as a Model of Perception,” Annals of the New York Academy of Sciences 1464, no. 1 (2020): 242–268.
11. Karl Friston, Christopher Thornton, and Andy Clark, “Free-Energy Minimization and the Dark-Room Problem,” Frontiers in Psychology 3 (2012): 130.
12. Matteo Colombo and Cory Wright, “First Principles in the Life Sciences: The Free-Energy Principle, Organicism, and Mechanism,” Synthese 198, suppl. 14 (2021): 3463–3488.
13. Jelle Bruineberg, Krzysztof Dołęga, Joe Dewhurst, and Manuel Baltieri, “The Emperor’s New Markov Blankets,” Behavioral and Brain Sciences 45 (2022): e183.
14. Wolfram Schultz, Peter Dayan, and P. Read Montague, “A Neural Substrate of Prediction and Reward,” Science 275, no. 5306 (1997): 1593–1599.
1 thought on “The Vanishing Act”