Essay · Artificial Intelligence · Geometry
Every Trap Is Near the Floor
To be stuck, every direction around you has to point up. In three dimensions that is easy to arrange. In a million it almost never happens, except near the bottom, where being stuck hardly matters. That one fact of geometry is why the machines we train do not behave the way our intuition says they should.
The principle
(½)d
A flat spot is a trap only if the ground rises in all d directions at once. Treat each direction as a coin flip and the odds halve with every dimension you add.
A picture, not a law. Section 03 says where it bends, and the ledger says where it breaks.
Fog · a ridge trail, late afternoon
You are walking down a mountain in fog so thick you can see only your boots. You have one rule: step downhill. It works for an hour. Then every step you try goes up. Left, up. Right, up. Forward, up. Back, up.
You are standing in a hollow, and as far as your feet can tell, you are done. You are not at the bottom. You are just stuck.
01The trap needs every direction
That hollow is what mathematicians call a local minimum, and it is the oldest fear in optimization. Training a neural network is the same walk in the same fog. The model is a point on a landscape. Height is how wrong the model is. The only rule is step downhill, a method called gradient descent. For decades the obvious worry was that the walk would end in a hollow like yours, somewhere high up, with a far better valley hidden behind a ridge it cannot see.
Look at what that worry quietly assumes. You were trapped because every direction available to you went up. On a hillside you have only two directions to test, north-south and east-west. For a flat spot to trap you, both have to curve upward. That is not much to ask of a landscape.
Now add directions. The landscape a large model walks does not have two. It has one for every adjustable number in the model, and there are billions. For a flat spot to trap it, the ground has to rise in every one of them at once. If each direction were a coin flip, the odds would halve with every dimension added. Three dimensions: one in eight. A hundred: about one in a million trillion trillion. A million: one in a number more than three hundred thousand digits long.
Figure 01
How rare a trap gets, in the coin-flip picture
The mathematics of real random landscapes adds a twist the coin misses. High up, the directions are not independent coin flips. They push against each other, and the odds of a trap fall even faster than halving. What you meet up there is not a hollow. It is a saddle: a place that curves up in some directions and down in others. A mountain pass. Flat enough to slow you down. Not a place anyone stays.
02Where the idea comes from
Richard Bellman gave us the phrase “curse of dimensionality” in 1957, and he had good reason. Search a space by checking its corners, and the corners multiply beyond any computer. For most of the century that was the whole story about high dimensions: more is worse.
Then the same geometry started showing its other face. In 1997 the mathematician Paul Kainen argued that high dimension could make some computations easier, not harder. In 2000 David Donoho gave the idea its popular name in a lecture on the curses and blessings of dimensionality. In 2007 the physicists Alan Bray and David Dean worked out what flat spots look like on large random landscapes. And in 2014 Yann Dauphin, Yoshua Bengio and their colleagues carried that result into neural networks with a simple claim: the obstacle to training is not a proliferation of bad hollows. It is a proliferation of saddles.
03Every trap is near the floor
Here is the part most tellings skip, and it is the best part. Bray and Dean found that height decides. Far up the landscape, almost every flat spot has many ways down. The lower you go, the fewer downhill directions each flat spot has, until near the bottom most flat spots really are hollows.
So traps exist. They are just not scattered at random. They collect near the floor. And a hollow near the floor is nearly as good as the floor itself. Analysis of large networks suggests the hollows they settle into sit in a narrow band just above the best possible answer, and experiments find many of those hollows joined to one another by long, curving valleys that barely rise at all.
High up, you cannot get stuck. You can only slow down. By the time you can get stuck, you are nearly there.
That is the whole reversal in one sentence. The fog walker’s fear was being trapped high on the mountain. In a million dimensions the mountain does not offer that trap. It offers passes, and passes lead down.
04The second strangeness: room
The same geometry does something else, and this is the part that delights me most. Pick two directions at random in three dimensions and they point every which way. Pick two at random in a thousand dimensions and they are almost exactly perpendicular: ninety degrees, give or take about two.
Figure 02
Two random directions, and how often they land nearly perpendicular
Perpendicular means independent. Two directions at right angles do not interfere with each other. In three dimensions you can fit exactly three of those. In high dimensions, if you allow “nearly perpendicular,” the number you can fit grows exponentially with the dimension. A space with a few thousand coordinates has room for vastly more than a few thousand nearly independent directions.
That is how a model can hold more ideas than it has dimensions. Researchers at Anthropic showed in 2022, in small toy models, that networks do exactly this when ideas rarely show up together: each one gets its own nearly perpendicular direction, overlapping the others a little, and the model pays a small price in interference. In July I wrote that dimensions give meaning room. The strange part is how much room. Not a larger room. An exponentially larger one.
05The turn: our intuition runs backwards
Put the two strangenesses together and you get the thing that makes AI feel uncanny. Every instinct we have was formed in three dimensions. In three dimensions, a complicated system with more parts gets stuck more often, and a crowded room gets crowded. In a very high-dimensional space both instincts invert. Adding adjustable numbers adds ways out, so the bigger model is the easier one to train. Adding dimensions adds room faster than anything can fill it, so the model can hold more distinctions without them trampling one another.
One precision matters here. The dimension that gets you unstuck is the number of adjustable parts, not the size of the input. A small network fed an enormous input can still get trapped. Give it more knobs than it strictly needs and the traps thin out. The blessing belongs to the machine’s own freedom, not to the data’s.
Stuck is a three-dimensional word.
A classroom · a few years from now
The slide shows the diagram every textbook has used for fifty years: a ball rolling down a bumpy curve and settling in a dip that is not the lowest one. The teacher starts to explain the danger. A student asks how many directions the picture has. One, the teacher says. And the real one? Billions.
Then the picture is upside down, she says. The only dips are at the bottom. The teacher, who learned it the other way, stops and agrees.
06What I expect to see
“It will get stuck” will keep being the wrong forecast.
When a large model stalls in training, the cause will keep turning out to be the data, the objective or the schedule, not a trap in the landscape. Whatever limits AI in the years ahead will be data, energy and money, not geometry.
Bigger stays easier to train.
For the kinds of models we build, adding adjustable parts will keep making training more reliable, not less. The first sign this is ending will be large models that fail to train for reasons no one can trace to data or setup.
Meaning becomes an address.
Finding things by what they mean rather than what they are called, across contracts, photos, messages and records, will become the default everywhere. In these systems meaning is a direction, and nearby directions are cheap to find. I take this one further in Meaning Becomes an Address.
Looking inside a model becomes geometry.
The way we inspect and steer AI will be to find directions, the direction for a concept, a tone, a refusal, and move along them. Reading a model will look less like reading code and more like surveying a landscape.
High-dimensional geometry enters ordinary education.
Near-perpendicular directions, saddles instead of hollows, volume that lives at the edge: within a decade these will be taught to non-specialists the way probability was taught in the last century, because they are the geometry of the tools everyone uses.
07Why I care about this
Very few people share my enthusiasm for high dimensionality. When I bring it up, most people hear a technical detail. I hear the reason. It answers the question everyone asks about AI and almost no one answers: why does this work at all? Why doesn’t a search through billions of settings simply get lost?
This week I finally found a video that carries the same delight, Why High-Dimensional Space Is So Strange, from the channel math_is_fun. It is the reason this piece exists. If you want to feel the strangeness before you reason about it, start there. And if the vocabulary of dimensions and parameters is new, Dimensions Before Parameters sets it up.
08The ledger
- Already true
- Large networks trained by nothing more than stepping downhill reach low error routinely, and experiments find their resting places joined by low, curving valleys. The mathematics of random landscapes shows that high up, flat spots where every direction rises are astronomically rare. That random directions in high dimensions are nearly perpendicular is a theorem, not a hunch.
- What has to happen
- For the predictions to hold, the models we build have to stay generously oversized for their tasks, and real training landscapes have to keep behaving enough like the random-landscape picture that “high up means there is a way down” stays true in practice.
- Where I am probably wrong
- The clean mathematics is about random landscapes, and real networks are not random. Their curvature is flat in most directions, which the tidy picture ignores. Small networks do get trapped, and have been proven to. The theorems that make bigger networks easy to train assume sizes far beyond anything practical, and they speak only to training error, not to how well the model does on new data. Escaping a saddle is nearly certain but can take a very long time. If a future kind of model turns out to be hard to train because of genuine traps high on its landscape, this piece is wrong at its root.
09Back in the fog
You are standing where every step you have tried goes up. But you have tried four. The landscape the machine walks has billions more, and almost always one of them goes down. You only thought you were in a hollow because you could not count the directions.
It was never a hollow. It was a pass.
Background
- Richard Bellman, Dynamic Programming, Princeton University Press, 1957 (origin of “curse of dimensionality”).
- Paul C. Kainen, “Utilizing geometric anomalies of high dimension: when complexity makes computation easier,” 1997.
- David L. Donoho, “High-Dimensional Data Analysis: The Curses and Blessings of Dimensionality,” AMS lecture, 2000.
- Alan J. Bray and David S. Dean, “Statistics of critical points of Gaussian fields on large-dimensional spaces,” Physical Review Letters 98, 2007.
- Yann Dauphin et al., “Identifying and attacking the saddle point problem in high-dimensional non-convex optimization,” NIPS 2014.
- Anna Choromanska et al., “The Loss Surfaces of Multilayer Networks,” AISTATS 2015.
- Timur Garipov et al., “Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs,” NeurIPS 2018; Felix Draxler et al., “Essentially No Barriers in Neural Network Energy Landscape,” ICML 2018.
- Itay Safran and Ohad Shamir, “Spurious Local Minima are Common in Two-Layer ReLU Neural Networks,” ICML 2018.
- Nelson Elhage et al., “Toy Models of Superposition,” Anthropic, 2022.
- math_is_fun, “Why High-Dimensional Space Is So Strange,” YouTube.
- John Rector, “Dimensions Before Parameters,” July 2026.
- John Rector, “Meaning Becomes an Address,” September 2026.
More at johnrector.me.
1 thought on “Every Trap Is Near the Floor”