Every Trap Is Near the Floor

Essay · Artificial Intelligence · Geometry

Every Trap Is Near the Floor

To be stuck, every direction around you has to point up. In three dimensions that is easy to arrange. In a million it almost never happens, except near the bottom, where being stuck hardly matters. That one fact of geometry is why the machines we train do not behave the way our intuition says they should.

The principle

(½)d

A flat spot is a trap only if the ground rises in all d directions at once. Treat each direction as a coin flip and the odds halve with every dimension you add.

A picture, not a law. Section 03 says where it bends, and the ledger says where it breaks.

Fog · a ridge trail, late afternoon

You are walking down a mountain in fog so thick you can see only your boots. You have one rule: step downhill. It works for an hour. Then every step you try goes up. Left, up. Right, up. Forward, up. Back, up.

You are standing in a hollow, and as far as your feet can tell, you are done. You are not at the bottom. You are just stuck.

01The trap needs every direction

That hollow is what mathematicians call a local minimum, and it is the oldest fear in optimization. Training a neural network is the same walk in the same fog. The model is a point on a landscape. Height is how wrong the model is. The only rule is step downhill, a method called gradient descent. For decades the obvious worry was that the walk would end in a hollow like yours, somewhere high up, with a far better valley hidden behind a ridge it cannot see.

Look at what that worry quietly assumes. You were trapped because every direction available to you went up. On a hillside you have only two directions to test, north-south and east-west. For a flat spot to trap you, both have to curve upward. That is not much to ask of a landscape.

Now add directions. The landscape a large model walks does not have two. It has one for every adjustable number in the model, and there are billions. For a flat spot to trap it, the ground has to rise in every one of them at once. If each direction were a coin flip, the odds would halve with every dimension added. Three dimensions: one in eight. A hundred: about one in a million trillion trillion. A million: one in a number more than three hundred thousand digits long.

Figure 01

How rare a trap gets, in the coin-flip picture

  • 2 directions1 in 4
  • 3 directions1 in 8
  • 10 directions1 in 1,024
  • 30 directions1 in ~10⁹
  • 100 directions1 in ~10³⁰
  • 1,0001 in ~10³⁰¹
  • 1,000,0001 in ~10³⁰¹⁰³⁰
Bar length is the number of zeros in the odds, on a linear scale where 100 directions fills the bar; the hatched rows run far off the chart. Values computed as ½ raised to the number of directions. This is the coin-flip picture: it shows the shape of the idea, and section 03 explains why the real landscape is stranger than a coin.

The mathematics of real random landscapes adds a twist the coin misses. High up, the directions are not independent coin flips. They push against each other, and the odds of a trap fall even faster than halving. What you meet up there is not a hollow. It is a saddle: a place that curves up in some directions and down in others. A mountain pass. Flat enough to slow you down. Not a place anyone stays.

02Where the idea comes from

Richard Bellman gave us the phrase “curse of dimensionality” in 1957, and he had good reason. Search a space by checking its corners, and the corners multiply beyond any computer. For most of the century that was the whole story about high dimensions: more is worse.

Then the same geometry started showing its other face. In 1997 the mathematician Paul Kainen argued that high dimension could make some computations easier, not harder. In 2000 David Donoho gave the idea its popular name in a lecture on the curses and blessings of dimensionality. In 2007 the physicists Alan Bray and David Dean worked out what flat spots look like on large random landscapes. And in 2014 Yann Dauphin, Yoshua Bengio and their colleagues carried that result into neural networks with a simple claim: the obstacle to training is not a proliferation of bad hollows. It is a proliferation of saddles.

03Every trap is near the floor

Here is the part most tellings skip, and it is the best part. Bray and Dean found that height decides. Far up the landscape, almost every flat spot has many ways down. The lower you go, the fewer downhill directions each flat spot has, until near the bottom most flat spots really are hollows.

So traps exist. They are just not scattered at random. They collect near the floor. And a hollow near the floor is nearly as good as the floor itself. Analysis of large networks suggests the hollows they settle into sit in a narrow band just above the best possible answer, and experiments find many of those hollows joined to one another by long, curving valleys that barely rise at all.

High up, you cannot get stuck. You can only slow down. By the time you can get stuck, you are nearly there.

That is the whole reversal in one sentence. The fog walker’s fear was being trapped high on the mountain. In a million dimensions the mountain does not offer that trap. It offers passes, and passes lead down.

04The second strangeness: room

The same geometry does something else, and this is the part that delights me most. Pick two directions at random in three dimensions and they point every which way. Pick two at random in a thousand dimensions and they are almost exactly perpendicular: ninety degrees, give or take about two.

Figure 02

Two random directions, and how often they land nearly perpendicular

  • 3 dimensions8.8%
  • 1020.1%
  • 10061.4%
  • 1,00099.3%
  • 10,000effectively all
Share of random pairs of directions that land within 5 degrees of a right angle. Computed by simulation, from 4,000 to 200,000 random pairs per row. The spread around 90 degrees shrinks like one over the square root of the dimension: about 33 degrees at 3 dimensions, under 2 degrees at 1,000.

Perpendicular means independent. Two directions at right angles do not interfere with each other. In three dimensions you can fit exactly three of those. In high dimensions, if you allow “nearly perpendicular,” the number you can fit grows exponentially with the dimension. A space with a few thousand coordinates has room for vastly more than a few thousand nearly independent directions.

That is how a model can hold more ideas than it has dimensions. Researchers at Anthropic showed in 2022, in small toy models, that networks do exactly this when ideas rarely show up together: each one gets its own nearly perpendicular direction, overlapping the others a little, and the model pays a small price in interference. In July I wrote that dimensions give meaning room. The strange part is how much room. Not a larger room. An exponentially larger one.

05The turn: our intuition runs backwards

Put the two strangenesses together and you get the thing that makes AI feel uncanny. Every instinct we have was formed in three dimensions. In three dimensions, a complicated system with more parts gets stuck more often, and a crowded room gets crowded. In a very high-dimensional space both instincts invert. Adding adjustable numbers adds ways out, so the bigger model is the easier one to train. Adding dimensions adds room faster than anything can fill it, so the model can hold more distinctions without them trampling one another.

One precision matters here. The dimension that gets you unstuck is the number of adjustable parts, not the size of the input. A small network fed an enormous input can still get trapped. Give it more knobs than it strictly needs and the traps thin out. The blessing belongs to the machine’s own freedom, not to the data’s.

Stuck is a three-dimensional word.

A classroom · a few years from now

The slide shows the diagram every textbook has used for fifty years: a ball rolling down a bumpy curve and settling in a dip that is not the lowest one. The teacher starts to explain the danger. A student asks how many directions the picture has. One, the teacher says. And the real one? Billions.

Then the picture is upside down, she says. The only dips are at the bottom. The teacher, who learned it the other way, stops and agrees.

06What I expect to see

  1. “It will get stuck” will keep being the wrong forecast.

    When a large model stalls in training, the cause will keep turning out to be the data, the objective or the schedule, not a trap in the landscape. Whatever limits AI in the years ahead will be data, energy and money, not geometry.

  2. Bigger stays easier to train.

    For the kinds of models we build, adding adjustable parts will keep making training more reliable, not less. The first sign this is ending will be large models that fail to train for reasons no one can trace to data or setup.

  3. Meaning becomes an address.

    Finding things by what they mean rather than what they are called, across contracts, photos, messages and records, will become the default everywhere. In these systems meaning is a direction, and nearby directions are cheap to find. I take this one further in Meaning Becomes an Address.

  4. Looking inside a model becomes geometry.

    The way we inspect and steer AI will be to find directions, the direction for a concept, a tone, a refusal, and move along them. Reading a model will look less like reading code and more like surveying a landscape.

  5. High-dimensional geometry enters ordinary education.

    Near-perpendicular directions, saddles instead of hollows, volume that lives at the edge: within a decade these will be taught to non-specialists the way probability was taught in the last century, because they are the geometry of the tools everyone uses.

07Why I care about this

Very few people share my enthusiasm for high dimensionality. When I bring it up, most people hear a technical detail. I hear the reason. It answers the question everyone asks about AI and almost no one answers: why does this work at all? Why doesn’t a search through billions of settings simply get lost?

This week I finally found a video that carries the same delight, Why High-Dimensional Space Is So Strange, from the channel math_is_fun. It is the reason this piece exists. If you want to feel the strangeness before you reason about it, start there. And if the vocabulary of dimensions and parameters is new, Dimensions Before Parameters sets it up.

08The ledger

Already true
Large networks trained by nothing more than stepping downhill reach low error routinely, and experiments find their resting places joined by low, curving valleys. The mathematics of random landscapes shows that high up, flat spots where every direction rises are astronomically rare. That random directions in high dimensions are nearly perpendicular is a theorem, not a hunch.
What has to happen
For the predictions to hold, the models we build have to stay generously oversized for their tasks, and real training landscapes have to keep behaving enough like the random-landscape picture that “high up means there is a way down” stays true in practice.
Where I am probably wrong
The clean mathematics is about random landscapes, and real networks are not random. Their curvature is flat in most directions, which the tidy picture ignores. Small networks do get trapped, and have been proven to. The theorems that make bigger networks easy to train assume sizes far beyond anything practical, and they speak only to training error, not to how well the model does on new data. Escaping a saddle is nearly certain but can take a very long time. If a future kind of model turns out to be hard to train because of genuine traps high on its landscape, this piece is wrong at its root.

09Back in the fog

You are standing where every step you have tried goes up. But you have tried four. The landscape the machine walks has billions more, and almost always one of them goes down. You only thought you were in a hollow because you could not count the directions.

It was never a hollow. It was a pass.

Background

More at johnrector.me.

Author: John Rector

John Rector is a Charleston-based entrepreneur, author, and AI strategist. He co-founded E2open, the supply-chain software company acquired for $2.1 billion in 2025, and in 2026 opened Charleston AI, a 3,000-square-foot lab that helps people and organizations understand and use artificial intelligence. He is the creator of The Reality Equation — a lecture series, book, and curriculum exploring attention, prediction, and how reality is experienced — and the author of more than two dozen books. He writes and speaks widely on artificial intelligence, attention, and the future of human work.

1 thought on “Every Trap Is Near the Floor”

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from John Rector

Subscribe now to keep reading and get access to the full archive.

Continue reading