Essay · Artificial Intelligence · Follow-on
Not Is a Nudge
In logic, “not” is a U-turn: it points a statement the opposite way. Inside the machines that search our documents, “not” is a short step to the side, and nearly the same step for every sentence. That one difference explains why AI keeps finding “not safe” when you asked for “safe.”
The angle logic says should be 180°
14°
measured between these two sentences
“This medication is safe during pregnancy.”“This medication is not safe during pregnancy.”
Measured in all-MiniLM-L6-v2, a widely used open embedding model with 384 dimensions. A larger 768-dimension model gave the same 14°. Two sentences on unrelated subjects typically sit about 87° apart in the same model.
Sunday · 9:15 p.m.
You ask your assistant to pull every listing from your files that allows dogs. It returns six. The third one says, in bold, “No pets of any kind.”
It did not misread the listing. It found exactly what you asked about. It just could not tell “about dogs” from “dogs allowed.”
01My first explanation was almost right
When I wrote yesterday that meaning becomes an address, one fact kept nagging at me. A sentence and its negation get addresses right next to each other. My first explanation was geometric. In a low-dimensional picture, “not” is a 180-degree turn, an arrow pointing the other way. In thousands of dimensions, I thought, that rotation somehow collapses in on itself, and the opposite ends up beside the original.
The instinct was right that geometry is the answer. The mechanism was wrong, and the way it was wrong turned out to be the interesting part. The exact opposite direction exists in every number of dimensions. Take any arrow, flip every coordinate, and you have an arrow pointing precisely the other way, 180 degrees off, in a million dimensions as easily as in two. Nothing collapses. So why doesn’t “not” go there?
Because “not” was never flipping every coordinate. It flips very little.
02Two kinds of opposite
Aristotle drew the line we need here. In the Categories he separated opposites like good and bad from the plainer opposite of a statement and its denial, where one of the two must be true and the other false. Good and bad are contraries. “It is safe” and “it is not safe” are contradictories.
Machines are bad at both, for the same root reason. Their sense of meaning is built on company, on Firth’s old line that you know a word by the company it keeps. And opposites keep the same company. “Hot” and “cold” show up in the same sentences about the same weather, ovens and coffee. Linguists who study this have noted for years that sharing contexts cannot, on its own, tell a synonym from an antonym. In the same model as the hero card, “hot” sits far closer to “cold” than to “banana,” and not so very far behind “warm.”
A statement and its denial keep even closer company. They are the same sentence, plus one small word.
03What the measurements show
Eight ordinary sentences and their negations, run through the same embedding model alongside twenty unrelated sentences, tell a clear story. In every case the negation was the single nearest neighbor of its original, out of 35 candidates. None came anywhere near 180 degrees.
Figure 01
How far “not” turns a sentence, on a scale where logic says 180°
One result stands out. An earlier run also tested “The seller will take the boat lift with them,” which means the same thing as “the seller shall not leave the boat lift.” It landed 35 degrees from the original. The negation, which means the opposite, landed 18 degrees away. The sentence with the opposite meaning was closer than the sentence with the same meaning. It shared more words, and it kept the same company.
Notice the other thing the figure shows. Nothing in the whole set got near 180 degrees. The farthest two sentences in the whole set were barely past a right angle. The far pole, the true opposite, is sitting empty.
04The principle: “not” is a small share of a big sentence
Here is the geometry that replaces my first explanation. A sentence in one of these systems is not a single arrow for a single idea. In the terms of the first piece in this series, a space with thousands of directions has room for a great many ideas at once. So the address of “This medication is safe during pregnancy” carries all of them together: medicine, pregnancy, safety, a certain clinical tone, the shape of a factual claim. The yes-or-no is one ingredient among many.
Flip one ingredient and leave the rest alone, and the arrow barely turns. How far it turns depends only on how big a share of the sentence that ingredient carried.
angle = arccos(1 − 2 × share)share = the fraction of the sentence’s total weight carried by the part that flips
Figure 02
The turn depends on how much of the sentence “not” carries
This is where my intuition survives, rescued. In a picture with room for only one idea, a sentence is nothing but its yes or no, so “not” has to reverse everything, and it points 180 degrees away. In a space with room for thousands of ideas, the sentence is mostly everything else. The 180-degree turn has not collapsed. It has been diluted.
05The turn: logic multiplies, the machine adds
There is one more finding, and I think it is the heart of the matter. For each of the eight pairs, take the small step from the sentence to its negation, and compare those steps with one another. They point in much the same direction: a typical pair of steps has a cosine of about 0.6, where two random directions in that space typically score within about 0.05 of zero.
So “not” is a direction, just as my original picture said. It is simply the wrong kind. In logic, “not” is multiplication by minus one: the whole statement is reversed. In the machine, “not” is addition: the same short step, tacked onto whatever sentence it appears in.
Figure 03
Two meanings of “not”
Addition changes what “not” is. It stops being an operation on the sentence and becomes one more thing the sentence is about. “This medication is not safe” is, to the machine, a sentence about medication, pregnancy, safety and negation. “This medication is safe” is about medication, pregnancy and safety. Three of four ingredients match. Of course they are neighbors.
The machine did not ignore the “not.” It filed it as a topic.
And now the empty far pole makes sense. The exact opposite of “This medication is safe during pregnancy” would be the arrow opposite in every ingredient at once: a different subject, a different setting, a different tone, reversed. That is not “it is not safe.” It is not a sentence anyone writes. The 180-degree direction is the address of a meaning that does not exist.
Part of this is even by design. A search system is trained to measure what a passage is about, and a warning that a medication is not safe is exactly what you want to see when you ask whether it is safe. Nearness measures aboutness. Negation is a question of truth. The trouble starts only when we treat the first as the second. Tests built for exactly this find the standard kind of search embedding does worse than random guessing at telling a passage from its negation. Systems that read the question and the passage together do better, and still get only about half right.
Your office · a year or two from now
You ask again for every listing that allows dogs. This time the assistant says: “Nine listings mention pets. Five allow dogs. Three do not, including the Folly Road house. One allows cats only.” Under each answer sits the exact sentence it relied on.
It found the nine the old way, by nearness. Then it read each sentence the way you would, word by word, and let the “not” count.
06What I expect to see
Finding and reading become two separate steps everywhere.
Every serious AI system that answers from documents will find by nearness and then check each candidate with a slower reader that treats “not” as a reversal. Systems that skip the second step will be the ones that make the news.
The costly errors will live in one word.
The expensive mistakes will cluster where a single “not,” “no,” “never,” “unless” or “except” carries the whole meaning: medication labels, contract clauses, inspection reports, safety instructions. Everywhere else, the nudge is harmless.
Every answer will show its sentence.
Answers built from documents will quote the exact sentence they relied on, because a person can see a “not” at a glance that a machine may have treated as a detail. The quote will become the receipt.
People will write positively for machine readers.
People who know their words will be read by AI will state things positively: “the seller will remove the boat lift,” not “the seller shall not leave it.” Wording that says what is true, instead of denying what is not, will become a quiet rule of good drafting.
Training will narrow the gap, not close it.
Models trained on examples that pair sentences with their negations already get better at “not,” and that will become standard. But as long as one arrow has to summarize a whole passage, “not” will stay a small share of it, and the second step in prediction one will stay necessary.
07Why this one matters to me
Yesterday’s piece ended with a rule: find by meaning, confirm by the original. It had a rule but no reason. Now I have the reason. A contract clause and its denial will always be neighbors in meaning space, because “not” is a nudge there, not a U-turn. The boat lift example I used was not an accident of one model. It is how these systems are built.
And I like that my own first explanation was wrong in a useful way. I reached for “the rotation collapses,” and the measurements handed back something better: the rotation is still there, 180 degrees, fully available. The machine just never uses it for “not,” because to the machine “not” is not a reversal at all. It is one more thing a sentence can be about. This series began with high dimensions making AI surprisingly good at things. This is the same geometry making it surprisingly bad at one. The next piece shows the nudge shrinking as a passage grows.
08The ledger
- Already true
- In widely used embedding models, a sentence and its negation sit closer together than most paraphrases, and nowhere near opposite. Standard word embeddings put antonyms close together. Published tests show common search embeddings do worse than chance at telling a passage from its negation, and training on negated examples measurably helps.
- What has to happen
- For the predictions to hold, the dominant way of finding things has to stay a single summary arrow per passage, and the second, slower reading step has to stay cheap enough to run on every answer that matters.
- Where I am probably wrong
- These measurements come from two open models and a handful of sentences. They show the shape, not a law, and the angles ranged from 14 to 50 degrees, not “a few.” The consistent “not” step rests on only eight pairs. This is mostly a matter of training, not geometric destiny: a model trained hard enough on contradictions can push negations further apart. Large language models, the kind that write rather than search, do carry internal directions that track negation and truth, though researchers find that structure messy. If a new kind of search model learns to treat “not” as a genuine reversal while staying fast, the two-step prediction will soften, and I will have described a phase rather than a principle.
09Back to Sunday night
The listing said “No pets of any kind.” The machine read every word. It heard “pets.” It heard “no,” too. It just filed “no” as one more thing the listing was about, and it was right that the listing was about pets.
It only missed which way the sentence pointed.
Background
- Aristotle, Categories, ch. 10; De Interpretatione, ch. 7.
- J. R. Firth, “A Synopsis of Linguistic Theory, 1930–1955,” in Studies in Linguistic Analysis, 1957.
- Saif Mohammad, Bonnie Dorr, Graeme Hirst and Peter Turney, “Computing Lexical Contrast,” Computational Linguistics, 2013.
- Nikola Mrkšić et al., “Counter-fitting Word Vectors to Linguistic Constraints,” NAACL 2016.
- Kawin Ethayarajh, “How Contextual are Contextualized Word Representations?” EMNLP 2019.
- Tianyu Gao, Xingcheng Yao and Danqi Chen, “SimCSE: Simple Contrastive Learning of Sentence Embeddings,” EMNLP 2021.
- Orion Weller, Dawn Lawrie and Benjamin Van Durme, “NevIR: Negation in Neural Information Retrieval,” EACL 2024.
- Amira Alhamoud et al., “Vision-Language Models Do Not Understand Negation,” CVPR 2025.
- Miriam Anschütz, Diego Miguel Lozano and Georg Groh, “This is not correct! Negation-aware Evaluation of Language Generation Systems,” INLG 2023.
- Samuel Marks and Max Tegmark, “The Geometry of Truth,” COLM 2024; Lennart Bürger, Fred Hamprecht and Boaz Nadler, “Truth is Universal: Robust Detection of Lies in LLMs,” NeurIPS 2024.
- Models measured: all-MiniLM-L6-v2 and all-mpnet-base-v2, sentence-transformers.
- John Rector, “Meaning Becomes an Address,” September 2026.
- John Rector, “Every Trap Is Near the Floor,” September 2026.
- John Rector, “Big Rooms Drown Small Words,” September 2026.
More at johnrector.me.
2 thoughts on “Not Is a Nudge”