"Rest your attention on the breath," the instruction goes, and for perhaps four seconds you do. Then you are somewhere else: a reply you never sent, the hum of the fridge, the ache in one knee. You notice you have drifted. You come back. The breath is still there, patient. You will leave it and return to it a hundred times in twenty minutes, and that returning, not the staying, is the whole of the exercise.
The word for what you are practicing is attention, and it is a strange word to have handed, a few years ago, to a machine.
In 2017 a design called the transformer changed how computers read language, and its central trick was named, without irony, attention. Here is what the name points to. When the machine takes in a sentence, each word or piece of a word sends out a kind of question: of everyone else here, who is relevant to me right now? Every other piece answers with an address. The questions and the addresses are compared, the matches are scored, and those scores become weights. Each piece is then rebuilt as a blend of the others, pulled in according to those weights.
That is the heart of it. Attention, in a machine, is a distribution of relevance across what it is allowed to see.
This is where the two worlds split.
Patañjali's yoga sutras give the practice a precise name: dhāraṇā, the binding of the mind to a single place. A later sutra describes ekāgratā, one-pointedness, as distraction wanes and a single focus rises. The meditator's labour is subtraction. To attend to the breath is to let the knee and the fridge and the unsent reply grow quiet, to narrow the field until one point is left, and lit.
The machine makes a different move. Within the sequence in front of it, each attention head assigns weights to the positions it is allowed to see. Some weights become vanishingly small. Masks can remove positions altogether. But concentration is not the point. The mechanism computes which pieces matter to this piece, and in what proportions. It does not hold one object against a desire to leave.
One practice attends by returning to one thing. The mechanism attends by distributing relevance across many.
So the same word names two different gestures, and the difference is not a technicality.
Here is the part that took me a while to see. Human attention is precious because it leaks. It wanders, it tires, it can be bought and stolen and thrown away. The reason a person can give you their attention, and mean something by it, is that they could have given it to anything at all and chose you. The offering is real because the supply is small, and the mind would rather be elsewhere.
Machine attention has a computational cost, but no felt one. Within a context, it calculates weights without fatigue or preference. There is no evidence of a mind wandering away or returning; those verbs describe us, not matrix operations. What we called attention when we named the mechanism was the arithmetic of relevance: which inputs, in what proportion, feed the next step.
Human attention also sorts relevance, but dhāraṇā asks for something more deliberate. The breath is not necessarily more urgent than the knee. You choose it anyway, and keep choosing it, against the drift of your own mind.
We say we pay attention, as though it were money, and the phrase is exacter than we treat it. Attention is spent. It comes from a small purse, it runs low, and whatever you lay it on is precisely what you have chosen not to lay it on.
The machine pays in computation, not in possibility. It gives nothing up when one word outweighs another. We do, and what we surrender may be the weight that makes attention worth giving.
Try the distinction
Many weights, one return
Inspect an illustrative machine distribution, then notice the different gesture of bringing one wandering focus home.
Machine attention
Select a word to inspect its share of this fixed example.
These numbers are illustrative, not the output of a real model and not a measure of global word importance.
Dhāraṇā
Choose where the mind wandered. Then return it to the breath.