Project status — July 2026: Fathom is a completed research project. This page describes its method. For the findings, see the Results and Conclusion.
The essay defines processing complexity as “how far a system draws its own past together into its present.” Fathom’s job was to turn that definition into a number that behaves properly: high where history is genuinely used, zero where it is not, and honest about the difference.
The measure: ΔAIS
The starting point is a quantity called active information storage: how much knowing a system’s recent past reduces uncertainty about its present. On its own this is not enough, because a system can score highly for trivial reasons, such as a slow drift or a repeating cycle. None of that is the integration the essay means. The essay’s own image is the river that “carries its whole history in the silt it moves and integrates none of it.” The past shapes the present without ever being genuinely gathered up.
So the measure is always a difference. For any signal, Fathom builds a set of surrogates: artificial versions with the same frequency content and the same distribution of values, but with any genuine temporal structure destroyed. These are the imitations that look the same on the surface but hold no real memory. The headline quantity is the active information storage of the real signal, minus the average across its surrogates.
If a signal’s apparent memory is only an artefact of its statistics, its surrogates score the same, and the difference is zero. Only genuine integration, structure that lives in the real signal but not in its matched shadows, produces a positive result. This one idea is what lets the instrument tell the essay’s integration apart from the river’s mere accumulation.
Judging significance
A positive result is not automatically meaningful, since it might be noise. Each measurement is therefore given a z-score: how many standard deviations the real signal sits above the spread of its surrogates. Fathom’s threshold is three, which is roughly a result that would arise by chance less than about once in a thousand times. Below that threshold a value is recorded but treated as uncertifiable, perhaps real but not demonstrable.
The gap between a quantity existing and being provable turns out to matter a great deal, and the Results page comes back to it.
Reading systems fairly
Complex systems have many parts. Fathom’s rule, fixed in advance, is to measure each part on its own and report the median across parts, never a single overall summary of the whole system’s activity. An overall summary flatters systems that happen to have one strong shared rhythm, which would quietly load the dice. The median across many separate parts is deliberately blunt and deliberately even-handed.
Deciding before looking
The most important feature of the method is not statistical but procedural. Before any data was analysed, Fathom fixed and sealed its hypotheses, its significance threshold, its parameters, and a table setting out what each possible outcome would mean for the theory. Once sealed, the plan could not be quietly adjusted to fit the data. This is pre-registration, a discipline borrowed from clinical trials, and it exists to prevent the natural human habit of deciding what a result means only after seeing which result you have.
The essay asks for this kind of seriousness. It states plainly what would force it to be given up, a clean neural line dividing the conscious from the rest, and a project testing it has to be held to the same standard: say in advance what each outcome will mean, then stand by it.
What was measured
The decisive test used a publicly released family of language models whose developers had kept snapshots throughout training, from the first random state to the finished model, more than 150 stages in all. Fathom watched a fixed set of individual artificial neurons, chosen at random by a rule settled before any results were seen, across every stage, as the model read a fixed passage of text. The text was held out: drawn from the same kind of material the model learned on, but never seen during training, so the measure could not be fooled by the model reciting from memory.
Two controls guarded the result. An untrained baseline, the model before any learning, showed how much apparent structure comes from the input alone. A shuffled-text control, the same words in scrambled order, gave a further check, since a trained model has nothing genuine to integrate in word salad. If the measure had lit up there, it would have been reading something other than genuine integration, and the test would have been void.
Every parameter above was fixed before the analysis ran. The results are on the next page.