Project status — July 2026: Fathom is a completed research project. This page reports its results. Background is on the overview and method pages.
The result, in plain terms
The instrument went quiet on the trained model.
ΔAIS found its clearest signal in the untrained network and in the first few dozen steps of training. After that it fell away, steadily, to the level of noise, and stayed there for the rest of training. By the time the model could actually handle language, the measure could no longer detect genuine integration in its activity at all. The finished, competent model was, on this measure, harder to read than the near-random one it had started as.
This matched an outcome that had been written down and sealed before the test was run: if the measure never reaches significance on the trained model, then it does not detect processing in models of this kind, and the test cannot be cited in support of the theory. That rule was fixed in advance, and it holds now that the result is known.
The numbers
Three systems were measured with one and the same instrument, at the same settings, judged against the same significance threshold. Significance is reported as a z-score: the number of standard deviations by which the real signal’s score exceeds the average of its surrogates. The project’s threshold for a certifiable result is z ≥ 3.
| System | What it is | Peak ΔAIS (normalised bits) |
Peak z | Onset (z ≥ 3) |
Certifiable? |
|---|---|---|---|---|---|
| Logistic map | A simple equation passing from order into chaos | ≈ 0.50 | ≈ 278 | r ≈ 3.591 | Yes, sharply |
| Echo state network | An artificial system with a tunable memory | ≈ 0.005 | 2.7 | never | No |
| Language model, early training |
A system that genuinely learns, near its random start | 0.0071 (step 64) | ≈ 4.0* | step 32* | Briefly, then lost |
| Language model, trained |
The same model, once competent | 0.0004 | 1.28 | never | No |
| Language model, shuffled-text control |
The trained model reading scrambled words | ≈ 0 | ≈ 0 | never | No (as intended) |
Two things in this table carry the whole result. The trained model never clears the significance threshold on real text: its highest z, once past the first few dozen steps, is 1.28, well below 3. And the shuffled-text control stays at zero, which confirms the instrument is sound: it answers to genuine structure and falls silent on nonsense. The silence on the trained model is therefore a real finding, not a fault in the tool.
Technical note: how the numbers are arrived at. For readers who want the machinery. This can be skipped without losing the argument.
Active information storage. The base quantity is the active information storage of a signal, introduced by Lizier, Prokopenko and Zomaya (2012). For a symbol sequence X, with the recent past written as the block of the previous k symbols Xt−k:t, it is the mutual information between that past and the present symbol:
AIS = I( Xt−k:t ; Xt ) = H(Xt) + H(Xt−k:t) − H(Xt−k:t+1)
where H is Shannon entropy in bits. In Fathom the history length is k = 3, and the result is normalised by log2 of the alphabet size so that scores across systems with different alphabets are comparable. Every entropy term carries a Miller–Madow correction (Miller, 1955), which offsets the downward bias of plug-in entropy estimates; without it, longer recordings would appear more complex purely because they are longer.
The symbol sequence. Continuous activations are converted to symbols by ordinal coding (Bandt and Pompe, 2002): each short run of consecutive values is replaced by the rank-order pattern it forms, which is robust to drift and monotonic distortion. Fathom uses embedding dimension 3.
The surrogate correction. A raw AIS is close to meaningless, because autocorrelation alone inflates it. The reported quantity is therefore a difference against a null that preserves the signal’s linear structure while destroying everything else:
ΔAIS = AIS(signal) − mean AIS(surrogates)
The surrogates are made by the iterative amplitude-adjusted Fourier transform, or IAAFT (Schreiber and Schmitz, 1996). Each surrogate has, to close tolerance, the same power spectrum and exactly the same distribution of values as the original, but any nonlinear or genuinely computational structure is destroyed. The question ΔAIS answers is therefore precise: is this system’s apparent memory anything more than autocorrelation? Fathom draws 12 surrogates per measurement.
Significance. The 12 surrogate scores form a null distribution with mean μ and standard deviation σ. The z-score is
z = ( AIS(signal) − μ ) / σ
and a result is treated as certifiable only when z ≥ 3, sustained across at least two adjacent points of a sweep. This threshold is the origin of the distinction, laboured elsewhere on these pages, between a quantity existing and being provable: a real but small ΔAIS that sits below the surrogate noise cannot be certified, and Fathom reports it as silence rather than as signal.
Reading many parts. A model has many artificial neurons. Each chosen neuron’s activation series is measured separately, and the median ΔAIS and median z across neurons are reported. A median, rather than a sum or a first principal component, is used deliberately, so that a single strong shared rhythm cannot dominate the result.
Primary references.
Lizier, J. T., Prokopenko, M., and Zomaya, A. Y. (2012). Local measures of information storage in complex distributed computation. Information Sciences 208, 39–54. doi:10.1016/j.ins.2012.04.016
Schreiber, T., and Schmitz, A. (1996). Improved surrogate data for nonlinearity tests. Physical Review Letters 77, 635–638. doi:10.1103/PhysRevLett.77.635
Bandt, C., and Pompe, B. (2002). Permutation entropy: a natural complexity measure for time series. Physical Review Letters 88, 174102. doi:10.1103/PhysRevLett.88.174102
Crutchfield, J. P., and Young, K. (1989). Inferring statistical complexity. Physical Review Letters 63, 105–108. doi:10.1103/PhysRevLett.63.105
Miller, G. (1955). Note on the bias of information estimates. In Information Theory in Psychology, 95–100.
What the silence means, and what it does not
This is the most important part of the page, because the result is easy to misread in two opposite and equally wrong ways.
It does not mean language models lack processing complexity. The instrument read neurons one at a time. A system that learns almost certainly does its real work in the patterns spread across many neurons at once, and a measure that reads them singly cannot see that. The silence tells us that the typical single neuron, read on its own, carries little. It does not tell us the network as a whole is empty. A blunt instrument finding nothing is not the same as there being nothing to find.
It also does not mean Fathom succeeded in measuring anything about minds. The measure went quiet exactly where the interesting question began. The test returned no verdict on its actual question, whether processing complexity switches on sharply or rises smoothly in a learning system, because the instrument could not reach the system.
What the result does show is more modest, and more useful. The method has a limit, and the limit can be located. ΔAIS works on small, discrete, low-dimensional systems and goes blind on large, continuous, learned ones. That boundary is not a nuisance to be tidied away. It may be the most substantial thing Fathom found, because if the physical trace of the essay’s quantity lives in exactly the high-dimensional, many-parts-at-once structure this instrument cannot capture, then the barrier is not a failure of cleverness but a feature of the ground itself.
How this sits with the essay
The essay predicted no sharp break in nature, and Fathom found none. An earlier stage of the project, sweeping the logistic map, did throw up what looked like a sudden threshold: the sharp onset at r ≈ 3.591 in the table above. But on inspection most of that jump was a property of the measuring, not of the world. Any instrument with a significance threshold will appear to switch on suddenly as a smooth signal climbs past the point where it can be proven, even when nothing underneath has jumped at all. The logistic map’s jump sits on top of a genuine change in the system’s dynamics, which is why its score really does leap there; strip that special case away, as the echo state network does, and the apparent threshold dissolves into a smooth rise that never certifies. So the apparent break was set aside. The essay’s central prediction was not overturned.
But nor was it confirmed, and honesty requires saying so. The essay had already granted the difficulty in its own words:
“The measurement problem is real and unsolved. But an unsolved measurement problem is not an untestable theory. It is an invitation to the kind of science that Darwin’s first sketch of natural selection eventually made possible.”
Fathom is one answer to that invitation, and what it found is that the measurement problem is harder and more specific than it first appears. The tool works where systems are simple, and loses its grip on precisely the systems, learned and richly interconnected, where the essay’s most interesting claims are staked. The theory stands where it stood. What has changed is that we now know, with some precision, where the current means of measuring it run out.
Why the project closes here
The pre-registered question was tested and answered, even if the answer was that the instrument cannot yet see this kind of system. Running the same method for longer would gather more of the same silence without addressing the reason for it. The sensible course is to stop, record what was learned, and design the next study to look where this one could not.
Two lessons carry forward. The first is that the gap between a quantity existing and being provable is not a technicality but may be the real shape of the whole problem. The second is that a measure which reads parts one at a time is the wrong tool for a system whose intelligence lies in how its parts act together. The obvious next step reads the parts jointly rather than singly, using the data already gathered, so no new model runs are needed. Whether it recovers a signal, stays silent, or fails its own controls is genuinely unknown, and, in keeping with the method, its meaning will be fixed in advance, before the first number is computed.
The theory in Consciousness and Complexity is untouched by this result, neither proven nor refuted. Fathom’s contribution is narrower and, in its way, more honest than a verdict would have been. It shows exactly where the instruments we have today stop being able to see.