Research analysis · Emergent computation

Prediction emerges in networks trained only to perceive

Train a recurrent network to clean up noisy music and nothing else. Afterwards, its internal states encode predictions about what comes next, its responses scale with surprise, and its error is linearly readable, even though no predictive objective was ever imposed. The catch: all of this appears only inside a specific band of input noise, and the band, not the emergence, is the transferable result.

Source: Prediction emerges in RNNs trained for perception (Gupta and Tabas), arXiv, September 2026. Primary source. Read in full via the arXiv HTML, including the results, the detection-theory analysis and the methods; training hyperparameters and supplementary robustness checks were read for the main design choices.

What the work claims

This is a simulation study in computational neuroscience, and its claim is a reframing of the predictive-processing debate. Predictive processing holds that the brain builds a generative model of its sensory world and uses it to predict incoming input; the longstanding objection is that prediction is expensive, so why would a system whose job is perceiving the present evolve to compute the future at all. The authors' answer, demonstrated rather than merely argued: it does not have to decide to. Optimise a recurrent network for perception alone, here operationalised as denoising, and the machinery of prediction shows up by itself, but only when the input is noisy enough that the present observation is insufficient and clean enough that structure can still be learned.1

The demonstration is specific. Networks trained at intermediate noise acquired three signatures that predictive-processing researchers treat as the empirical markers of a predicting brain: hidden states from which the next sensory token can be predicted better than by a first-order Markov model; population response magnitudes that track prediction error; and states from which the prediction-error vector itself is linearly decodable. Networks trained at very low or very high noise showed none of these. The authors also derive, with detection theory, why the band should exist: below it, single observations are nearly always self-sufficient, above it, observations no longer support learning the statistics prediction would need to exploit.

How it works

The substrate is deliberately plain: single-layer gated recurrent networks trained with binary cross-entropy to recover latent musical tokens from noisy observations. The stimuli are 700 unique Bach compositions, drawn from 1,432 MIDI files and tokenised into 12-dimensional binary chroma vectors, split 490 for training, 70 for validation and 140 for testing. Gaussian noise at 14 amplitudes, from nearly clean to a signal-to-noise ratio of 0.5, corrupts the tokens. Every combination of noise level and network size, from 8 to 256 units, was trained five times from independent initialisations, 420 networks in total, so the effects are population-level rather than single-model anecdotes.

The test for emergent prediction is disciplined. After denoising training, the recurrent weights are frozen and a linear readout is trained on prediction alone. The readout counts as evidence of an internal prediction only if it beats a conservative Markovian benchmark, a first-order model given the ground-truth current token. In the case-example network, 64 units trained at noise standard deviation 0.25, the readout beat that benchmark with a small but highly reliable effect, while remaining well below a network trained end-to-end on prediction, which the authors treat as the empirical ceiling. For the response signature, they correlate the squared magnitude of the hidden-state update triggered by a new observation against the squared prediction error, controlling for the change in the latent token itself, and the partial correlations were positive in every one of the 140 test compositions. For the readout signature, ridge regression decoding the error vector from hidden states beat even the most stringent benchmark, a decoder given current and previous observations.

The band structure is the theoretical core. Prediction emerged for training noise between roughly 0.15 and 0.4 in units of the binary signal amplitude. The authors' calculation makes the bounds intuitive: at the lower edge, a single observation already identifies the latent token with probability above 0.99, so there is nothing to gain from a generative model; at the upper edge, that probability falls below 0.27 and the observations stop supporting the learning of statistical structure at all. Prediction is worth computing only in between, exactly as filtering theory would predict.

Where a skeptic should push

The load-bearing assumption is that denoising tokenized Bach is a fair metaphor for perception. It is a structured, discrete, low-dimensional domain with stable long-range regularities, chosen so that prediction is possible yet not trivially Markovian. Natural sensory worlds differ in dimensionality, nonstationarity and noise statistics, and the paper's single domain cannot tell us how the band moves when they do. The right stance is that the study establishes existence, not ubiquity: emergence under one well-characterised regime, with a theory for where the regime's edges should lie.

Second, the effect sizes are modest where they matter most. The signature-of-prediction readout advantage over the conservative benchmark is statistically overwhelming but small in absolute terms, and the emergent predictor sits closer to a first-order model than to a network actually trained to predict. These networks are doing something predictive, but they are not doing it well by the standards of a system built for the purpose. Third, everything here is in silico with rate-like units and linear probes; nothing yet says a spiking, biological substrate with plasticity dynamics of its own would land in the same band. The detection-theory bounds are the most portable element precisely because they make no reference to the network at all.

What organoid training objectives can omit

The obvious implication is that a predictive internal model, the thing many organoid-intelligence programmes quietly hope their cultures contain, may not need to be engineered in. If a substrate is optimised for a perceptual task under the right input statistics, predictive structure is a free by-product rather than a separate achievement. One prized computation, next-state prediction, can arrive uninvited when the training setup is right. The same logic reframes spontaneous activity: a culture left to its own dynamics may already encode predictions about whatever regularities its stimulation history contains, whether or not its experimenters designed for that.

The non-obvious implication is the band, and it is a direct design constraint. Predictive structure emerged only when single observations were insufficient but structure remained learnable. Closed-loop organoid protocols typically deliver clean, low-entropy stimulation: precisely repeated spatiotemporal patterns, tightly voltage-controlled pulses, deterministic game inputs. By this paper's logic, that is the noiseless regime, the one where there is no pressure to build a generative model and no predictive signature to find. A protocol that wants its tissue to model the world's statistics has to put noise in on purpose, at a level tuned so the present input is ambiguous but the sequence statistics are learnable. That is a concrete, testable protocol change: sweep stimulation entropy and look for the emergence window, rather than assuming the most controlled stimulus is the best teacher. The flip side is a warning about over-noising: past the upper edge, the substrate cannot learn the structure at all, and more noise just degrades.

The paper also hands the field an assay battery, and this may be its most useful export. Whether an organoid has actually built a model of its input statistics is not answerable from evocative raster plots; it requires benchmark-beating comparisons, exactly the discipline the authors model. The three-step check translates directly: freeze the tissue's state, train a linear decoder for the next stimulus element, and require it to beat a first-order model given the same input; test whether response magnitude scales with stimulus surprise, controlling for the stimulus change itself; and test whether the surprise vector is linearly decodable from population activity against a decoder given the raw stimulus history. Passing all three would be far stronger evidence of predictive computation in tissue than any mismatch response on its own. The threat implicit here is symmetrical: if signatures of prediction can arise without a predictive objective, then seeing mismatch-like responses in a dish is weak evidence of much at all, and the burden of proof moves to the benchmark comparisons.

The bottom line

Established in simulation: recurrent networks trained solely to denoise structured sequences acquire encodable predictions, surprise-scaled responses and decodable prediction error, without any predictive training signal, and these signatures appear only at intermediate input noise, with detection-theory bounds of roughly 0.99 single-observation recoverability at the lower edge and 0.27 at the upper edge. Not established: that biological or organoid tissue would do the same, that the band sits at the same noise levels outside tokenized music, or that emergent prediction is strong prediction, since the effect sizes are modest against prediction-trained ceilings. For organoid intelligence the durable contribution is twofold: a design principle, stimulation noise is a control knob that decides whether a culture is pressured to build a model of its input statistics at all, and an assay, benchmark-beating frozen-state decoders rather than evocative responses, as the evidentiary bar for claiming predictive computation in living tissue. What would confirm the transfer is an organoid study that sweeps closed-loop stimulation entropy and reports where, if anywhere, the three signatures emerge; what would break it is finding that tissue plasticity under those regimes produces none of the signatures even inside the predicted window.

Frequently asked questions

What exactly emerged in these networks?

Three signatures of predictive processing, none of which the networks were trained for: hidden states from which the next token was predictable better than by a first-order Markov model, hidden-state updates whose magnitude tracked prediction error, and states from which the prediction-error vector was linearly decodable. All three appeared only for networks trained at intermediate noise levels.

Why does the noise level matter so much?

Because prediction is only worth computing in a middle regime. When single observations almost always identify the latent state, there is nothing for a generative model to add; when noise swamps the observations, the statistics prediction needs cannot be learned. The authors compute recoverability above 0.99 below the band and below 0.27 above it, with emergence in between.

Does this prove the brain uses predictive processing?

No. It is a simulation study on one well-structured domain, tokenized music, with rate-like units and linear probes. It demonstrates that predictive machinery can arise from perception training alone, and offers a theory for when it should, but it does not record from animals and does not settle the biological debate.

How should an organoid lab use this?

Two ways. Treat stimulation entropy as a designed variable: overly clean closed-loop stimuli sit in the regime where nothing pressures tissue to build a model, so sweep the noise and look for an emergence window. And adopt the assay: freeze the substrate's state, train a linear next-stimulus decoder, and require it to beat a first-order benchmark before claiming the culture predicts.

Are the effects large?

Statistically decisive, practically modest. The predictive-readout advantage over the conservative Markovian benchmark was small in absolute terms, and the emergent predictors sat far below networks trained directly on prediction. The value of the paper is the mechanism and the band structure, not benchmark-scale performance.

What would falsify the transfer to living tissue?

An organoid closed-loop study that sweeps stimulation noise through the predicted window and finds none of the three signatures at any level, or finds them uniformly at all levels with no band structure. Either result would say the silicon band does not transfer, and would redirect the search toward what differs, most likely the plasticity rules.

References

  1. A. Gupta and A. Tabas. Prediction emerges in RNNs trained for perception. arXiv:2609.02739 [q-bio.NC], 2026. https://arxiv.org/abs/2609.02739. Accessed 2026-09-30.