Predictive coding survives the death of the Gaussian neuron
The most influential theory of cortical computation, the free-energy principle implemented as predictive coding, has always depended on an embarrassing assumption: that neurons behave like Gaussian variables, including firing at negative rates. This theory paper from OIST removes that assumption, showing the predictive-coding correspondence holds across the exponential family of distributions. The payoff is a framework where heterogeneous, nonlinear, spiking neurons are not a deviation from the theory but its natural substrate, and where learning runs on local plasticity rules rather than backpropagation.
Source: Extended predictive coding framework as variational free-energy minimisation under exponential-family assumption, arXiv:2605.30882, 2026. Primary source. Read the full HTML of the preprint, including the derivation, the simulation section, and the discussion of biological plausibility.
What the work claims
This is a theory paper with a small existence-proof simulation, not an experimental study, and it should be weighted accordingly. The central claim is a generalization: the known correspondence between the free-energy principle (FEP) and predictive coding (PC), previously derived only under Gaussian assumptions with a Laplace-approximated posterior, holds when the approximate posterior and prior belong to the exponential family of distributions, which includes but is far exceeds the Gaussian. The correspondence is exact up to the second cumulant of the posterior; the derivation neglects the third.1
Two consequential properties follow. First, the framework admits arbitrary nonlinear, neuron-specific frequency-current curves, as long as they are monotonically increasing and differentiable, so sigmoidal, exponential, and hyperbolic neurons can coexist in one population, and firing rates never go negative while the normative objective is preserved. Second, the learning rules derived from the same variational objective are local: a Hebbian-like term plus a term depending on total input intensity, which the authors connect to dendritic plateau potentials and BAC firing in the apical tufts of pyramidal neurons. A two-layer, 768-neuron spiking implementation trained on sequential MNIST with these local rules shows stable reduction of variational free energy and structured internal dynamics.
How it works
The free-energy principle casts perception as variational inference: instead of computing an intractable posterior, the brain adjusts the parameters of an approximate posterior to minimize variational free energy, an upper bound on sensory surprise. Predictive coding is the candidate neural implementation, with two functional cell classes: representational neurons carrying the current belief, and error-coding neurons computing the difference between top-down predictions and bottom-up signals. Under Gaussian assumptions the two map onto each other cleanly, but at the cost of linear neurons that can fire at negative rates.1
The technical move here is to parameterize posterior and prior as exponential-family distributions, whose log-partition function links natural parameters to expectation parameters through a Legendre transform. The gradient of that log-partition function is the neuron's activation function. This gives a concrete dictionary: a neuron's internal state plays the role of membrane potential, its sampled activity the role of spikes, and its measured frequency-current curve directly determines which distribution it encodes. Because every monotone differentiable curve is admissible, measured neuronal heterogeneity becomes a representational resource: diverse activation curves let the population encode richer, non-Gaussian posteriors.
The simulation makes this concrete. A two-layer network, 512 neurons then 256, uses heterogeneous Bernoulli posteriors, so sampled activities are literal binary spikes. Each neuron's activation is a sigmoid with its own gain, drawn uniformly from 0.5 to 3.5, and its own shift, drawn uniformly from 1 to 4. It is shown sequences of MNIST digits, one image per 20 timesteps, in a fixed incrementing order, and runs inference and learning simultaneously with no separate backpropagation phase, the plasticity time constant set 50 times slower than the inference time constant. After 100 epochs, test-phase dynamics show firing patterns that switch with each digit change, first-layer free energy that spikes at image switches and rapidly re-descends, a second layer whose representation is more stable, and occasional anticipatory shifts toward the next digit before it appears. The third-cumulant approximation is not proven but empirically supported by the monotone free-energy decrease.
Where a skeptic should push
The single most load-bearing assumption is that error-coding neurons as the mathematics requires them can exist in biological tissue. The authors are unusually candid about this in their own discussion: their error units must have no intrinsic dynamics, must propagate analogue, signed, non-spiking signals that can be negative, must receive one-to-one hard-wired connections from representational neurons, and must have feedback synapses perfectly symmetric with their feedforward counterparts. They note the tension with Dale's law, which constrains a neuron to be consistently excitatory or inhibitory, and propose splitting error units into excitatory-inhibitory pairs and invoking feedback alignment as partial fixes. These are patches, not demonstrations. For anyone hoping to port this framework onto self-wired tissue, this is the load-bearing wall: the theory needs precisely specified error channels that no known self-organizing circuit delivers.
Second, the empirical support is one toy: two layers, 768 neurons, ordered MNIST sequences, no performance comparison against any trained baseline, no fit to real spike recordings, and no test of whether the inferred latent structure corresponds to anything behaviorally meaningful. The demonstration shows free energy goes down, which it is mathematically constructed to do; it does not show the network infers well by any external criterion. Third, the likelihood stays Gaussian, and the correspondence is shown only up to the second posterior cumulant, so the theory's claimed breadth is itself bounded. The information-thermodynamics detour, connecting the log-partition function to canonical ensembles, is suggestive but explicitly speculative.
What this means for organoid computation
The non-obvious implication runs against a quiet prejudice in biological computing. Organoid cultures are embarrassingly heterogeneous: neurons with different gains, thresholds, and response nonlinearities, assembled by development rather than by design. Standard Gaussian predictive-coding accounts implicitly treat that heterogeneity as noise that complicates the theory. This paper inverts the relationship: within an exponential-family posterior, each neuron's activation curve is the Legendre transform that determines which slice of the distribution it encodes, so a population of diverse nonlinearities represents a richer posterior than a homogeneous one could. The heterogeneity of self-wired tissue stops being a manufacturing defect and becomes, in principle, a representational asset. That is a genuinely different posture for the field, and it is grounded in a specific mechanism, the gradient-of-log-partition to frequency-current mapping, not in vibes.
The mapping also yields a concrete assay. Because the theory reads a neuron's activation function as a statement about what it encodes, patch-clamp or MEA characterization of frequency-current curves in a cortical organoid becomes theory-laden measurement rather than mere phenomenology: a population with mixed sigmoidal and accelerating curves is, on this account, a population configured to encode a non-Gaussian posterior. Combined with the model's second prediction, that perceptual inference should show error-like signals and free-energy-reducing relaxation dynamics after perturbation, an organoid lab can attempt a real falsification: drive the tissue with structured input sequences, perturb it, and look for the spike-then-settle error signature. This is the rare grand theory that hands experimentalists a cheap, specific test.
The genuine threat is the same wall named in the skeptic section, seen from the organoid side. The framework's error units require analogue signed channels, one-to-one hard wiring, and synaptic symmetry, and an organoid's stochastic self-assembly provides none of these by construction. If predictive coding of this kind is what cortex does, then random organoid wiring may implement at best a degenerate version of it, and no amount of stimulation sophistication will conjure the required error-channel architecture. That would narrow what organoid substrates can compute without developmental patterning, which is an uncomfortable but important constraint for the field to internalize. There is also a hype-correction duty here: the free-energy principle has been criticized as unfalsifiably broad, and while this paper makes it more concrete, a two-layer MNIST toy cannot rehabilitate the whole program, and organoid marketing that borrows the vocabulary of variational inference should be held to the paper's own standard of mechanistic specificity.
The bottom line
Established: the free-energy-to-predictive-coding correspondence generalizes from Gaussian to exponential-family posteriors and priors, holding up to the second posterior cumulant, with local plasticity rules and admissible heterogeneous nonlinear spiking neurons, and a 768-neuron two-layer implementation trained without backpropagation shows stable free-energy reduction on sequential MNIST. Not established: that biological or organoid neural tissue implements this scheme, that the required error-coding circuitry exists anywhere outside the mathematics, or that the model infers in any externally validated sense. What would confirm it: recordings from cortex or organoids showing the predicted error-and-relaxation dynamics after structured perturbation, and a derivation or mechanism for signed error transmission compatible with Dale's law. What would weaken it: evidence that the third-cumulant approximation fails in richer regimes, or that no distinct error-coding population exists in real tissue. For organoid intelligence, the lasting value is a license to treat heterogeneity as compute and a specific, cheap experimental signature to hunt for.
Frequently asked questions
What is the free-energy principle in one paragraph?
It is the hypothesis that the brain performs perception by variational Bayesian inference: rather than computing an intractable posterior over world states, it adjusts an approximate posterior to minimize variational free energy, an upper bound on sensory surprise. Predictive coding is the proposed neural implementation, with representational neurons carrying beliefs and error neurons propagating prediction errors up a hierarchy.
What exactly was shown before this paper, and what is new?
Previously the correspondence between free-energy minimization and predictive-coding dynamics was derived only under Gaussian assumptions for prior, posterior, and likelihood, which forces linear neurons that can emit negative firing rates. This paper shows the correspondence holds when posterior and prior are any exponential-family distributions, up to the second posterior cumulant, while the likelihood remains Gaussian.
How can heterogeneity be a good thing computationally?
In the framework, a neuron's activation function is the gradient of its distribution's log-partition function, which determines which aspect of the posterior it encodes. Allowing each neuron its own monotone nonlinear curve lets the population represent richer, non-Gaussian distributions than a homogeneous linear population can, so measured diversity of frequency-current properties becomes representational capacity.
How is the network trained without backpropagation?
Inference and learning run simultaneously from the same variational objective, with local rules: a Hebbian-like term plus a term depending on total input intensity, run 50 times slower than the inference dynamics. The authors relate the second rule to dendritic plateau potentials and BAC firing in pyramidal apical tufts.
What did the simulation actually demonstrate?
A two-layer network of 512 and 256 spiking Bernoulli neurons with heterogeneous gains and shifts, shown ordered sequences of MNIST digits changing every 20 timesteps, reduced its variational free energy stably after 100 epochs of training, switched internal firing patterns with digit changes, and showed a more stable second-layer representation that occasionally anticipated the next digit.
What is the weakest link for biological plausibility?
The error-coding neurons must have no intrinsic dynamics, transmit analogue signed signals that can go negative, receive one-to-one hard-wired connections, and enjoy perfectly symmetric feedback synapses, all of which strain known biology and Dale's law. The authors propose excitatory-inhibitory splitting and feedback alignment as partial remedies, but these remain conjectural.
References
- A. Kataoka and K. Doya. Extended predictive coding framework as variational free-energy minimisation under exponential-family assumption. arXiv:2605.30882 [q-bio.NC]. 2026. https://arxiv.org/abs/2605.30882. Accessed 2026-10-09.