A normative brain theory that survives contact with real neurons
The free-energy principle says the brain does variational Bayesian inference, and under Gaussian assumptions that maps neatly onto predictive coding. But Gaussian predictive coding demands neurons with linear input-output curves, which implies firing rates that go negative, and it forces every neuron to behave alike. Kataoka and Doya prove the correspondence survives a far broader class of distributions, the exponential family, recovering nonlinear, heterogeneous, strictly non-negative firing, with learning rules that are local and Hebbian-like. A 768-neuron two-layer spiking demonstration on sequential MNIST shows the dynamics doing what the theory promises.
Source: Extended predictive coding framework as variational free-energy minimisation under exponential-family assumption, A. Kataoka and K. Doya, arXiv:2605.30882v2, 2026. Primary source. Read in full via the arXiv HTML rendering, including the derivation to the second posterior cumulant, the local plasticity rules, and the heterogeneous Bernoulli simulation.
What the work claims
This is a theory paper with a supporting simulation. The free-energy principle (FEP) holds that perception is variational inference: the brain minimizes an upper bound on surprise called variational free energy. Earlier work showed that, if the brain's internal distributions are Gaussian and treated under the Laplace approximation, free-energy minimization takes exactly the form of predictive coding (PC): representational neurons encoding beliefs, error-coding neurons propagating mismatches, local dynamics that descend the free energy. The claim here is that the Gaussian cage was unnecessary. If the approximate posterior and prior belong to the exponential family of distributions, a vastly broader class, and the likelihood stays Gaussian, then gradient descent on variational free energy still takes predictive-coding form, up to the second cumulant of the posterior; the third cumulant is the price of the generalization and is neglected.1
Three consequences matter. First, the dynamics of a representational neuron become the sum of an attraction drive toward the prior and a prediction-error drive, exactly the PC architecture, but its activation function is now the derivative of the distribution's log-partition function, so any monotone increasing, differentiable input-output curve is admissible. That single move dissolves two standing embarrassments: firing rates can never go negative, and neurons no longer need to be identical. Second, the synaptic learning rules derived from the same objective are local, Hebbian-like plasticity with an uncertainty-weighted decay for prediction synapses, and prior-regulating plasticity whose requirement that a synapse compare against total input is mapped, speculatively but concretely, onto apical tuft integration and dendritic plateau potentials in pyramidal cells, with bottom-up error signals arriving at basal dendrites. Third, in a factorized setting each neuron can carry its own activation function, so a single population can mix sigmoidal, exponential, and hyperbolic response curves, and the authors argue this heterogeneity is representational: different tuning curves encode different member distributions of the exponential family.
How it works
The exponential family is the class of distributions whose log-density is linear in a set of natural parameters, which makes it the natural home of tractable variational inference. Kataoka and Doya parameterize each cortical layer's beliefs by such a distribution, assign each neuron an internal state (the natural parameter) and an emitted activity sampled from the distribution (spikes, if the distribution is Bernoulli), and let free energy descend. Two implementation choices do the heavy lifting. Gradient descent in natural parameters requires the Fisher metric, and the authors show the natural-gradient update can be computed by simple linear input reception within a single neuron, avoiding the biologically implausible requirement that each neuron integrate every other neuron's state. And for factorized posteriors, the log-partition function separates per neuron, which is the mathematical license for heterogeneous tuning curves: each neuron's strictly convex component function determines its own frequency-current curve, and the population still collectively descends the same free energy.
The demonstration is deliberately modest: a two-layer network, 512 neurons then 256, of heterogeneous Bernoulli units with gains sampled uniformly from 0.5 to 3.5 and shifts from 1 to 4, trained for 100 epochs on sequential MNIST with the digit changing every 20 timesteps. Sampling from the posterior makes it a spiking network. The reported dynamics match the theory: variational free energy at the first layer spikes at each digit switch and rapidly re-descends, the deeper layer's free energy is more stable across switches, absolute prediction error falls faster than free energy after each switch, confirming the objective is not mere error, and predictions of the current digit sharpen at the sensory layer. A few switch points show the second layer's prediction drifting toward the new digit before bottom-up evidence of the switch could have arrived, suggestive but explicitly not consistent across all switches, of anticipatory dynamics.
Where a skeptic should push
The single most load-bearing assumption is the error-coding neuron. The derivation assigns these units demanding properties: one-to-one, hard-wired connections from lower-layer representational neurons, and reciprocal synapses that are perfectly symmetric with their forward counterparts. The authors are admirably direct about this, listing the requirements and offering the standard split-into-excitatory-and-inhibitory-pops patch for sign, but symmetric weights remain the same biological implausibility that has dogged predictive coding since its inception, and nothing in this paper relaxes it. The third-cumulant neglect is a controlled approximation, but it means the correspondence is proven only up to second order, and the supplementary discussion concedes that going beyond a Gaussian likelihood runs into estimators with no closed-form neural implementation. The simulation is illustrative, not evaluative: no accuracy benchmark against a non-variational baseline, 768 units, one dataset, qualitative figures. And the pyramidal-cell mapping, basal error, apical prior, is posited, not tested. What is demonstrated is mathematical consistency and plausible dynamics; what is asserted is that cortex works this way.
Free energy as a dish-native training signal
The persistent engineering problem of organoid intelligence is not readout, it is objective. You can record a culture and decode it, but what should the culture optimize? Backprop is not available to tissue, and most closed-loop training schemes fall back on task error routed through a silicon controller, which makes the living substrate a peripheral, not a computer. This paper offers the most dish-compatible normative objective on the table. Variational free energy of a sensory stream is computable from exactly the signals an MEA setup already has: spikes in, prediction errors estimated from stimulation history, and the descent dynamics require only local, Hebbian-like updates plus a modulatory prior signal. No symmetric backward pass is needed for the learning rules themselves, the asymmetry problem is confined to the error units, not spread across every synapse. If a stimulated organoid's plasticity can be shaped so its dynamics approximate free-energy descent, training becomes something the tissue does to incoming spatiotemporal patterns, not something done to the tissue.1
Second, the paper quietly converts a liability into an asset. Every dish is a snowflake: gain, threshold, and response curves vary neuron to neuron, across batches, and across donors, and that variability is usually treated as noise to be engineered away. Here heterogeneity is load-bearing: mixed tuning curves mean a richer approximate posterior, and the theory guarantees the objective still descends. That is a rare alignment between what living tissue does spontaneously and what a normative theory rewards, and it suggests OI quality metrics should stop penalizing cellular diversity and start measuring whether the diversity is organized.
The threat is the mirror image of the promise. The error-coding requirements show how much architectural scaffolding predictive coding still smuggles in, and a dish does not ship with hard-wired, symmetric error units; if implementing the objective requires that scaffolding, the normative story becomes a silicon fantasy wearing biological clothes. There is also a hype-correction duty: a 768-neuron MNIST demo with qualitative figures is a proof of dynamics, not of competence, and the field has a habit of quoting free-energy language as if it were performance. Until someone shows an organoid whose free energy measurably descends on a structured input stream and whose predictions improve because of it, this remains a hypothesis about what tissue could compute, not evidence of what it does.
The bottom line
Established: the free-energy-to-predictive-coding correspondence extends from Gaussians to the exponential family, holding to the second posterior cumulant, with admissible activation functions any monotone increasing differentiable curve, strictly non-negative firing, neuron-level heterogeneity, and local plasticity rules derived from the same objective. Open: the biological realizability of the error-coding machinery, behavior beyond second order, and any demonstration that real tissue descends free energy. For biological computing the actionable hypothesis is that variational free energy is a trainable, dish-native objective: computable from MEA signals, descendable by local rules, and friendly to the heterogeneity cultures already exhibit. What would confirm it: a closed-loop organoid experiment showing free-energy descent and improving predictions on a structured input stream, with learning that survives pharmacological blockade of candidate pathways. What would break it: evidence that no biologically available mechanism can supply the required error signals without symmetric wiring, in which case the objective stays silicon-side and the tissue remains a readout, not a computer.
Frequently asked questions
What is the free-energy principle?
The claim that the brain performs variational Bayesian inference by minimizing an upper bound on surprise, variational free energy. Under Gaussian assumptions this minimization was previously shown to map onto predictive coding: representational neurons, error neurons, and local error-descending dynamics.
What does this paper add?
It shows the correspondence holds when posterior and prior belong to the exponential family, far beyond Gaussians, up to the second posterior cumulant. Activation functions become any monotone increasing differentiable curve, firing rates stay non-negative, and neurons may have heterogeneous tuning while still descending the same free energy.
Are the learning rules biologically plausible?
The derived rules are local and Hebbian-like: prediction synapses update with pre-post activity plus an uncertainty-weighted decay, and prior-regulating synapses compare against integrated input, which the authors map onto apical tuft plateau mechanisms. The error-coding neurons themselves, however, require hard-wired one-to-one and symmetric connectivity.
What was demonstrated in simulation?
A two-layer spiking network of 512 and 256 heterogeneous Bernoulli neurons, trained on sequential MNIST, exhibited free energy that spikes at input switches and re-descends, stable deeper-layer representations, sharpening digit predictions, and occasional anticipatory shifts. The demonstration is qualitative; there is no baseline accuracy comparison.
Why does this matter for organoid intelligence?
Variational free energy is an objective a dish could in principle descend with local plasticity, computed from signals a microelectrode array already records, and the theory rewards rather than punishes the cellular heterogeneity cultures naturally exhibit.
What are the biggest caveats?
Symmetric, hard-wired error-coding circuitry, a second-order approximation, a tiny qualitative simulation, and an untested mapping to real pyramidal-cell physiology. The mathematics is solid; its biological implementation remains asserted rather than shown.
References
- A. Kataoka and K. Doya. Extended predictive coding framework as variational free-energy minimisation under exponential-family assumption. arXiv:2605.30882. 2026. https://arxiv.org/abs/2605.30882. Accessed 2026-09-27.