Research analysis · Learning rules and credit assignment

Local learning rules converge, but only inside a low-rank cage, and that cage is the real budget for training living tissue

A dish of neurons has no backpropagation. Every serious proposal for training organoid intelligence therefore relies on some local learning rule, one in which each synapse changes using only information physically available to it. A new dynamical-systems analysis draws the boundary of that whole family of rules in a tractable setting: random-feedback local learning converges, but the solutions it can reach are confined to low-rank changes of the initial wiring, with the rank capped by the number of output dimensions.

Source: Dynamics and Representation Structure of Local Approximations to Gradient-Based Learning in Linear Recurrent Neural Networks, Williams, Payeur, and Lajoie, arXiv:2606.00243, Universite de Montreal, Mila, and Imperial College London, 29 May 2026. Primary source. Read: the full arXiv HTML, including methods, all results sections, the representation-structure proposition, and the discussion.

What the work claims

This is a theory paper, not yet peer-reviewed, and it says so implicitly by staying in the one setting where the mathematics is exact: linear recurrent networks trained on a student-teacher task, in a data-aligned regime where the network's dynamics separate into independent orthogonal modes. Within that setting the authors compare three learning rules. Full backpropagation through time (BPTT) is the non-local reference. One-step truncated BPTT (tBPTT) keeps only the most recent temporal contribution. Random feedback local online learning (RFLO), in the family introduced by Murray and closely related to e-prop from Bellec and colleagues, replaces the precise gradient signal at each synapse with a random fixed projection of the output error, combined with a synapse-local eligibility trace.1

Three claims emerge. First, all three rules share the same optimal fixed-point manifold, the set of solutions that minimize the loss, but BPTT and tBPTT additionally admit an entire non-optimal manifold where the input and output pathways collapse to zero while positive loss remains. RFLO lacks this spurious manifold, because the random feedback term forbids the collapse. Second, RFLO is the slowest of the three near almost every point of the optimal manifold, and in trajectory simulations it can miss the nearest branch entirely, wander through parameter space, and converge to the opposite branch, inflating learning time drastically. Third, and most consequential, the authors prove (their Proposition 3.1) that whenever the random feedback matrix is a scaled identity, a class that includes the original RFLO formulation, learning is restricted to low-rank perturbations of the initial recurrent and input weights, with the attainable rank bounded by the output dimension.1

How it works

The reason exact gradient training is biologically implausible, and equally implausible for an organoid rig, is that computing the true gradient through a recurrent network requires information that no synapse possesses: the future error and the sensitivity of the current activity to every past parameter change, propagated back through time. RTRL does this locally in space but at prohibitive memory cost; BPTT does it exactly but non-locally. RFLO and its relatives escape by approximation: keep a trace at each synapse of how its recent activity would influence the output now, and multiply that trace by a random but fixed stand-in for the true error derivative. Each synapse then updates using only its own history and a global signal, which is precisely the shape of dopamine-gated plasticity in real tissue.1

The technical engine is a change of coordinates. In data-aligned linear networks each mode of the teacher maps onto one mode of the student, and the learning dynamics of every mode reduce to a small system of ordinary differential equations that can be solved for fixed points, stability, and convergence rates. The authors then test the theory outside its assumptions: in nonlinear student networks that are not data-aligned, the predicted mode dynamics still track experiment reasonably well for the two dominant modes, in runs averaged over six seeds. The rank result is the paper's structural contribution. Plotting the eigenvalue spectrum of the weight change learned by each algorithm shows BPTT and tBPTT filling out high-rank updates, RFLO leaving a spectrum that collapses after a handful of components (spectra averaged over 13 converged RFLO runs, 12 e-prop runs, and 20 runs for the other algorithms), and e-prop, whose eligibility traces carry heterogeneous or plastic time constants, escaping toward higher-rank solutions.1

Where a skeptic should push

The load-bearing assumption is that a linear, data-aligned student-teacher setting tells you something about the nonlinear, never-aligned networks you actually want to train. The authors are honest about this and provide one bridge experiment, but that bridge has limits: the theory degrades for the smaller modes, the inputs are white noise, and the teacher is constructed to be mode-aligned. In genuinely nonlinear networks the paper's own discussion concedes that RFLO's slow, wandering convergence could tip into trapping in local minima or outright divergence, the failure modes a linear analysis cannot see.

Second, a subtle and easily missed point cuts both ways: one-step tBPTT behaves almost identically to full BPTT in this regime, and the authors infer that data-aligned problems never required temporal credit assignment in the first place. That means the tasks where local rules look competitive are exactly the tasks where the hard problem, assigning credit across time, was absent by construction. Extrapolating RFLO's adequacy from such tasks to sequence problems that genuinely demand long-range credit would be exactly the overreach this paper warns against, even though the paper itself stops short of saying so bluntly. Third, the rank ceiling is proved only for feedback matrices that are scaled identities; the authors note e-prop's heterogeneous time constants seem to loosen it, but leave the mechanism open. The result is a bound for one member of the family, not yet for the family.1

What locality means for organoid training

Organoid intelligence has a credit assignment problem with no escape hatch: there is no way to backpropagate through living tissue, so every training protocol ever deployed on a dish, from spike-timing-dependent plasticity schemes to closed-loop stimulation, is a member of the local-rule family this paper characterizes. That makes the rank result the closest thing the field has to a budget. If the rule governing plasticity in a culture behaves like RFLO, then training can only sculpt the connectivity within a low-dimensional subspace around whatever wiring the tissue grew on its own, and the dimension of that subspace is set by the number of independent output channels the rig can read and reinforce, not by the number of electrodes and not by the number of synapses. The practical reading is uncomfortable and useful at once: adding more stimulation channels does not add trainable capacity unless it adds independent readout dimensions, and the slow-convergence finding says the field should expect training to take orders of magnitude longer than silicon baselines, with trajectories that can detour wildly before settling. Several published organoid-learning results look far better than this budget allows, which is precisely why the audit matters: under a strictly local rule, claims of rich learned behavior carry the burden of explaining where the high-rank degrees of freedom came from.

The opportunity is equally specific, and it comes from the paper's own exception. e-prop escapes the rank cage through heterogeneous eligibility time constants, and heterogeneity of intrinsic timescales is the one thing neural cultures have in surplus. A biological synapse does not need an engineered trace; its plasticity kernel is shaped by its receptor composition, its neuromodulator exposure, and its recent history, all of which vary cell to cell. If culture heterogeneity functions the way e-prop's engineered heterogeneity does, living tissue may be less rank-limited than silicon implementations of the same rule, an inversion worth testing head to head. The threat is the flip side: if the effective biological rule is closer to plain RFLO, the expressive ceiling is low, and the honest architecture for organoid computing is a frozen, spontaneously rich reservoir with a trained readout, not a trainable brain. One design lever falls out cleanly either way: because rank is bounded by output dimension, engineering diverse, independent readout targets is the highest-value investment a tissue-training program can make, cheaper by far than any attempt to make the tissue more homogeneous or more numerous.1

The bottom line

Established, within the analyzed regime: local random-feedback rules converge to the same optimal solutions as backpropagation but through slower, less stable dynamics, and the reachable solution set is confined to low-rank weight changes bounded by output dimension. Hypothesis: that these facts carry over to nonlinear networks and to biological substrates, and that heterogeneous eligibility dynamics loosen the cage in tissue. What would confirm the transfer: rank measurements of connectivity change after training in actual cultures or neuromorphic arrays, compared against output dimension, and a head-to-head of homogeneous versus heterogeneous plasticity kernels. What would break it: evidence that biological plasticity kernels carry enough state to evade the scaled-identity condition everywhere that matters, which would dissolve the cage but leave the convergence problem, the harder one, fully intact.

Frequently asked questions

What is RFLO learning?

Random feedback local online learning, a biologically plausible training rule in which each synapse keeps a local eligibility trace of its recent influence on the output and multiplies it by a fixed random projection of the output error. Synapses update without knowing the true gradient, using only locally stored information and a global signal.

What does low-rank mean here?

The change in the recurrent weight matrix that learning produces can be decomposed into independent components; a low-rank change has only a few such components. The paper proves the attainable rank is bounded by the number of output dimensions, so a network with one readout can only be trained along a one-dimensional slice of its possible wirings.

Why can't backpropagation be used on living tissue?

Backpropagation through time requires propagating error information backward through every past activity state of the network. No physical mechanism in a culture delivers a synapse the precise future-error derivative it would need, and no external computer can write per-synapse updates into tissue at the required precision.

Does this mean local learning is useless?

No. The rules converge to optimal solutions; they are slow and structurally constrained, not broken. The paper also shows e-prop, a close relative with heterogeneous eligibility time constants, reaches higher-rank solutions, which hints at designs, possibly biological ones, that loosen the constraint.

What was actually proven versus simulated?

The fixed-point manifolds, stability properties, and convergence rates are derived analytically in the linear data-aligned setting, and the rank restriction is a proved proposition. The tests beyond that regime, in nonlinear non-aligned networks, are numerical experiments over six seeds, and the authors state their limits plainly.

What should an organoid training program do differently after this paper?

Count independent readout dimensions as the budget for what training can express, expect long training horizons, and test whether biological heterogeneity of plasticity kernels acts like e-prop's engineered heterogeneity. Claims of rich learned behavior in cultures should be accompanied by evidence of where the high-rank degrees of freedom came from.

References

  1. E. Williams, A. Payeur, G. Lajoie. Dynamics and Representation Structure of Local Approximations to Gradient-Based Learning in Linear Recurrent Neural Networks. arXiv:2606.00243. 2026. https://arxiv.org/abs/2606.00243. Accessed 2026-09-29.