A memory-limit theory of cortex and subcortex doubles as a job description for hybrid organoid-silicon computers
A dish of cortical tissue has no dopamine system, no striatum, and no reward circuitry. Pairing it with a silicon controller is usually framed as a workaround for that lack. A theory paper from RIKEN and the University of Tokyo suggests the pairing is not a workaround but the correct architecture, and it specifies which side should own the reward.
Source: Cortex and subcortex play distinct roles over learning when cortical memory is limited, Farrell and Toyoizumi, arXiv:2606.00667, RIKEN Center for Brain Science and University of Tokyo, 30 May 2026. Primary source. Read: the full arXiv HTML, including methods, all results sections, the experimental-proposal section, and the discussion. Code was listed as being prepared for release and was not yet public at access time.
What the work claims
This is a theory paper in a deliberately toy setting: a binary decision tree of depth two or three where reward sits only at the leaves and relocates halfway through training, with 200 episodes per phase. A model-free learner, SARSA(0) with a backward-propagating update, plays the role of subcortical habit learning. A model-based learner tracks the environment's transition probabilities but is granted only a small memory, four to eight slots, far fewer than the number of edges in the tree. The paper's question is how that scarce memory should be allocated, and it compares two antipodal strategies. MAXREWARD spends every slot tracking the edges that lead to the currently rewarded leaves. MAXREACH spends them maximizing arbitrary goal-reaching, tracking the edges that make the most destinations reachable, regardless of where reward sits now.1
The finding is clean. When memory is ample, the two strategies coincide and both work. When memory is scarce, MAXREWARD wins while the reward stays put and MAXREACH wins in the window right after the reward moves, and the size of that window grows with how often the world changes. A planner that ignores reward and learns the structure of the environment instead adapts faster to nonstationarity, because when reward relocates, structure knowledge converts instantly into policy while reward knowledge has to be rebuilt from repeated experience. The authors draw the functional dissociation explicitly: cortex, as the expensive memory-limited module, should specialize in general structure learning, while subcortical circuits, cheap and reward-dense, should own reward-based learning.1
How it works
The mechanism is a conversion-rate asymmetry. The model-free learner needs reward to be delivered repeatedly at a location before its values shift, so after a reward move it is slow by construction. The model-based learner, if it spent its memory on structure, already knows the transition graph; when the reward signal flips, it recomputes values from its model in one step, with no new experience required. Reward information is used by the two modules in fundamentally different ways: the model-free module consumes reward as training signal, while the structure-tracking module needs reward only to read out a policy from a model it would have built anyway. This is not the familiar exploration-exploitation trade-off at the level of action choice; it is a budget allocation problem one level up, at the level of what a memory-constrained system bothers to represent.1
The paper then does what a good theory paper should: it reaches toward biology. It notes that cortex receives sparser dopamine innervation than striatum, consistent with cortex not needing dense reward signal to keep learning useful structure; that humans are observed to learn transition statistics without explicit reward; and that the hot-cache hybrid, where a transient reward-tracking allocation spins up for a new task and is later pruned as habit learning catches up, matches the observed dynamics of cortical spines, which form transiently in response to a newly rewarded task and are subsequently eliminated. Simulations averaged over 40,000 trials per condition (4,000 with a variance-reducing policy-average estimator) support the qualitative patterns, and the authors sketch a consistency-analysis framework for testing which strategy an animal, or an experiment, is actually using.1
Where a skeptic should push
The gap between the toy world and any real deployment is wide, and the authors say so themselves. A regular binary tree with two actions, tabular transition tracking, and a clean single reward relocation is the simplest environment on which the distinction could be demonstrated; whether the ordering of strategies survives continuous state spaces, parametric function approximators, partial observability, or stochastic reward is untested. The model-free learner is deliberately weakened to SARSA(0); eligibility traces or replay would narrow the conversion-rate gap that produces the result, possibly eliminating it. There is also an acknowledged structural flaw the reader should not skip: untracked edges are assigned probability zero, so the model-based planner systematically underestimates reward that lies off its tracked subgraph, a bias that would matter in any sparse-reward real task. Finally, the code was not yet released at access time, so the simulation details rest on the text alone.1
On the biology side, the dopamine-sparsity and spine-pruning correspondences are suggestive, not evidence; the review-level literature the paper cites is itself contested, as the authors note. The right summary of the epistemic status: a clear proof of concept that memory scarcity can invert which knowledge a planner should store, plus a hypothesis-generating map onto cortical-subcortical anatomy, with the animal-level predictions explicitly framed as testable rather than established.
Division of labor for hybrid wetware systems
Read as organoid engineering rather than neuroscience, the paper hands the field a job description. The cortical organoid is the expensive, memory-limited, slowly plastic module; it has rich recurrent dynamics and no reward system. The silicon controller is the cheap, fast, reward-soaked module; it computes error signals trivially and never forgets. The naive hybrid design, and most published closed-loop organoid training, pipes task reward into the culture as directly as possible, pairing stimulation with success. The paper says that is MAXREWARD, the strategy that wins only while the task never changes, and it wastes the tissue's scarcest resource, its limited plasticity budget, on exploiting the current objective instead of learning the task's structure. The non-obvious prescription: let the tissue learn reward-free dynamics, the transition structure of the input world it is embedded in, and let silicon own the objective entirely. Structure learned by the substrate transfers across objectives instantly at the silicon boundary; reward burned into plasticity has to be unlearned every time the task changes, and unlearning in tissue is the slow, unreliable direction.1
There are two further transfers. The hot-cache result maps directly onto stimulation gating: a brief window of reward-paired plasticity when a new task arrives, then consolidation and release of the allocation, is both what the hybrid should implement and what the culture's own biology plausibly does, so the controller should treat plasticity windows as a schedulable resource with explicit time constants rather than a permanent connection. And the probability-zero flaw is a warning specific to hybrid systems: when the silicon controller models what the tissue has learned, anything untracked reads as impossible, so the controller will systematically overestimate the culture's competence gaps and can misplan interventions; a hybrid stack needs an explicit uncertainty model over the tissue's untracked state. The governance angle is real too: once the objective lives entirely in silicon and the competence lives entirely in tissue, the question of who or what owns the policy stops being philosophical, and an audit of the silicon reward module becomes an audit of the hybrid's behavior. The opportunity is an architecture that plays to each substrate's costs; the threat is that a mis-specified reward boundary in the controller silently becomes the trained character of the living part.1
The bottom line
Established, in simulation: under memory scarcity, reward-blind structure tracking outperforms reward tracking immediately after environmental change, and the advantage scales with nonstationarity. Hypothesis: that this inversion explains a cortical-subcortical division of labor and should guide hybrid designs that pair living tissue with silicon control. What would confirm it: the consistency-analysis tests the authors propose on animal data, and a hybrid organoid experiment in which structure-trained tissue transfers across reward reassignments faster than reward-trained tissue from the same batch. What would break it: showing that with realistic plasticity dynamics and richer model-free baselines the conversion-rate gap closes, at which point the inversion, and the job description built on it, disappears.
Frequently asked questions
What are model-based and model-free learning?
Model-free learners store values for states or actions directly from experienced reward, like habit learning. Model-based learners build an internal model of how actions change the world and plan through it, which lets them recompute optimal behavior immediately when goals change, if the model itself is retained.
What are MAXREWARD and MAXREACH?
Two strategies for spending a scarce memory budget on transition knowledge. MAXREWARD tracks the edges leading to currently rewarded locations; MAXREACH tracks edges that maximize how many goals are reachable at all, ignoring where reward sits now.
When does each strategy win?
With ample memory they tie. With scarce memory, MAXREWARD wins while the reward location is stable, and MAXREACH wins in the period after the reward moves, with the advantage growing the more often the world changes.
How solid is the evidence?
It is simulation in a deliberately simple environment: a binary decision tree, tabular learning rules, and reward relocated once between two 200-episode phases, averaged over tens of thousands of trials. The authors state plainly that scaling to richer settings is untested and the code was not yet released.
What does this say about how to train an organoid?
Pipe reward through silicon rather than into the tissue. Let the culture spend its limited plasticity learning reward-free structure of the input world, keep the objective in the controller, and use brief reward-paired plasticity windows, the hot-cache pattern, only when a new task arrives.
What could go wrong in a hybrid built this way?
The model-based module in the paper assigns untracked transitions probability zero, so a silicon controller that models the culture's knowledge will treat everything it has not measured as impossible, overestimating gaps and misplanning interventions. An explicit uncertainty model over untracked tissue state is the fix the paper's flaw implies.
References
- M. Farrell, T. Toyoizumi. Cortex and subcortex play distinct roles over learning when cortical memory is limited. arXiv:2606.00667. 2026. https://arxiv.org/abs/2606.00667. Accessed 2026-09-29.