A spiking agent that keeps its learned self after the buffer is pulled
A single-author preprint reports that a minimal spiking agent develops a durable, self-shaped behaviour, and that the durability appears only when a narrow, agency-gated credit signal is allowed to do slow structural work. Strip that channel out and the behaviour evaporates the moment the memory buffer is removed. For the organoid-intelligence field, the interesting part is the test it implies, not the agent it builds.
Source: From Detecting Agency to Doing Work: Self-Caused Credit Builds a Durable Behavioral Self in a Minimal Spiking Agent, arXiv preprint (cs.AI), June 2026. Primary source. Read: full text of the version-1 PDF, including the ablation tables and the continual-learning experiment.
What the work claims
The paper asks a developmental question in engineering terms: what turns a system that can merely tell self from world into one that is durably shaped by that distinction. Its answer is a specific credit-assignment mechanism.1 The author introduces what he calls agency-gated slow credit: a conjunctive term, written Own times Agency times Salience, that drives a slow update to the network's parameters. In plain terms, a parameter changes durably only when three conditions coincide at once, that the outcome was the agent's own, that the agent had agency over it, and that it mattered. Because the term is multiplicative rather than additive, any one factor going to zero vetoes the update.
The headline result is a dissociation. In a spiking model built from leaky integrate-and-fire neurons trained with a biologically flavoured local rule, a learned self-preserving choice survives removal of the agent's episodic memory buffer, with a reported retained fraction of 0.96 across fifty runs. The same behaviour collapses to zero when the slow decoders are reset or the agency gate is removed. Holding the fast reward pathway matched, durable behaviour develops only when the self-credit channel is permitted to do slow work: post-unload self-preservation is reported as 1.00 with the channel on and 0.00 with it off, and a harder twenty-four dimensional partially-observed control task shows the same split, 0.74 against 0.00. The author is careful to disclaim any assertion of consciousness, and frames the retained behaviour as an operational behavioural self, a residue you can measure rather than a mental state you must infer.
How it works
The design separates two timescales. A fast pathway handles moment-to-moment action selection and reward, the kind of in-context adjustment that guides behaviour only while the relevant signal is present. A slow pathway changes the decoders, the learned weights that map neural population activity onto outputs, and it is this slow change that can persist. The agency gate decides when the slow pathway is allowed to write. The author reproduces an agency comparator from prior work, a circuit that estimates whether an observed change was self-caused, and then toggles only the slow-credit channel, so the comparison isolates the effect of self-credit doing structural work from the effect of simply detecting agency.
The most conceptually useful contribution is a bookkeeping identity: the paper reports a plastic-work analysis in which the deformation of the behavioural attractor's basin equals the net self-credit work done. An attractor basin here is the set of states that flow toward a stable behaviour, and deepening that basin is what makes a choice durable rather than fleeting. Casting durability as accumulated work gives a physical-style ledger for something usually described only qualitatively. The claim is then stress-tested in a continual-learning setting: across eight tasks learned in sequence under exogenous interference, the multiplicative veto retains old tasks with a final post-unload accuracy of 0.88 and forgetting of 0.13, while an additive pooling of the same signals collapses to chance, an ablation with no agency gate falls below chance, and episodic or replay baselines sit near chance once the buffer is unloaded. All of this is reported with no replay buffer and no mechanism that depends on knowing where one task ends and the next begins.
Where a skeptic should push
This is a simulation and a position paper wearing the clothes of an empirical result, and it should be weighed as such. The single most load-bearing assumption is that the agency comparator, the circuit that decides an outcome was self-caused, is available and correct. In this work that comparator is engineered and reproduced from a cited construction; it is not learned from scratch and it is not derived from a biophysical substrate. If the gate that decides agency is itself the hard problem, then the paper has shown that a correct gate plus slow credit yields durability, which is a conditional result, not a demonstration that durability emerges from raw dynamics.
Second, the numbers are clean to the point of being suspicious for a stochastic learning system: post-unload scores of exactly 1.00 against exactly 0.00 suggest a task engineered so the dissociation is near-binary, which is legitimate for isolating a mechanism but tells us little about graded, noisy, real-world learning. The twenty-four dimensional result at 0.74 is the more informative one precisely because it is not saturated. Third, the framing leans on loaded vocabulary, self, agency, durable self, that does real persuasive work while the underlying quantities are decoder retention and basin depth. The author's explicit refusal to claim consciousness is welcome, but a reader still has to keep translating the vocabulary back into the measured variables. Finally, this is a single-author preprint with no independent replication; the fifty-run sample sizes bound within-study variance but say nothing about whether the effect survives a different substrate or a different comparator.
The learning signal a trained culture still lacks
Closed-loop training of living neural cultures, the paradigm popularised by embodied dish-based agents that play a simulated game under a free-energy-style feedback rule, delivers a global, ungated signal.2 Predictable stimulation follows good actions and unpredictable stimulation follows bad ones, and the whole culture receives it. This preprint, read against that paradigm, makes an uncomfortable prediction. If durable, self-shaped behaviour requires a slow-credit channel that fires only on self-caused, salient outcomes, then an undifferentiated feedback signal matches the paper's ungated conditions, whether one reads it as the additive pooling that collapses to chance or the missing agency gate that falls below chance. Either mapping loses the behaviour after unload, so the conclusion is robust to which one is right: the mechanism that would make a dish's learning persist is the one current rigs do not implement.
The non-obvious implication is a test, and it is cheap. The paper's central move is the unload: train, then remove the signal and look for residue. Organoid and neuronal-culture training almost never reports this. Learning is typically scored while the closed loop is running, which cannot distinguish a durable change in the tissue from an in-context adjustment that will vanish when stimulation stops. Borrowing the post-unload probe would let the field separate the two, and it would do so without any new hardware. That is the genuine opportunity: a design principle, gate slow plasticity on self-caused salient events, plus a falsification protocol, remove the loop and measure what remains.
The genuine threat is obsolescence of a claim rather than of a platform. Much of the excitement around biological learning rests on the intuition that living tissue adapts by construction. This paper, grounded in a spiking model rather than vibes, argues that adaptation which survives is not a generic property of a plastic substrate; it is the product of a specific conjunctive gate doing specific structural work. A culture has abundant plasticity, but there is no demonstrated biological mechanism that implements an Own times Agency times Salience veto, and the paper's own agency comparator is hand-built. So the honest reading cuts against tissue as much as for it: the result hands the field a target to engineer toward and a reason to doubt that the target is already met. The dual-use and ethics dimension is quieter here than in some work, but note one thing, an operational behavioural self is defined behaviourally and explicitly severed from any consciousness claim, which is the right move and also a reminder that a durability metric is not a welfare metric.
The bottom line
Treat this as a hypothesis with a clean in-silico demonstration, not an established fact about living tissue. The established part is narrow and internal: in this spiking model, on these tasks, durable post-unload behaviour tracks an agency-gated slow-credit channel and disappears without it, and a basin-work identity ties durability to accumulated self-credit. The speculative part is everything about biology: nothing here shows that a culture implements such a gate, and the comparator that makes the whole thing work is engineered. What would confirm the transfer is a wet demonstration that contingency-gated plasticity produces measurable post-unload residue in a culture where ungated feedback does not. What would break the borrowed lesson is a showing that ordinary, ungated closed-loop training already yields durable residue under a proper unload test, which would mean the gate is unnecessary in tissue. Either way, the reusable gift is the unload protocol, and any group claiming a dish has learned should be running it.
Frequently asked questions
What does the phrase agency-gated slow credit actually mean?
It is a rule for when a network is allowed to change its durable weights. The change is permitted only when an outcome was the agent's own, was under its control, and was salient, and because the three factors are multiplied, any one of them being absent blocks the update entirely.
Why does removing the memory buffer matter so much here?
Removing the buffer, which the paper calls the unload, tests whether learning left a structural trace or was only being held in a temporary store. A behaviour that survives the unload has changed the network itself; one that vanishes was merely context that the buffer was supplying.
Is this a result in living tissue?
No. It is a simulation using spiking leaky integrate-and-fire neurons and a local learning rule. The relevance to living cultures is inferential, and the paper's engineered agency comparator has no demonstrated biological counterpart.
How could a wet-lab group use this tomorrow?
By adopting the unload probe. Train a culture in a closed loop, then remove the feedback and measure whether the trained behaviour persists, which separates durable plasticity from adjustments that only exist while the loop is running.
Does the paper claim the agent is conscious or sentient?
It explicitly does not. The author defines an operational behavioural self as a measurable residue in behaviour and separates that quantity from any claim about experience, which is the correct and conservative framing.
References
- Han H. From Detecting Agency to Doing Work: Self-Caused Credit Builds a Durable Behavioral Self in a Minimal Spiking Agent. arXiv preprint. 2026. arXiv:2606.30191. Accessed 2026-07-29.
- Kagan BJ, Kitchen AC, Tran NT, et al. In vitro neurons learn and exhibit sentience when embodied in a simulated game-world. Neuron. 2022. doi:10.1016/j.neuron.2022.09.001. Accessed 2026-07-29.