Research analysis · Neural decoding

Computation on a few bits per token: a stored-program decoder for scarce neural data

Most neural decoding work chases scale. This preprint goes the other direction: a memory-augmented Transformer whose feed-forward weights are synthesized on the fly, token by token, from a small bank of low-rank instructions read off a slow latent state. On three public motor-cortex recordings the model beats a matched modern Transformer exactly where real neural data lives, at small budgets, and its program channel is measurable: the network runs on about 3 bits of instruction per token.

Source: The Von-Neumann State-Space Transformer for Neural Decoding, arXiv:2608.25088, 2026. Primary source. Read the full HTML of the preprint, including all tables and the language-modeling control experiment.

What the work claims

This is a methods-and-results preprint: a new architecture plus a scaling study on public benchmarks, no new animal data. The central claim is that for neural decoding, the right inductive bias is a stored-program style of computation. In a standard Transformer the feed-forward block applies one fixed operator to every token; here a controller reads a slowly varying low-dimensional state and uses it to synthesize the operator itself, per token, as a shared base matrix plus a small weighted sum of low-rank instructions.1

Three specific results support the claim. First, on three motor-cortex benchmarks the model, called VN-SST, beats a modern Transformer of matched depth and matched parameter window under a limited-data budget, most decisively on the scarcest dataset. Second, on that scarce dataset, lengthening the training context helps VN-SST and hurts the Transformer, a crossover the author argues comes from the model's carried state. Third, the program is a measurable control channel: given a bank of 32 instructions, the trained network uses only about 2.5 to 3.2 bits of code per token, well under the 5-bit ceiling, while accuracy stays flat as the bank grows. A control experiment on two small text corpora suggests the mechanism is not specific to neural data.

How it works

The benchmarks are the three Neural Latents Benchmark recordings hosted on DANDI: MC_RTT (primary motor cortex, random-target reaches, finger velocity, 130 units), MC_Maze (motor and premotor cortex, maze reaches, hand velocity, 182 units), and Area2_Bump (somatosensory area 2, perturbation responses, hand velocity, 65 units). Spikes are binned at 50 ms, smoothed, z-scored, and framed as a joint codec: from one shared hidden state the model must both continue the population spike train autoregressively and decode behavior through a linear head, with the behavioral term up-weighted tenfold in the loss.1 The recordings are the standard Neural Latents Benchmark testbed for latent-dynamics models of motor cortex.2

The architecture replaces both halves of a Transformer layer. Attention becomes three parallel carried pathways: causal-band local attention for short-range structure, a diagonal selective state-space model for slow dynamics, and a delta-rule fast-weight matrix memory that writes prediction errors and reads them back by content address. The feed-forward SwiGLU block becomes programmable: its weight matrix at token t is W(t) = W0 plus a per-token linear combination of K rank-r instruction matrices, where the combination coefficients, bounded by a tanh, are decoded from the token embedding together with an m-dimensional readout of the carried state-space memory. A slow latent trajectory thus acts as an instruction pointer: it selects which operator executes at each token. Default bank size is 8 instructions of rank 8 with an 8-dimensional manifold readout.

The protocol is careful in the ways that matter for small data. Depth is fixed at 4 layers and only width is scaled, with both architectures fitted into the same roughly 64K to 270K non-embedding parameter window; every point is a mean over 3 seeds; and the context-length sweep equalizes the number of optimizer steps across conditions so that longer context is not confounded with more training.

Where a skeptic should push

The most load-bearing assumption is that the gains come from programmability rather than from the three memory pathways, which any memory-augmented architecture could carry. The paper never ablates the instruction bank against the pathways alone, so the reader cannot tell how much of the scarce-data margin is the von Neumann mechanism and how much is ordinary recurrence and fast weights. That matters because the neuroscience-flavored claim, that low-dimensional latent dynamics route cortical computation, rides on the bank specifically.

Second, the evidence base is narrow: three recordings, all motor cortex and somatosensory cortex, all behavioral velocity decoding, one species, with single-trial spike prediction sitting near its noise floor for both models, so the comparison metric is entirely the behavioral readout. The gains also shrink with data, as the author honestly reports: at full recording the two models nearly converge, for example 0.61 vs 0.52 on MC_RTT at the largest budget, 0.85 vs 0.85 on MC_Maze. Third, this is a single-author industry preprint, and the code is described as available on request subject to internal review, which is not reproducibility. Treat the exact margins as provisional until the implementation is public and independently rerun. The language-modeling control (perplexity 49.4 vs 60.6 on tiny-Shakespeare, 45.1 vs 52.1 on WikiText-2 at the top of the ladder) suggests a generic mechanism, but small-text perplexity is a weak proxy for anything biological.

What this changes for organoid readout

The non-obvious implication is about what scarce-data decoding is, because scarce data is the permanent condition of organoid intelligence. A dish of living neural tissue gives you short recordings, few independent sessions, drift between days, and no possibility of the million-trial datasets that Vision-scale models enjoy. This paper demonstrates, on public neural data, that the winning strategy in exactly that regime is not a bigger decoder but a decoder with a carried low-dimensional state that makes longer context pay. That flips an experiment-design assumption the OI field inherits from deep learning: with a state-carrying decoder, longitudinal recordings of the same tissue become a first-class asset, because accuracy rises with accumulated context rather than saturating. An organoid substrate that stays alive and stable for weeks is worth measurably more to this class of readout than a sequence of one-shot preparations, and experiments should be designed to exploit that.

The control-bits diagnostic is the second gift. Because the operator is addressed rather than blended, the authors can count how much program the network actually uses: about 2.5 to 3.2 bits per token, three to six effective operators out of thirty-two, consistently across three brain areas. For organoid work this points at a quantitative answer to the diffuse question of whether a culture is computing. Fit a decoder with an explicit program channel to MEA recordings and the number of effective operators its latent state addresses becomes a statistic you can compare across preparations, developmental ages, and training interventions, with a memoryless linear decoder as the floor, exactly the control this paper uses. Zero effective operators beyond the floor is evidence against structured computation, and it is a result a negative-results-friendly field badly needs.

The genuine threat cuts the other way: the decoder can substitute for the substrate. A sufficiently strong prior model will extract behavior-level structure from weak or even generic neural activity, and as sample-efficient silicon decoders improve, an unremarkable organoid plus a very good decoder will look like a computing organoid. The safeguard is already in the paper's design and mostly ignored in the field: always report the memoryless linear reference, and demand that any claimed substrate capability survive decoder ablation, meaning the gain must shrink when the programmable channel and carried state are removed. The second threat is quieter: the paper's advantage lives at small data and vanishes at large data, so results demonstrated on organoid-scale data cannot be assumed to scale to data-rich regimes, and marketing them as such would be a category error.

The bottom line

Established, within the limits of a non-public implementation: on three public motor-cortex benchmarks, a per-token programmable feed-forward operator with carried state beats a matched Transformer under scarce data and matched budgets (for example decode R-squared 0.351 vs 0.207 on the scarcest dataset), turns longer context into rising rather than falling accuracy there, and compresses a 32-instruction bank to roughly 3 bits of program per token while accuracy stays flat in bank size. Not established: that biological cortex implements anything like fetch-decode-execute; the von Neumann framing is an analogy, and the memory pathways rather than the instruction bank may carry the gains. What would confirm the mechanism: an ablation isolating the bank from the pathways, public code, and replication on recordings from other brain areas. What would weaken it: a plain state-space or fast-weight model without the bank matching the full result. For organoid intelligence the durable residue is twofold: design longitudinal readout experiments because context now compounds, and adopt the effective-operators count as a standard audit statistic for claims of computation in living tissue.

Frequently asked questions

What is VN-SST in one sentence?

A Transformer variant whose feed-forward weight matrix is synthesized at every token from a shared base operator plus a small set of learned low-rank instructions, with the per-token selection code read from a carried low-dimensional state.

How big are the decoding gains?

Under a limited-data budget, peak behavioral decode R-squared is 0.351 vs 0.207 for a matched Transformer on MC_RTT, 0.716 vs 0.655 on MC_Maze, and 0.696 vs 0.632 on Area2_Bump, all 3-seed means. The margins narrow as training data grows and nearly vanish at full recording.

What does it mean that longer context helps?

On the scarce MC_RTT dataset the Transformer's decode declines as training context grows from 16 to 128 tokens while VN-SST's rises from 0.60 to 0.68. The author attributes this to the carried state-space and fast-weight memories, which integrate long-range structure without widening the local attention window.

What are the control bits?

A diagnostic measuring how much of the instruction bank the trained network actually uses, via the entropy of the per-token code distribution. Given 32 instructions the network uses about 2.5 to 3.2 bits per token, three to six effective operators, while decode accuracy stays essentially flat as the bank grows from 1 to 32.

Is this specific to neural recordings?

Probably not, based on the control experiment: the same architecture beats a matched Transformer on tiny-Shakespeare and WikiText-2 language modeling, reaching the same perplexity with roughly 2 to 3 times fewer parameters. But small-corpus perplexity is a weak stand-in for biological relevance.

What is the biggest caveat?

Single author, industry preprint, code available only on request subject to internal review, three motor-related recordings in one species, and no ablation separating the instruction bank from the carried-memory pathways that any modern sequence model could carry. The direction is promising; the exact numbers await independent reproduction.

References

  1. M. Sarafyazd. The Von-Neumann State-Space Transformer for Neural Decoding. arXiv:2608.25088 [cs.LG]. 2026. https://arxiv.org/abs/2608.25088. Accessed 2026-10-09.
  2. F. Pei et al. Neural Latents Benchmark '21: Evaluating latent variable models of neural population activity. NeurIPS Datasets and Benchmarks Track. 2021. https://arxiv.org/abs/2109.04463. Accessed 2026-10-09.