Solving for network weights instead of descending toward them
A physics-informed neural network that never runs backpropagation matches or beats deep gradient-trained networks on nonlinear differential equations while training one to two orders of magnitude faster. For everyone trying to train computation into physical or living substrates, that is a direct challenge to the assumption that iterative gradient descent is the price of doing business.
Source: Physics-Informed Stochastic Configuration Machine: A Backpropagation-Free Neural Network with Fast Training for Nonlinear Differential Equations, Song, Chen, Cen, Wu, and Zhang, arXiv:2608.26549, August 2026. Primary source. Read: full text PDF.
What the work claims
This is a methods and theory paper. Physics-informed neural networks (PINNs) embed a differential equation into the training loss so a network can solve or fit the equation from sparse data; the standard approach trains them with backpropagation and gradient descent, which is slow, memory-hungry, and prone to non-convex failure modes.2 The authors' Physics-Informed Stochastic Configuration Machine (PI-SCM) removes gradient descent from the loop entirely: it analytically linearizes the physics loss around the current solution and then solves for the network weights in closed form with a sequence of generalized linear least squares problems.1
The claimed results are strong. Across forward problems (the Van der Pol oscillator, the two-dimensional Helmholtz equation, the Allen-Cahn equation) and inverse problems that jointly reconstruct state and identify physical parameters, PI-SCM achieves competitive or lower error than deep gradient-trained PINNs while cutting training time by one to two orders of magnitude. Each configuration is repeated 100 times. The authors also prove universal approximation properties for their three algorithmic variants, named localized construction, sliding-window updating, and global updating.1
The boldness is in the framing: iterative optimization is presented not as a fundamental cost of training but as an artifact of refusing to linearize, one that can be engineered away.
How it works
The building blocks are stochastic configuration machines: networks whose hidden nodes are assigned random parameters and whose only trainable quantities are the output weights. That design normally still needs gradient descent to fit those output weights against a nonlinear physics loss. The paper's core move is to remove that last gradient step too.
The trick is a first-order Taylor expansion of the differential operator. At each stage of a progressive construction, the authors evaluate the local Jacobians of the physics residual with respect to every derivative component of the solution, up to the operator's order. The nonlinear residual is then approximated around the previously established state as a linear function of the new weights, which projects the whole physics loss into a linearized algebraic subspace. In that subspace the optimal weights are not searched for; they are the explicit solution of a generalized linear least squares solve. A regularization parameter is chosen by the L-curve method, and solving a sequence of these linearized problems, each anchored at the last, recovers a full nonlinear fit. The sliding-window and global variants update and re-solve rather than rebuild, trading construction cost against solution fidelity.1
Dynamically, each iteration is a Gauss-Newton-like step with an analytic Jacobian: linearize, solve exactly, re-linearize. Gradient descent with small steps is replaced by exact solves in a moving linear subspace, which is where the speed comes from.
Where a skeptic should push
The most load-bearing assumption is that the differential operator is known and smooth enough to differentiate analytically. Every result in the paper rests on having an explicit F to linearize: the Van der Pol oscillator, Helmholtz, Allen-Cahn. That is exactly the setting where PINNs are already comfortable, and it is not the setting that matters most. The authors are commendably explicit about the boundaries: future work must handle higher-dimensional geometries, more severe multiscale stiffness, and noisy parameter identification, none of which appear here. The inverse problems are run on clean synthetic data; the moment parameter identification must cope with real measurement noise, the linearized subspace inherits that noise and the closed-form solve can amplify it.
Second, part of the speedup is structural rather than conceptual. A random-feature network with solved output weights is shallow compared with a deep PINN; comparing training times without normalizing for model capacity overstates the gain. A fairer claim is the accuracy-to-cost trade-off, and on that the paper does deliver: a competing extreme learning approach is faster still on one Helmholtz benchmark, but its RMSE sits more than three orders of magnitude above the global variant, and the authors report that plainly.1
Demonstrated: on low-dimensional, operator-known, noise-free benchmarks, sequential linearized least squares trains physics-informed models one to two orders of magnitude faster than gradient descent with equal or better accuracy. Asserted: that this class of methods scales to the stiff, high-dimensional, noisy regimes where it would actually matter.
Backprop-free training meets living substrates
Organoid intelligence and biological computing rest on a training asymmetry: you cannot backpropagate through living tissue, so the field invests in local plasticity rules, error-feedback circuits, and in-situ perturbation methods, while justifying the effort partly by the claim that iterative training is intrinsically expensive. This paper attacks the second half of that syllogism on digital hardware. If a useful class of models can be trained by exact linear solves in seconds, one of the economic arguments for wetware, that training is costly and biology does it cheaply, weakens. The case for computing on living tissue will have to rest on what it does at inference and at steady state: self-repair, adaptation to unmodeled chemistry, energy per operation in the femtojoule range, and computation that is its own sensor. None of those is touched by a faster least squares solver, but the lazy version of the training argument is now demonstrably lazy.
The opportunity is a design spec. The PI-SCM step is, mechanistically, a sequential linearization around a measured operating point: evaluate local Jacobians, project the error into a linear subspace, solve for the correction exactly. Any training scheme for a physical or biological substrate that can measure its own local response to perturbations can in principle implement an analog of this step, with perturbation playing the role of the Jacobian and local plasticity the role of the weight solve. Methods like equilibrium propagation already gesture in this direction; this paper supplies a digital benchmark for what such physics-coupled training must match to be worth its fragility.
There is also a near-term practical consequence. Fitting dynamical models to tissue recordings, mean-field reductions of organoid activity, closed-loop stimulation policies, all involve repeated solves of physics-constrained fits. A one to two order of magnitude cut in that inner loop changes which real-time experiments are feasible on commodity hardware, and it costs nothing biological.
The genuine threat runs the other way too: if the hard part of biological computing is recast as repeatedly solving well-posed inverse problems against known kinetics, then hybrid systems that keep the tissue as a sensor and hand the training mathematics to silicon become more attractive than fully wet training. The paper is about differential equations, but its implication for the field is a division of labor it never mentions.
The bottom line
Established: for nonlinear differential equations with known operators, replacing gradient descent with sequential analytic linearization and least squares solves trains physics-informed networks one to two orders of magnitude faster at competitive or better accuracy, with a proof of approximation power. Hypothesis: the same division, measure local Jacobians, project, solve, can guide how we train or calibrate biological substrates, and it recalibrates what living hardware must be better at than linear algebra. What would confirm the transfer: a physical or hybrid implementation approximating the linearized-solve step against measured tissue responses, beating digital baselines on a real closed-loop task. What would break it: the method stalling on stiff, high-dimensional, noisy systems, which the authors themselves list as untested.
Frequently asked questions
What does backpropagation-free training actually mean here?
No gradients are ever computed. The physics loss is linearized analytically around the current solution using local Jacobians of the differential operator, and the network weights fall out as the explicit solution of a linear least squares problem, repeated as the solution improves.
How large is the measured speedup?
The paper reports training one to two orders of magnitude faster than deep gradient-based physics-informed networks across benchmarks including the Van der Pol oscillator, two-dimensional Helmholtz, and Allen-Cahn equations, with each configuration repeated 100 times and competitive or lower error.
Why does this matter for organoid intelligence?
Training living tissue is the field's central bottleneck and a core justification for the approach. Showing that a broad class of models can be trained exactly and quickly on digital hardware forces the biological case onto inference-time properties like self-repair and energy efficiency, and provides a concrete mathematical template for what physical learning rules must approximate.
What are the method's limits?
It needs an explicit, differentiable differential operator, so it applies where the governing physics is known. The benchmarks are low-dimensional and noise-free; higher-dimensional geometries, multiscale stiffness, and noisy parameter identification are listed by the authors as future work, and real measurement noise is where closed-form solves can misbehave.
Could a biological substrate implement this kind of training step?
Mechanically, the step is measure local responses, project the error into a linear subspace, apply the exact correction. A substrate that can probe its own response to perturbations and adjust accordingly implements an analog of it, with perturbation in place of the Jacobian. Whether any living tissue can do this accurately enough to beat silicon is an open experimental question.
References
- Song Y, Chen Z, Cen L, Wu L, Zhang K. Physics-Informed Stochastic Configuration Machine: A Backpropagation-Free Neural Network with Fast Training for Nonlinear Differential Equations. arXiv:2608.26549. 2026. https://arxiv.org/abs/2608.26549. Accessed 2026-10-03.
- Raissi M, Perdikaris P, Karniadakis GE. Physics-informed neural networks. Journal of Computational Physics. 2019;378:686-707. Background on the gradient-trained baseline the primary source improves upon.