TL;DR
Get school and study supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Q Labs Research reports that Dust, a zeroth-order training method, can pretrain transformer language models without a backward pass. Its experiments found closer agreement with backpropagation as the method’s population grew, but the researchers say this required substantially more compute, and the reported efficiency comparisons with another method rely on extrapolations.
Q Labs Research says it has tested Dust, a method for pretraining transformer language models without backpropagation, and found that its results can approach those of backpropagation when given a larger perturbation population. The October 2026 report describes an alternative way to assign learning signals, but the authors say its stronger results come with substantially more compute.
Dust is a zeroth-order optimization method: rather than calculating backpropagation gradients, it perturbs a model’s activations and uses changes in loss to estimate an update direction. The report says these perturbations are applied independently at each token. The authors call each token a “virtual population member,” allowing one forward pass to evaluate many perturbations in parallel without separately materializing a population of models.
In its experiments, Q Labs reports that Dust’s gradient estimates align more closely with backpropagation as population size increases. The researchers say this alignment remained strong across the scales they tested, up to 1 billion tokens. They also report that a 243-million-parameter model outperformed a model 120 times smaller at most tested population sizes. These are results reported by the authors; the supplied material does not provide independent replication.
The report compares Dust with EGGROLL, an evolution-strategy method that perturbs weights. Q Labs estimates that, from 1 million tokens onward, Dust is roughly 1,000 to 10,000 times more efficient in its transformer implementation comparison. The authors explicitly describe this as an extrapolation, so it should not be read as a measured advantage across all hardware, workloads or training budgets. They also say Dust can approximate backprop closely at large populations and exceeds it in some settings.
A Different Route to Language Model Training
If the reported results hold up, Dust could broaden the range of learning methods researchers can test on transformer models. Backpropagation depends on differentiable operations and provides a direct gradient calculation; Dust instead uses reward-weighted activation perturbations to estimate how changes affect loss. That makes the study relevant to researchers exploring whether more computation can substitute for some analytic structure in model training.
The practical trade-off remains central. Q Labs says Dust approaches backprop as population size grows, which also means spending more compute on the search. The report’s results do not establish that Dust is cheaper or more effective for a typical production training run. Its comparison with EGGROLL is an extrapolation, and its claims about performance beyond tested settings remain a research proposition.
high performance GPU for AI training
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Dust Differs From Weight Search
Evolution strategies typically explore parameter space by perturbing model weights and evaluating the resulting candidates. Q Labs says this approach becomes costly as the population grows because each candidate must be represented and evaluated. Dust takes a different route: it perturbs activations rather than weights, and uses tokens within a forward pass as its virtual population.
The report frames the work against a longstanding assumption that zeroth-order methods do not scale well to large networks. It argues that, in its tests, larger models were more population-efficient than smaller ones. That finding challenges the assumption in the settings studied, but does not by itself show that the method will scale across architectures, data or training budgets that were not tested.
“Dust approximates backprop closely at large population (i.e. substantially more compute) and in multiple settings even exceeds it.”
— Q Labs Research, in its report’s TL;DR
large memory server for machine learning
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Evidence Beyond the Reported Tests
The supplied report material does not establish whether Dust’s results have been independently reproduced, or how performance changes across other datasets, model designs and hardware. It also does not provide enough detail here to assess the absolute compute cost of matching backpropagation in a practical training run. The EGGROLL efficiency figures are explicitly extrapolated, and the authors’ suggestion that compute-rich training might surpass backprop remains a possibility they raise, not a demonstrated outcome across general workloads.
As an affiliate, we earn on qualifying purchases.
Replication and Larger Training Runs
The next evidence to watch for is whether other researchers can reproduce the reported alignment and training results, and whether Dust remains competitive when evaluated under clearly matched compute budgets. Further experiments could also show how the method behaves across larger models, longer training runs and different data. Q Labs’ report presents initial research findings; the available source material does not state a follow-up schedule or identify a planned release date for further results.
tensor processing unit (TPU) for deep learning
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Dust?
Dust is a zeroth-order training method that perturbs model activations and uses changes in loss to estimate updates, rather than calculating gradients through backpropagation.
Did Dust outperform backpropagation?
Q Labs reports that Dust approached backpropagation at larger population sizes and exceeded it in some experimental settings. The report also says those larger populations required substantially more compute; the supplied material does not establish a general advantage.
How does Dust use tokens as a population?
The method applies perturbations independently at each token. Q Labs describes each token as a virtual population member, evaluated together in one forward pass.
How firm is the efficiency comparison with EGGROLL?
The report estimates a 1,000-to-10,000-fold efficiency difference from 1 million tokens onward, but says the comparison is based on extrapolations. It is not presented as a measured result for every workload or hardware setup.
Source: hn
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
