Dust, a new zeroth-order method, competes with backpropagation in pretraining transformer language models, according to research published on qlabs.sh. The approach perturbs activations independently at every token, allowing parallel evaluation of a virtual population in a single forward pass. Dust’s efficiency and scalability were demonstrated on models up to 243 million parameters and datasets exceeding one billion tokens.
The Dust method works by applying node perturbation at each token, treating each as a member of a virtual population. This allows it to approximate backpropagation gradients closely, especially as the population size grows. The research shows Dust is between 1,000 and 10,000 times more efficient than EGGROLL, a state-of-the-art evolutionary strategy method implemented on transformers, when processing over one million tokens. Larger models showed better population efficiency, challenging the belief that zeroth-order methods do not scale well.
This development is significant because backpropagation has been the foundational algorithm for training deep neural networks, including transformers, due to its ability to compute first-order gradients. Zeroth-order methods like Dust do not require differentiability and have traditionally been seen as less scalable. Dust’s performance suggests that in compute-rich environments, it could match or even surpass backpropagation, potentially influencing future training paradigms for large language models.
The research highlights that Dust’s gradient estimates maintain strong alignment with backpropagation across all tested scales, up to one billion tokens. This finding supports the method’s viability for large-scale transformer training and opens avenues for further exploration of zeroth-order optimization techniques in deep learning, as detailed on qlabs.sh.