Loading...
Thumbnail Image
Item

Byzantine-Robust Optimization under (L0, L1 )-Smoothness

Bolatov, Arman
Horvath, Samuel
Takac, Martin
Gorbunov, Eduard
Citations
Altmetric:
Supervisor
Department
Machine Learning
Embargo End Date
Type
Conference proceeding
Date
License
http://rightsstatements.org/page/InC/1.0/
Language
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
We consider distributed optimization under Byzantine attacks in the presence of (L<inf>0</inf>, L<inf>1</inf>)-smoothness, a generalization of standard L-smoothness that captures functions with state-dependent gradient Lipschitz constants. We propose Byz-NSGDM<sup>1</sup>, a normalized stochastic gradient descent method with momentum that achieves robustness against Byzantine workers while maintaining convergence guarantees. Our algorithm combines momentum normalization with Byzantine-robust aggregation enhanced by Nearest Neighbor Mixing (NNM) to handle both the challenges posed by (L<inf>0</inf>, L<inf>1</inf> )-smoothness and Byzantine adversaries. We prove that Byz-NSGDM achieves a convergence rate of O(K<sup>−1/4</sup> ) up to a Byzantine bias floor proportional to the robustness coefficient and gradient heterogeneity. Experimental validation on heterogeneous MNIST classification, synthetic (L<inf>0</inf>, L<inf>1</inf> )-smooth optimization, and character-level language modeling with a small GPT model demonstrates the effectiveness of our approach against various Byzantine attack strategies. An ablation study further shows that Byz-NSGDM is robust across a wide range of momentum and learning rate choices.
Citation
A. Bolatov, S. Horvath, M. Takac, E. Gorbunov, "Byzantine-Robust Optimization under (L0, L1 )-Smoothness," 2026, pp. 826-854.
Source
Proceedings of Machine Learning Research
Conference
Third Conference on Parsimony and Learning (CPAL 2026)
Keywords
Subjects
Source
Third Conference on Parsimony and Learning (CPAL 2026)
Publisher
ML Research Press
DOI
Full-text link