Research Projects

Project I: Zeroth-Order Optimization

Understanding Learning Dynamics of Zeroth-Order Optimiztaion

Loading calculation...

Despite the growing success of zeroth-order (ZO) optimization in large-scale learning, a fundamental question remains: why can ZO work effectively for models with billions of parameters? Classical convergence analyses often predict a strong dependence on parameter dimension, but provide limited insight into how ZO actually learns. In our ZO Learning Dynamics [ICML 2026], we study ZO optimization from a kernel perspective and show that its learning dynamics are governed by a randomly projected empirical Neural Tangent Kernel. Leveraging the Johnson-Lindenstrauss lemma, we show that this random projection can preserve the first-order kernel structure with sufficiently many perturbations, and that the approximation quality depends primarily on the number of perturbations and the output dimension, rather than the massive parameter dimension. This provides a new explanation for why ZO optimization can remain effective even for highly overparameterized models.

Zeroth-Order Optimization for Federated LLM Fine-Tuning

Scaling federated learning (FL) to LLMs is fundamentally constrained by the cost of communicating high-dimensional model updates. To achieve dimension-free communication in FL, we first address this challenge with DeComFL [ICLR 2025], which exploits the compact structure of ZO updates: each gradient scalar paired with a random seed used to generate and reconstruct the corresponding perturbation vector. Building on this principle, we further develop HiSo [ICLR 2026] to overcome the slow convergence of conventional ZO methods through Hessian-informed optimization, and ZO-SignSGD [Asilomar 2025] to improve communication efficiency and Byzantine robustness through sign-based updates. Together, these methods establish ZO optimization as a new communication primitive for scalable, efficient, and robust FL.

Zeroth-Order Optimization for Resource-Constrained Training

Training LLMs across resource-constrained devices is fundamentally limited by memory, communication, and system heterogeneity. We first address these challenges with MU-SplitFed [NeurIPS 2025], which combines ZO optimization with unbalanced server-client updates to reduce communication rounds and mitigate straggler-induced latency in split federated learning. Building on the potential of ZO for resource-constrained training, we develop HOSL [WiOpt 2026], a hybrid-order framework that combines memory-efficient ZO optimization on constrained clients with first-order optimization on powerful servers. We further extend ZO to model-parallel LLM fine-tuning with SparQ [TMLR 2026], which integrates ZO optimization with activation sparsity to jointly reduce memory and communication overhead. Together, these methods demonstrate how ZO optimization can enable scalable training across heterogeneous and resource-constrained computing environments.

Project II: Federated Learning

Federated Learning under Arbitrary Client Participation

Loading focus...

FL in practice often operates with only a subset of clients participating in each communication round. When client availability is heterogeneous and unpredictable, such arbitrary client participation can introduce persistent optimization bias and prevent conventional FL methods from converging to the desired global solution. We address this challenge with FAST [IJCAI 2025], a lightweight mechanism for FL under arbitrary client participation. We further develop FOCUS [NeurIPS 2025], which reinterprets FL through the lens of decentralized optimization over time-varying graphs. By modeling client participation, local updates, and aggregation as stochastic matrix operations, FOCUS introduces a push-pull strategy that transforms participation-induced bias into delayed gradient information. This perspective enables exact linear convergence under arbitrary client participation without knowing or estimating client participation probabilities, and provides a principled connection between federated and decentralized optimization.

Communication-Efficient Federated Learning

See Project I: Zeroth-Order Optimization for Federated LLM Fine-Tuning.

Project III: Understanding Learning and Optimization Mechanisms.

Loading sl...

This project is driven by a fundamental question: WHY do certain learning/optimization approaches work? First, we ask why zeroth-order optimization, despite its classical dimension-dependent limitations, can effectively fine-tune LLMs with billions of parameters. We answer this question by connecting ZO learning dynamics to first-order learning through random projection theory, showing that the approximation quality can depend primarily on the number of perturbations and output dimension rather than the massive parameter dimension [ICML 2026]. Second, we ask why subliminal learning works: how can a student acquire task knowledge from seemingly unrelated outputs without explicit task supervision? Using the empirical Neural Tangent Kernel, we characterize how learning signals propagate from ghost-output supervision to task predictions, revealing the mechanism behind this unexpected knowledge transfer and the key conditions that determine when it occurs [Arxiv]. Together, these works seek to explain not only whether a learning approach works, but fundamentally why it works.