Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Auto-FL-Research: Agentic Search for Federated Learning Algorithms

About

Federated learning (FL) research often depends on many small but consequential algorithmic choices: optimizer variants, server aggregation rules, local training schedules, normalization, regularization, and model architecture. These choices are expensive to explore manually and difficult to compare fairly when candidate changes can also alter the FL training or evaluation path. In this work, we present Auto-FL-Research (AFR), a constrained coding-agent workflow for FL algorithmic recipe search. Agents may propose and implement candidate training algorithms, including server aggregation rules, client update schedules, local objectives, and registered model variants, while task profiles fix the mutation surface, compute budget, communication contract, and final model evaluation. Each campaign records candidate scores, runtime, edited files, artifacts, and failure status. We evaluate AFR on five healthcare cross-silo FLamby tasks and on grouped-client profiles for the five fixed LEAF datasets plus the LEAF synthetic task. Five-seed repeat evaluations support gains on four FLamby tasks and five of six LEAF profiles, while also exposing seed-sensitive and search-selected failure cases. Same-budget controls show that several gains correspond to FL-recipe changes, whereas other improvements are recovered by fixed-surface scalar controls or fail under repeat or held-out evaluation. These mixed outcomes are part of the contribution: they show how agent-generated candidates can be separated into repeated FL mechanisms, fixed-surface tuning effects, and selected single-run artifacts.

Holger R. Roth, Ziyue Xu, Chester Chen, Daguang Xu, Peter Cnudde, Andrew Feng• 2026

Related benchmarks

TaskDatasetResultRank
Brain Image SegmentationFLamby IXI (repeat)
Dice Score98.95
3
ClassificationFLamby Heart Disease (repeat)
Accuracy79.4
3
Whole Slide Image classificationFLamby Camelyon16 (repeat)
ROC AUC0.7494
3
Skin lesion classificationFLamby ISIC2019 (repeat)
Balanced Acc.64
3
Survival AnalysisFLamby TCGA-BRCA (repeat)
C-index0.808
3
ClassificationLEAF Synthetic
Accuracy98.9
2
Image ClassificationLEAF FEMNIST
Accuracy87.3
2
Next-Character PredictionLEAF Shakespeare
Next-character Accuracy57.5
2
Next-token predictionLEAF Reddit
Next-token Accuracy15.6
2
Sentiment AnalysisLEAF Sent140
Accuracy74.9
2
Showing 10 of 11 rows

Other info

Follow for update