Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Model Parallelism With Subnetwork Data Parallelism

About

Pre-training large neural networks at scale imposes heavy memory demands on accelerators and often requires costly communication. We introduce Subnetwork Data Parallelism (SDP), a distributed training framework that partitions a model into structured subnetworks trained across workers without exchanging activations. We study two complementary masking regimes: backward masking, which applies sparsity only in the backward step to retain unbiased gradients, and forward masking, which also removes parameters in the forward pass to deliver stronger efficiency gains while providing additional regularization. We further explore two subnetwork construction strategies: neuron level and block level, applied across both transformers and CNNs. In experiments spanning 1B LLaMA pre-training on FineWeb to ResNet-18 on CIFAR, SDP reduces per device memory usage by 28%-60% while maintaining or improving performance under FLOP-matched settings.

Vaibhav Singh, Zafir Khalid, Pietro Cagnasso, Edouard Oyallon, Eugene Belilovsky• 2025

Related benchmarks

TaskDatasetResultRank
Language ModelingLLaMA 1B (val)
Validation Loss2.434
4
Downstream evaluationlm-evaluation-harness ARC-E, BoolQ, HellaSwag, OBQA, SciQ
ARC-E Accuracy0.381
4
Showing 2 of 2 rows

Other info

Follow for update