PyTorch Distributed: Experiences on Accelerating Data Parallel Training

About

This paper presents the design, implementation, and evaluation of the PyTorch distributed data parallel module. PyTorch is a widely-adopted scientific computing package used in deep learning research and applications. Recent advances in deep learning argue for the value of large datasets and large models, which necessitates the ability to scale out model training to more computational resources. Data parallelism has emerged as a popular solution for distributed training thanks to its straightforward principle and broad applicability. In general, the technique of distributed data parallelism replicates the model on every computational resource to generate gradients independently and then communicates those gradients at each iteration to keep model replicas consistent. Despite the conceptual simplicity of the technique, the subtle dependencies between computation and communication make it non-trivial to optimize the distributed training efficiency. As of v1.5, PyTorch natively provides several techniques to accelerate distributed data parallel, including bucketing gradients, overlapping computation with communication, and skipping gradient synchronization. Evaluations show that, when configured appropriately, the PyTorch distributed data parallel module attains near-linear scalability using 256 GPUs.

Shen Li, Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith, Brian Vaughan, Pritam Damania, Soumith Chintala• 2020

Related benchmarks

Task	Dataset	Result
Visual Question Answering	TextVQA	--	1455
Physical Interaction Question Answering	PIQA	Accuracy79.3	462
Multi-discipline Multimodal Understanding	MMMU	--	422
Visual Question Answering	DocVQA	--	205
Commonsense Reasoning	HellaSwag	HellaSwag Score72.1	62
Visual Question Answering	ChartQA	Score21.6	32
Commonsense Reasoning	WinoGrande	Score64.2	22
Visual Question Answering	InfographicVQA	ANLS45.9	19
Language Modeling	WikiText-103 (fine-tuning)	Training Time (s)1.13e+4	11
Social Interaction Question Answering	SIQA	Normalized PLL Score49.8	10

Showing 10 of 23 rows

Other info

Follow for update

@wizwand_team Discord