Notes on distributed training, GPUs, and learning in public. More on the about page.
-
Inside NCCL's all-reduce: ring, double binary tree, or neither?
In the last post I said the standard ring all-reduce is literally a reduce-scatter and an all-gather run back to back. That’s true, and it’s also where most explanations stop, mine included. It left me with a picture of one algorithm, th...
-
FSDP collectives 101: why reduce-scatter, and why not broadcast?
If you learned distributed training through DDP, you probably carry two instincts: after the backward pass, all-reduce the gradients; and if only one place has the freshest weights, broadcast them out. I carried both into FSDP and they c...