xtrain

gahow/xtrain

Fork 0

Commit Graph

Author	SHA1	Message	Date
Gahow Wang	0131f05b26	docs: Phase T8 — distributed data parallel Design doc for the NCCL DDP path: comm bootstrap (rank-0 UniqueId + grouped CommInitRank), thread-per-GPU launch model (Var is !Send), all-reduce-then- local-step scheme (in-place fp32 AllReduce on .grad() + /world, each rank steps its own GpuAdamW), why params stay consistent (NCCL bit-identical reduce + same init/state), batch sharding math vs single-GPU, verification plan + scaling table. Lists TP/PP/ZeRO/bf16-comm as out-of-scope follow-ups. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-15 17:15:49 +08:00

Author

SHA1

Message

Date

Gahow Wang

0131f05b26

docs: Phase T8 — distributed data parallel

Design doc for the NCCL DDP path: comm bootstrap (rank-0 UniqueId + grouped
CommInitRank), thread-per-GPU launch model (Var is !Send), all-reduce-then-
local-step scheme (in-place fp32 AllReduce on .grad() + /world, each rank steps
its own GpuAdamW), why params stay consistent (NCCL bit-identical reduce + same
init/state), batch sharding math vs single-GPU, verification plan + scaling
table. Lists TP/PP/ZeRO/bf16-comm as out-of-scope follow-ups.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

2026-06-15 17:15:49 +08:00

1 Commits