xtrain/docs at 6465a2d5ceb2d72a6e580d05d63ca865a679fc8d - xtrain - Local Gitea

gahow/xtrain

Files

History

Gahow Wang 4379868f2d docs: M2d — ragged-batching lever, 9× measured, step bottleneck → rollout

Records the M2d lever (batch the GRPO training-side forwards), the right-pad-is-free
insight, both exact gates, the end-to-end no-OOM smoke, and the 9× throughput.

The honest decomposition correction: M2c claimed the training forwards "dominate" the
step; the clean per-component bench falsifies the strong form — they were ~2.5 s of
the ~8.5 s step (~30%), worth the 9×, but the rollout (~6 s) was always the larger
share. After M2d the step is ~95% rollout, so the next step-level lever is full B×G
rollout batching (today only the G samples of each prompt decode in lockstep; the B
prompts are still sequential). Same measure-first lesson, once more.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

2026-06-30 23:03:28 +08:00

..

docs: v12 — 1.05B long-ctx base + chat-alpha SFT quality check

2026-06-29 16:19:12 +08:00

00-build-chain.md

docs: backfill T1 build-chain

2026-06-15 15:12:55 +08:00

01-tensor.md

docs: Phase T2 — tensor abstraction

2026-06-15 15:12:55 +08:00

02-gemm-autodiff.md

docs: Phase T3 — GEMM fwd/bwd + finite-diff

2026-06-15 15:27:03 +08:00

03-autograd-engine.md

docs: Phase T4 — autograd engine

2026-06-15 15:53:55 +08:00

04-tiny-transformer.md

docs: Phase T5 — tiny transformer

2026-06-15 16:09:30 +08:00

05-training-loop.md

docs: Phase T6 — training loop

2026-06-15 16:30:14 +08:00

06-performance.md

docs: Phase T7 — performance

2026-06-15 17:00:29 +08:00

07-distributed.md

docs: Phase T8 — distributed data parallel

2026-06-15 17:15:49 +08:00

08-export-xserv.md

docs: T9 verification results (xserv == xtrain, dash5)

2026-06-15 17:37:46 +08:00

09-batched-forward.md

docs: Phase T10 — batched forward

2026-06-16 00:44:50 +08:00

10-caching-allocator.md

perf: KI-5 FIXED — single-GPU 40K->93K tok/s, DDP scaling 1.3x->5x@8

2026-06-16 11:15:02 +08:00

11-bf16-mixed-precision.md

perf: KI-2 FIXED — dim768 bf16 fits batch 32, tok/s 31.5K→40.8K

2026-06-16 14:28:20 +08:00

12-activation-recompute.md

perf: KI-3 fixed — dim1024 batch32 fits, mem 31.1→14.6GB, tok/s 39.7K→31.5K

2026-06-17 09:50:29 +08:00

13-flash-attention.md

docs: T14 flash-attention results + evolution/README rows

2026-06-17 23:34:10 +08:00

14-gqa.md

docs: T15 GQA results + evolution row (模型架构) + README build-journey row

2026-06-18 01:44:58 +08:00

15-grad-accum.md

docs: Phase T16 — gradient accumulation design

2026-06-17 23:41:17 +08:00

16-process-per-gpu.md

docs: T17 process-per-GPU results — measured throughput-neutral

2026-06-18 18:03:14 +08:00

17-dropout.md

docs: T21 — record DDP-dropout wiring gap + fix (known-issues / evolution / dropout doc)

2026-06-18 21:22:49 +08:00

18-post-training-rl-sft.md

docs: M2d — ragged-batching lever, 9× measured, step bottleneck → rollout

2026-06-30 23:03:28 +08:00

evolution.md

docs: M2d — ragged-batching lever, 9× measured, step bottleneck → rollout

2026-06-30 23:03:28 +08:00

known-issues.md

docs: T21 — record DDP-dropout wiring gap + fix (known-issues / evolution / dropout doc)

2026-06-18 21:22:49 +08:00