4.0 KiB
EXP-SIMFID-Q235-V020-PORTABILITY: large FP8 MoE boundary case
Status: initial compatibility failure diagnosed on 2026-07-18; the vLLM 0.20 profile path was subsequently repaired and the four-case matrix completed on 2026-07-19. Historical vLLM 0.10.2 SLO data remains excluded.
Scope
This is the requested Qwen3-235B-A22B-FP8 case, kept separate from the Qwen30 surface. Its purpose is first to establish whether the current Frontier stack can even represent the same large-FP8-MoE runtime; it is not evidence that a Qwen30 profile generalizes.
| Item | Frozen proposal |
|---|---|
| model | Qwen3-235B-A22B-FP8: FP8 e4m3/dynamic-block checkpoint; BF16 compute/KV contract recorded at runtime |
| engine | community vLLM 0.20.0 (88d34c640…) |
| simulator | Frontier deadc4a…, with frontier.__file__ recorded by each profiler |
| candidate plans | TP4/EP1 and TP8/EP8, each MNS 64/128, MBT=8192 |
| first workload | Fixed-P: ISL=2048, OSL=1, prefix off, provisional global QPS=1.6 |
| objective | mean/p90 TTFT and E2E; TPOT=N/A |
TP1/TP2 do not fit the 239-GB checkpoint. TP8/EP1 is not assumed legal: only the two plans above have existing historical memory/topology evidence.
Gates
- A fresh vLLM 0.20 TP4 and TP8 server must load the checkpoint and return an exact 2048→1 response without FP8 fallback. This is a compatibility smoke, not a latency sample.
- New model-specific vLLM 0.20/deadc4a profiles must cover linear, router, FP8 MoE, prefill attention, and TP collectives for TP4/TP8. Old vLLM 0.10.2 CUDA-event CSVs are forbidden inputs.
- The profile manifest binds model config/quantization, vLLM, exact Frontier import path/commit, hardware, and finite shape coverage. CPU replay must emit legal request metrics before a real surface is launched.
- Only then run the four-plan real/simulator Fixed-P surface with three fresh trials; a missing simulator cell is coverage failure, not a poor config.
This ordering prevents a profile-runtime mismatch from being misreported as a simulator tuning failure.
Initial result: the real plans ran before Frontier had an admissible counterpart
Both real vLLM 0.20 portability gates completed one exact Fixed-P request
(ISL=2048, OSL=1, prefix off) on H20. These are startup/semantic gates,
not estimates of mean or p90 latency.
| Real plan | Runtime-resolved MoE state | Completed request | TTFT / E2E |
|---|---|---|---|
| TP4/EP1, MNS64 | FP8 MoE TRITON, 128 experts per rank |
1/1 | 1655.99 / 1656.04 ms |
| TP8/EP8, MNS64 | FP8 MoE FLASHINFER_CUTLASS, 16/128 experts per EP rank |
1/1 | 720.54 / 720.58 ms |
The matching Frontier deadc4a profile gate failed before any GPU measurement:
Qwen3-235B-A22B.jsonexposestorch_dtype=BF16but no FP8quantization_config/[128,128]block shape, while the checkpoint uses dynamic block-FP8 weights. Frontier rejects the valid request as a profile configuration mismatch.- Its MoE profiler then imports the removed vLLM symbol
invoke_fused_moe_kernel, so it is not ABI-compatible with the installed community vLLM 0.20 runtime.
This was an admissible intermediate coverage failure. Filling it with historical vLLM 0.10.2 profiles would have changed both the runtime and the selected kernels and would not have answered the stated comparison question.
Final result after profile repair
The later campaign added graph-compatible kernel-only decode profiles and true
mixed attention coverage, then completed all four workloads on the frozen
{TP4/EP1, TP8/EP8} x MNS{64,128} surface. All 48 real fresh-server trials
and all 16 Frontier cells completed. The decision-level selection regrets are
maintained in the consolidated
Qwen30/Qwen235 report.
The decisive counterexample is Fixed-PD: Frontier's selected configuration is 33.0--37.2% slower for real TPOT and 30.7--34.6% slower for real E2E than the real-hardware optimum. Thus repairing the coverage gate enabled a valid comparison but did not establish tuning fidelity.