12 KiB
Community Qwen235B / Frontier smoke report
Bottom line
The community-vLLM TP4 feasibility smoke passed, but the unmodified Frontier profiles are not valid inputs for a formal simulator-versus-runtime ranking test yet. The smoke exposed an execution-semantic mismatch inside the MoE profiler: serving selects vLLM's tuned H20 block-FP8 Triton config, while the standalone Frontier profiler omits the FP8 dtype key and silently falls back to a default config. At the tested TP4/EP1/16-token point, fixing this lookup reduces measured grouped-GEMM time by about 19%, with the direction consistent across five paired routing seeds.
This is precisely why merely installing the same vLLM release on both sides is not a sufficient alignment contract.
Change
- Built an isolated Python 3.12 environment around community vLLM 0.10.2 and
Frontier commit
d9cfeb6d8791fbf2f295dd9744c56a666171776e. - Pinned Transformers 4.55.2 and tokenizers 0.21.4 after reproducing an incompatibility with the unconstrained Transformers 5.13.1 resolver result.
- Added Frontier metadata for the local block-FP8 checkpoint without changing the original Hugging Face config.
- Collected representative FP8 linear, FlashInfer prefill-attention, FP8 MoE, and TP4 NCCL all-reduce measurements.
- Loaded the full Qwen3-235B-A22B-FP8 checkpoint on four H20s through community vLLM and completed one real request.
- Ran a 2x2 MoE diagnostic that independently toggles the runtime FP8 tuning key and BF16 compute type.
Expected effect
The smoke was intended to determine whether community vLLM can provide a shared, reproducible profiler/serving substrate before paying for the complete response surface. It was not intended to establish Frontier ranking accuracy.
Success required all representative operator and collective paths to execute, the TP4 server to become ready, one request to return successfully, and all GPU processes to be released. A discovered profiler/runtime semantic mismatch was treated as a gate on the later simulator comparison rather than calibrated away with end-to-end serving data.
Frozen environment
- Host:
dash0; 8 NVIDIA H20 GPUs. - Model:
/home/admin/cpfs/wjh/models/Qwen/Qwen3-235B-A22B-FP8. - Original model-config SHA256:
702c46d431bb984db9035a1225186bbfdb52c0d19c82104df4a37cd005e0369e. - Python 3.12.3, torch 2.8.0+cu128, vLLM 0.10.2, flashinfer-python 0.3.1.post1, Transformers 4.55.2, tokenizers 0.21.4.
- Ubuntu 24.04.2, kernel 5.10.134, CUDA runtime 12.8, CUDA toolkit 12.9, NVIDIA driver 580.95.05, GCC 13.3.0, glibc 2.39.
- Two 80-core Intel Xeon sockets (160 online CPUs), 8 H20 GPUs with 97,871 MiB each and all-to-all NV18 links; GPU0--3 share NUMA node 0.
- Serving: TP4, FlashInfer, eager execution, custom all-reduce disabled, FP8 weights, BF16 KV cache, MNS=64, MBT=8192, max model length 40,960, prefix caching and speculative decoding disabled.
The first environment resolved Transformers 5.13.1 because vLLM 0.10.2 only
declares transformers>=4.55.2. That environment failed before model loading
because vLLM accesses all_special_tokens_extended, an API absent from the
new tokenizer implementation. Pinning the declared minimum restored the API;
the original and corrected freezes are both retained.
Verification
End-to-end smoke
| Stage | Result |
|---|---|
| FP8 linear, TP4, 16 tokens | Passed; QKV projection mean 0.2232 ms and output projection mean 0.1105 ms |
| FlashInfer prefill attention, TP4 | Passed for sequence lengths 32/64/96/128; attention mean 0.0571--0.0665 ms |
| FP8 MoE, TP4/EP1, 16 tokens | Passed; unmodified grouped-GEMM mean 0.3075 ms in the initial point |
| TP4 NCCL all-reduce, 128 KiB/rank | Passed; mean 0.03694 ms, p50 0.03138 ms, p95 0.05843 ms |
| Community-vLLM TP4 model load | Passed; 55.1328 GiB of weights/rank, 627 seconds |
| KV-cache sizing | 18.72 GiB/rank, 417,616 tokens; reported 10.20x concurrency at length 40,960 |
| API readiness and one-token completion | Passed; /v1/models and /v1/completions returned HTTP 200 |
| Cleanup | Passed; all eight GPUs returned to zero compute processes and zero MiB used |
The successful serving-only run started at 09:09:52 UTC and completed at 09:22:29 UTC (12 minutes 37 seconds). The one-request response contains one completion choice and one completion token. The copied artifacts match the remote SHA256 manifests.
FP8 path audit
The Frontier confirmation UI prints FP8 Quantization: N/A and labels the
operations BF16 because it displays the unset CLI override rather than the
model-config-derived operation precision. This is misleading, not evidence of
a BF16-weight profile:
- the quantization manager configured
attn_pre_proj,attn_post_proj, andmoe_grouped_gemmas FP8 from the model config; - the linear implementation invokes
apply_w8a8_block_fp8_linearwith a 128x128 weight block; - the MoE implementation quantizes weights and activations and invokes vLLM's
fused kernel with
use_fp8_w8a8=True; FRONTIER_FP8_GEMM_SURROGATEwas not enabled;- each CSV carries
method=fp8|act=dynamic|serialized=True|block=128x128separately from its BF16 output/compute dtype metadata.
MoE runtime-alignment factorial
Point: Qwen3-235B-A22B-FP8, TP4, EP1, 16 input tokens, uniform routing seeds 0--4. Values are the mean of each row's CUDA-event samples, then averaged across the five paired routing seeds.
| Variant | Grouped GEMM mean (ms) | Paired delta vs original | Routing-seed 95% CI |
|---|---|---|---|
| Original: default config + FP16 compute type | 0.3100 | 0.00% | -- |
| FP8 config key only + FP16 compute type | 0.2508 | -19.12% | [-21.26%, -16.98%] |
| Default config + BF16 compute type only | 0.3088 | -0.40% | [-1.31%, 0.51%] |
| FP8 config key + BF16 compute type | 0.2512 | -18.99% | [-21.14%, -16.85%] |
These intervals describe variation across the five routing seeds, not independent process/server-run uncertainty. Even with that limitation, the factorial localizes the dominant error at this point to kernel tuning-config selection rather than FP16-versus-BF16 compute type.
The concrete lookup difference is:
- Frontier calls
get_config_dtype_str(base_dtype). For BF16 this returnsNone, so it searches for a config without an FP8 dtype component and uses the default when that file is absent. - vLLM serving calls the same helper with
use_fp8_w8a8=True, obtainsfp8_w8a8, and loadsE=128,N=384,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128,128].json.
Frontier also hard-codes tl.float16 as the FP8 kernel compute type, while the
serving path uses BF16 for this model. That difference was not material at this
single point, but it should still be corrected to make the counterfactual
execution semantics identical.
Validity audit
Headline claims and benchmark surface
| Claim | Verdict | Scope |
|---|---|---|
| Community vLLM 0.10.2 can load and serve this model on TP4 H20 | PASS | One frozen TP4 configuration and one real request |
| Unmodified Frontier chooses a different MoE tuning config, causing a material operator-time error at the tested point | PASS | TP4/EP1, 16 tokens, five paired routing seeds |
| Frontier is or is not sufficient as a Qwen235B config ranker | NEEDS EVIDENCE | No aligned full profiles, frozen simulator surface, or real config surface yet |
The mechanism diagnostic uses the unmodified Frontier wrapper as baseline, grouped-GEMM CUDA-event latency as its direct metric, and the same model, hardware, vLLM binary, token count, TP/EP point, and routing seeds in all four cells. Absolute times, paired relative differences, across-seed standard deviations, and a paired routing-seed interval are retained. The data range is deliberately narrow and cannot support an end-to-end or cross-config claim.
| Benchmark issue | Verdict | Severity | Evidence / required action |
|---|---|---|---|
| Microbenchmark presented as end-to-end performance | PASS | -- | The 19% number is reported only as mechanism evidence; ranking remains blocked. |
| Simplified profiler matches the real execution path | FAIL | Blocking | Tuning-key and compute-type semantics differ. Apply the alignment patch and test exact config dictionaries before simulation. |
| Statistical significance / repeat protocol | NEEDS EVIDENCE | Major | Five paired routing seeds are not independent process repeats. Repeat in reverse/randomized variant order and at separated times. |
| Selective parameter range | NEEDS EVIDENCE | Major | Only TP4/EP1/16 tokens is measured. Cover the full decision-relevant token and TP range, including boundary values. |
| Calibration set equals evaluation set | PASS | -- | No serving result calibrated the profile; the protocol freezes simulator outputs before ground truth. |
| Proper diagnostic baseline | PASS | -- | Each ablation is compared with the exact unmodified Frontier commit on the same point. |
| Platform and absolute values missing | PASS | -- | OS, kernel, CPU, GPU topology, driver/toolkit/runtime versions, commands, absolute metrics, logs, and hashes are retained. |
| Full ranking baseline/SOTA comparison | N/A | -- | This smoke makes no ranking or tuner-superiority claim; it becomes required in the formal evaluation. |
Overall audit decision: the feasibility smoke and localized mismatch result are reportable, but any Frontier-sufficiency or config-ranking claim is Block until the required reruns and end-to-end comparison are complete.
Result
The shared community stack is feasible: profiles can be collected and the 235B checkpoint can serve on TP4 within memory. The stronger result, however, is a falsification of the current alignment assumption. Same model, hardware, vLLM package, quantization signature, and nominal backend did not imply the same kernel configuration. A small wrapper-level omission produced a roughly 19% operator-time error before any scheduling approximation was involved.
Therefore no Frontier config ranking should be generated from the original MoE CSV and presented as a fair simulator-versus-real comparison. The correct next boundary is: first make the profiler reproduce the serving kernel selection, then freeze profile-only simulator outputs, and only then collect the real response surface.
Remaining risk and next gate
- The 19% result is one TP4/EP1/16-token point. It must be checked across the decision-relevant token counts and TP4/TP8 before claiming its effect on config ordering.
- Frontier's shuffling helper still performs a separate BF16/full-expert-width
config lookup and reports a missing
E=128,N=1536H20 file. At 16 tokens its defaultBLOCK_SIZE_Mmatches the tuned path, but this must be validated over the full token grid. - The derived Frontier model config adds
is_checkpoint_fp8_serialized=true, which Frontier requires to parse the checkpoint metadata. This adaptation must remain explicit and hash-tracked. - The representative attention and collective smokes do not provide the profile coverage needed by the prefill candidate grid. Decode-dominant EP8 remains blocked on all-to-all measurement/consumption.
- A TP4 cold start spends about 10.5 minutes reading weights. Using that time as a lower-bound proxy, eight TP4/TP8 cells already cost roughly 8.4 H20-GPU-hours in weight loading alone; a real one-pass response surface will exceed 10 H20-GPU-hours after initialization and load probes. It requires a separate launch approval after simulator outputs are frozen.
Recommended next step: upstream the two semantic fixes with regression tests
that compare the exact vLLM config dictionary selected by profiler and serving
for each (M, TP, EP, quantization) point; align the shuffling block-size
lookup; regenerate TP4/TP8 profile closure; then run and hash the Frontier
prefill predictions. Do not start the eight-cell real serving sweep before
that gate passes.