Files
aituner/runs/frontier-multicase-sufficiency-v0/results/community-qwen235b-smoke/report.md

12 KiB

Community Qwen235B / Frontier smoke report

Bottom line

The community-vLLM TP4 feasibility smoke passed, but the unmodified Frontier profiles are not valid inputs for a formal simulator-versus-runtime ranking test yet. The smoke exposed an execution-semantic mismatch inside the MoE profiler: serving selects vLLM's tuned H20 block-FP8 Triton config, while the standalone Frontier profiler omits the FP8 dtype key and silently falls back to a default config. At the tested TP4/EP1/16-token point, fixing this lookup reduces measured grouped-GEMM time by about 19%, with the direction consistent across five paired routing seeds.

This is precisely why merely installing the same vLLM release on both sides is not a sufficient alignment contract.

Change

  • Built an isolated Python 3.12 environment around community vLLM 0.10.2 and Frontier commit d9cfeb6d8791fbf2f295dd9744c56a666171776e.
  • Pinned Transformers 4.55.2 and tokenizers 0.21.4 after reproducing an incompatibility with the unconstrained Transformers 5.13.1 resolver result.
  • Added Frontier metadata for the local block-FP8 checkpoint without changing the original Hugging Face config.
  • Collected representative FP8 linear, FlashInfer prefill-attention, FP8 MoE, and TP4 NCCL all-reduce measurements.
  • Loaded the full Qwen3-235B-A22B-FP8 checkpoint on four H20s through community vLLM and completed one real request.
  • Ran a 2x2 MoE diagnostic that independently toggles the runtime FP8 tuning key and BF16 compute type.

Expected effect

The smoke was intended to determine whether community vLLM can provide a shared, reproducible profiler/serving substrate before paying for the complete response surface. It was not intended to establish Frontier ranking accuracy.

Success required all representative operator and collective paths to execute, the TP4 server to become ready, one request to return successfully, and all GPU processes to be released. A discovered profiler/runtime semantic mismatch was treated as a gate on the later simulator comparison rather than calibrated away with end-to-end serving data.

Frozen environment

  • Host: dash0; 8 NVIDIA H20 GPUs.
  • Model: /home/admin/cpfs/wjh/models/Qwen/Qwen3-235B-A22B-FP8.
  • Original model-config SHA256: 702c46d431bb984db9035a1225186bbfdb52c0d19c82104df4a37cd005e0369e.
  • Python 3.12.3, torch 2.8.0+cu128, vLLM 0.10.2, flashinfer-python 0.3.1.post1, Transformers 4.55.2, tokenizers 0.21.4.
  • Ubuntu 24.04.2, kernel 5.10.134, CUDA runtime 12.8, CUDA toolkit 12.9, NVIDIA driver 580.95.05, GCC 13.3.0, glibc 2.39.
  • Two 80-core Intel Xeon sockets (160 online CPUs), 8 H20 GPUs with 97,871 MiB each and all-to-all NV18 links; GPU0--3 share NUMA node 0.
  • Serving: TP4, FlashInfer, eager execution, custom all-reduce disabled, FP8 weights, BF16 KV cache, MNS=64, MBT=8192, max model length 40,960, prefix caching and speculative decoding disabled.

The first environment resolved Transformers 5.13.1 because vLLM 0.10.2 only declares transformers>=4.55.2. That environment failed before model loading because vLLM accesses all_special_tokens_extended, an API absent from the new tokenizer implementation. Pinning the declared minimum restored the API; the original and corrected freezes are both retained.

Verification

End-to-end smoke

Stage Result
FP8 linear, TP4, 16 tokens Passed; QKV projection mean 0.2232 ms and output projection mean 0.1105 ms
FlashInfer prefill attention, TP4 Passed for sequence lengths 32/64/96/128; attention mean 0.0571--0.0665 ms
FP8 MoE, TP4/EP1, 16 tokens Passed; unmodified grouped-GEMM mean 0.3075 ms in the initial point
TP4 NCCL all-reduce, 128 KiB/rank Passed; mean 0.03694 ms, p50 0.03138 ms, p95 0.05843 ms
Community-vLLM TP4 model load Passed; 55.1328 GiB of weights/rank, 627 seconds
KV-cache sizing 18.72 GiB/rank, 417,616 tokens; reported 10.20x concurrency at length 40,960
API readiness and one-token completion Passed; /v1/models and /v1/completions returned HTTP 200
Cleanup Passed; all eight GPUs returned to zero compute processes and zero MiB used

The successful serving-only run started at 09:09:52 UTC and completed at 09:22:29 UTC (12 minutes 37 seconds). The one-request response contains one completion choice and one completion token. The copied artifacts match the remote SHA256 manifests.

FP8 path audit

The Frontier confirmation UI prints FP8 Quantization: N/A and labels the operations BF16 because it displays the unset CLI override rather than the model-config-derived operation precision. This is misleading, not evidence of a BF16-weight profile:

  • the quantization manager configured attn_pre_proj, attn_post_proj, and moe_grouped_gemm as FP8 from the model config;
  • the linear implementation invokes apply_w8a8_block_fp8_linear with a 128x128 weight block;
  • the MoE implementation quantizes weights and activations and invokes vLLM's fused kernel with use_fp8_w8a8=True;
  • FRONTIER_FP8_GEMM_SURROGATE was not enabled;
  • each CSV carries method=fp8|act=dynamic|serialized=True|block=128x128 separately from its BF16 output/compute dtype metadata.

MoE runtime-alignment factorial

Point: Qwen3-235B-A22B-FP8, TP4, EP1, 16 input tokens, uniform routing seeds 0--4. Values are the mean of each row's CUDA-event samples, then averaged across the five paired routing seeds.

Variant Grouped GEMM mean (ms) Paired delta vs original Routing-seed 95% CI
Original: default config + FP16 compute type 0.3100 0.00% --
FP8 config key only + FP16 compute type 0.2508 -19.12% [-21.26%, -16.98%]
Default config + BF16 compute type only 0.3088 -0.40% [-1.31%, 0.51%]
FP8 config key + BF16 compute type 0.2512 -18.99% [-21.14%, -16.85%]

These intervals describe variation across the five routing seeds, not independent process/server-run uncertainty. Even with that limitation, the factorial localizes the dominant error at this point to kernel tuning-config selection rather than FP16-versus-BF16 compute type.

The concrete lookup difference is:

  • Frontier calls get_config_dtype_str(base_dtype). For BF16 this returns None, so it searches for a config without an FP8 dtype component and uses the default when that file is absent.
  • vLLM serving calls the same helper with use_fp8_w8a8=True, obtains fp8_w8a8, and loads E=128,N=384,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128,128].json.

Frontier also hard-codes tl.float16 as the FP8 kernel compute type, while the serving path uses BF16 for this model. That difference was not material at this single point, but it should still be corrected to make the counterfactual execution semantics identical.

Validity audit

Headline claims and benchmark surface

Claim Verdict Scope
Community vLLM 0.10.2 can load and serve this model on TP4 H20 PASS One frozen TP4 configuration and one real request
Unmodified Frontier chooses a different MoE tuning config, causing a material operator-time error at the tested point PASS TP4/EP1, 16 tokens, five paired routing seeds
Frontier is or is not sufficient as a Qwen235B config ranker NEEDS EVIDENCE No aligned full profiles, frozen simulator surface, or real config surface yet

The mechanism diagnostic uses the unmodified Frontier wrapper as baseline, grouped-GEMM CUDA-event latency as its direct metric, and the same model, hardware, vLLM binary, token count, TP/EP point, and routing seeds in all four cells. Absolute times, paired relative differences, across-seed standard deviations, and a paired routing-seed interval are retained. The data range is deliberately narrow and cannot support an end-to-end or cross-config claim.

Benchmark issue Verdict Severity Evidence / required action
Microbenchmark presented as end-to-end performance PASS -- The 19% number is reported only as mechanism evidence; ranking remains blocked.
Simplified profiler matches the real execution path FAIL Blocking Tuning-key and compute-type semantics differ. Apply the alignment patch and test exact config dictionaries before simulation.
Statistical significance / repeat protocol NEEDS EVIDENCE Major Five paired routing seeds are not independent process repeats. Repeat in reverse/randomized variant order and at separated times.
Selective parameter range NEEDS EVIDENCE Major Only TP4/EP1/16 tokens is measured. Cover the full decision-relevant token and TP range, including boundary values.
Calibration set equals evaluation set PASS -- No serving result calibrated the profile; the protocol freezes simulator outputs before ground truth.
Proper diagnostic baseline PASS -- Each ablation is compared with the exact unmodified Frontier commit on the same point.
Platform and absolute values missing PASS -- OS, kernel, CPU, GPU topology, driver/toolkit/runtime versions, commands, absolute metrics, logs, and hashes are retained.
Full ranking baseline/SOTA comparison N/A -- This smoke makes no ranking or tuner-superiority claim; it becomes required in the formal evaluation.

Overall audit decision: the feasibility smoke and localized mismatch result are reportable, but any Frontier-sufficiency or config-ranking claim is Block until the required reruns and end-to-end comparison are complete.

Result

The shared community stack is feasible: profiles can be collected and the 235B checkpoint can serve on TP4 within memory. The stronger result, however, is a falsification of the current alignment assumption. Same model, hardware, vLLM package, quantization signature, and nominal backend did not imply the same kernel configuration. A small wrapper-level omission produced a roughly 19% operator-time error before any scheduling approximation was involved.

Therefore no Frontier config ranking should be generated from the original MoE CSV and presented as a fair simulator-versus-real comparison. The correct next boundary is: first make the profiler reproduce the serving kernel selection, then freeze profile-only simulator outputs, and only then collect the real response surface.

Remaining risk and next gate

  • The 19% result is one TP4/EP1/16-token point. It must be checked across the decision-relevant token counts and TP4/TP8 before claiming its effect on config ordering.
  • Frontier's shuffling helper still performs a separate BF16/full-expert-width config lookup and reports a missing E=128,N=1536 H20 file. At 16 tokens its default BLOCK_SIZE_M matches the tuned path, but this must be validated over the full token grid.
  • The derived Frontier model config adds is_checkpoint_fp8_serialized=true, which Frontier requires to parse the checkpoint metadata. This adaptation must remain explicit and hash-tracked.
  • The representative attention and collective smokes do not provide the profile coverage needed by the prefill candidate grid. Decode-dominant EP8 remains blocked on all-to-all measurement/consumption.
  • A TP4 cold start spends about 10.5 minutes reading weights. Using that time as a lower-bound proxy, eight TP4/TP8 cells already cost roughly 8.4 H20-GPU-hours in weight loading alone; a real one-pass response surface will exceed 10 H20-GPU-hours after initialization and load probes. It requires a separate launch approval after simulator outputs are frozen.

Recommended next step: upstream the two semantic fixes with regression tests that compare the exact vLLM config dictionary selected by profiler and serving for each (M, TP, EP, quantization) point; align the shuffling block-size lookup; regenerate TP4/TP8 profile closure; then run and hash the Frontier prefill predictions. Do not start the eight-cell real serving sweep before that gate passes.