5.0 KiB
Frontier coverage audit for Qwen3.6-27B
Verdict (2026-07-17):
FAIL: original Frontier coverage. This is a static capability result, not a latency-fidelity result. It says that the examined existing Frontier cannot represent the real Qwen3.6 execution contract; it does not say that a future extended simulator will be wrong.
Question
Before spending profile or replay GPU-hours, can the checked Frontier baseline simulate the same model/engine state that the real server executes?
Compared contracts
| Side | Observed contract |
|---|---|
| Real server | ModelScope Qwen/Qwen3.6-27B@cea40373b9214dd387123e68841890af30dcd469; vLLM 0.20.2, BF16; Qwen3_5ForConditionalGeneration; 64-layer non-MoE hybrid with three Gated DeltaNet/linear-state layers followed by one gated-attention layer per block. The release vLLM run used a FlashInfer GDN prefill kernel and changed requested --block-size 16 to a 784-token attention/state page at TP1/2/4. |
| Frontier baseline | /home/admin/cpfs/wjh/frontier-qwen30-vllm020-profile-v1/Frontier, commit d9cfeb6d8791fbf2f295dd9744c56a666171776e. Its model-spec directory has Qwen2/Qwen3-MoE and all-attention dense examples, but no Qwen3.6 or Qwen3_5 spec. Source search across *.py, *.json, *.md returned no hit for qwen3_5, qwen3.6, GatedDeltaNet, DeltaNet, mamba, linear attention, state cache, or gdn. |
Frontier source model-spec inventory
This is the complete frontend/configs/model_configs/*.json inventory at the
checked commit. It establishes that the simulator can instantiate these
structural specifications; it is not evidence that a matching vLLM
profile, runtime contract, or validated tuning result exists for each model.
| Class | Source specs |
|---|---|
| All-attention, non-MoE | Llama-3.1-405B-Instruct-FP8, Llama-3.2-1B-Instruct, llama2_7b_dense_example, llama3.1-405b, llama3.1-8b, llama3.3-70b, qwen2_dense_test |
| MoE | Phi-tiny-MoE-instruct, Qwen3-235B-A22B, Qwen3-235B-A22B__layers96, Qwen3-30B-A3B-tiny, deepseek-v3, mixtral_8x7b_moe, qwen2_moe_example, qwen3-a3b-30b-moe, qwen3-next-80b-a3b-instruct-reduced-l2, qwen3-next-80b-a3b-instruct-reduced-l20, Step2Mini-tiny, step-moe-noquant-small, step-moe-noquant, step-moe |
The exact JSON architecture labels are llama, qwen2, phimoe,
qwen3_moe, deepseek_v3, mixtral, qwen2_moe, qwen3_next,
step2_mini, and step3_text. The separately held experimental artifacts
in this project establish profile-and-surface evidence only for
Qwen3-30B-A3B and Qwen3-235B-A22B; the remaining names are not
automatically runnable or validated on H20/vLLM merely because a JSON exists.
Why the missing model file is semantic, not administrative
The Frontier execution-time taxonomy separates attention_prefill_execution_time
and attention_decode_execution_time; its operator context and parameter
accounting are built from attention Q/KV heads, attention head dimension, FFN,
and optional MoE routing. No Gated-DeltaNet/linear-state operator, state-cache
capacity rule, or its mixed prefill/decode scheduling state was found.
Adding a JSON whose layer count merely matches Qwen3.6 would therefore map three quarters of its text layers onto an all-attention approximation. Reusing Qwen30/MoE or Qwen3.5 profiles has the same defect. Neither is a valid execution of the existing baseline.
Decision
- Record Frontier as coverage failure for the Qwen3.6 non-MoE-hybrid branch. Do not rank configs, fabricate high latency, or spend GPU-hours on an all-attention profile substitute.
- A new Gated-DeltaNet/state-cache model and profiling implementation would be an extended Frontier intervention. It must be evaluated separately from the capability of the original simulator and cannot support the claim that existing Frontier already solves tuning.
- The Qwen3.6 result cannot be presented as evidence for a conventional pure dense-transformer class. It is a useful non-MoE/hybrid boundary case.
- No runnable local checkout/version of Vidur or APEX was found; only a historical Vidur patch exists. Their coverage remains unassessed, not passed or failed, until their exact released artifacts are frozen and checked against this model/runtime contract.
Next critical experiment
Do not extend Frontier before establishing the requested no-SLO metric on a
model it already represents. The next direct test is the existing
Qwen3-30B-A3B MoE branch under one same-trace Fixed-P replay surface, scored
by mean/p90 TTFT and E2E (TPOT N/A), with simulator-only coverage checked
before any real 600-s sweep. This separates a known architectural coverage
failure from a fidelity failure on an in-scope model.
Evidence locations
- Real topology smoke and resolved page state:
artifacts/qwen36/smoke-v4-20260717/in this run directory. - Real model/runtime audit: dense-preflight.md.
- No-SLO latency protocol and selection gates: experiment-card.md.