# Qwen3.6-27B non-MoE/hybrid preflight and topology-smoke record > **Status (2026-07-17): `TOPOLOGY-SMOKE PASS; FRONTIER COVERAGE BLOCKED`.** > The exact snapshot and official community-vLLM release are accepted through > a fresh-server TP1/TP2/TP4 GPU smoke. No Frontier Qwen3.6 contract/profile, > trace replay, or latency-matrix cell exists. The smoke is compatibility > evidence only, not simulator-fidelity or latency-selection evidence. ## Decision enabled by this preflight The latency-matrix experiment may add its dense branch only after the exact model snapshot, a runnable open-source vLLM environment, and a Qwen3.6 Frontier contract/profile root exist on the same stack. Until then, all dense matrix cells are **`NOT RUN`**. They are neither simulator failures nor evidence for dense+MoE generalization. ## Executed staging record (after explicit approval) | Step | Evidence | Result | |---|---|---| | Immutable snapshot staging | ModelScope `Qwen/Qwen3.6-27B@cea40373b9214dd387123e68841890af30dcd469` was downloaded to `/home/admin/cpfs/wjh/models/Qwen/Qwen3.6-27B`. The downloader recorded `download_start` and `download_complete` at `2026-07-17T07:42:00Z` and `2026-07-17T07:44:45Z`, respectively, in `/home/admin/cpfs/wjh/models/.staging/qwen36-27b-20260717/download.log`. | **PASS**; 29 snapshot files, approximately 52 GiB. | | Snapshot integrity | Sorted per-file SHA-256 inventory: `/home/admin/cpfs/wjh/models/.staging/qwen36-27b-20260717/files.sha256`; inventory digest `031cea103726cbf27732b6971e898e210aee77f0c32e4bc703c947e0372349c4`. It includes `config.json`, tokenizer files, and all 15 `model-*.safetensors` shards; no `*.incomplete`, `*.lock`, or `*.tmp` files were found. | **PASS**. | | CPU config/tokenizer parser | The preinstalled system `transformers==4.57.5` rejects `model_type=qwen3_5`; this environment is excluded. A dedicated CPU-only environment with `transformers==5.5.3` successfully parsed the snapshot with `CUDA_VISIBLE_DEVICES=""`: `Qwen3_5ForConditionalGeneration`, text model type `qwen3_5_text`, 64 layers, hidden size 5120, `mamba_ssm_dtype=float32`, tokenizer length 248077. | **PASS**, but only for checkpoint/tokenizer parsing; it is not the serving runtime. | | Community vLLM release runtime | The initial source-build attempt was stopped before acceptance and is excluded. Official PyPI wheel `vllm==0.20.2` was staged at `wheel-probe-proxy/vllm-0.20.2-cp38-abi3-manylinux_2_35_x86_64.whl`, SHA-256 `22a7dd06eb03371298e13d6100f3dedbf307352342aaf08e87c929c60aae9b4d`, then installed into the clean environment `/home/admin/cpfs/wjh/venvs/qwen36-vllm-0.20.2`. `pip check` reports no broken requirements. | **PASS (CPU-only)**: vLLM `0.20.2`, Torch `2.11.0`, Transformers `5.14.1`; with `CUDA_VISIBLE_DEVICES=""`, `vllm.transformers_utils.config.get_config(..., trust_remote_code=False)` parses the exact snapshot as `Qwen3_5ForConditionalGeneration` / `qwen3_5`. | | Fresh GPU topology smoke | Fresh text-only servers on H20 TP1, TP2, TP4 used the exact release runtime with requested `MNS=8`, `MBT=8192`, BF16, chunked prefill on, prefix cache off, and one `2048 → 1` request per topology. Every response reported exactly 2048 prompt and 1 completion token; all four used GPUs were zero-memory after cleanup. Evidence root: `artifacts/qwen36/smoke-v4-20260717/`; artifact inventory SHA-256 `5b09c94e61801bfb7085257cfeea88049dfa81001f50810d85b9f87468088be0`. | **PASS (GPU compatibility)**. It establishes that TP1/2/4 can serve this model on the intended release, not that any topology has lower latency. | ## Confirmed model identity and architecture | Item | Confirmed fact | Consequence for the experiment | |---|---|---| | Canonical source | ModelScope [`Qwen/Qwen3.6-27B`](https://www.modelscope.cn/models/Qwen/Qwen3.6-27B) | Use this repository, not an inferred Hugging Face mirror or an old Qwen3.5 checkpoint. | | Immutable revision | Read-only `git ls-remote --symref` on 2026-07-17 resolved `master`/`HEAD` to `cea40373b9214dd387123e68841890af30dcd469`. | Pin this commit hash in the download record. Do **not** record only the mutable name `master`. Re-resolve it immediately before an approved download; a changed hash requires a new review. | | Model metadata | ModelScope API reports `Architectures=["Qwen3_5ForConditionalGeneration"]` and `ModelType=["qwen3_5"]`. The official card says the artifact is a causal language model with a vision encoder. | This is dense with respect to MoE routing, but it is **not** a conventional all-attention dense transformer. The text benchmark must send no media; the model runner must still be the architecture selected from the exact checkpoint config. | | Text architecture (official card) | 27B parameters; hidden size 5120; padded token embedding / LM output 248320; 64 layers; `16 × (3 × (Gated DeltaNet → FFN) → 1 × (Gated Attention → FFN))`; DeltaNet has V=48/QK=16 heads of 128; gated attention has Q=24/KV=4 heads of 256 and RoPE dimension 64; FFN intermediate size 17408; native context 262144. | A Frontier profile must cover both DeltaNet/linear-state and gated-attention paths, in prefill, decode, and true mixed execution. An all-attention Qwen3.5 or Qwen3-30B/MoE profile is not a valid substitute. | | Tokenizer | The staged snapshot contains the tokenizer artifacts in the immutable file inventory. The earlier dedicated CPU parser reported tokenizer length `248077`; `248320` above is still only the padded embedding/output size, not the tokenizer contract. | Request input/output lengths must be tokenized with this staged tokenizer. Before trace materialization, freeze a tokenizer/config subset digest and do not borrow counts from a Qwen3.5 tokenizer. | The official ModelScope card states that these artifacts are compatible with Transformers, vLLM, SGLang, and KTransformers. That is useful compatibility evidence, not a successful run on the selected `dash0` stack. ## Runtime provenance and excluded alternatives | Evidence | Read-only observation | Verdict | |---|---|---| | System vLLM | `/usr/local/bin/vllm` reports `0.13.0rc2.dev2111+gb44b43f43.d20260309`, with Torch `2.8.0a0+5228986c39.nv25.6` and Transformers `4.57.5`. Its installed source contains `qwen3_5.py`, `qwen3_5_mtp.py`, registry entries for `Qwen3_5ForConditionalGeneration`, and branches for `qwen3_5_text`. | Architecture-family support is present, but this development build is not yet the frozen community experiment runtime. | | Uninstalled local source | `/home/admin/cpfs/wjh/agentic-kv/third_party/vllm_v20_build` is `v0.20.0` source with build by-products; `/home/admin/cpfs/wjh/venvs/vllm-0.20.0` does not contain a runnable vLLM/Transformers pair. | **Excluded.** It is not the experiment runtime and must not be cited as an alternative same-stack result. | | Qwen-specific state/page rule | The v0.20 Qwen3.5 implementation obtains Mamba/DeltaNet state settings from the checkpoint when cache dtype is auto; the actual v0.20.2 GPU smoke then resolved its attention page to 784 tokens. | The state/page rule is a model-state contract, not an optional benchmark flag. The profile and simulator must consume the observed physical page setting. | Therefore the defensible statement now is: **the exact Qwen3.6-27B snapshot is CPU-parsed and GPU-served by the frozen official community vLLM release on `dash0` at TP1/2/4. It is a non-MoE hybrid model, not a conventional all-attention dense-transformer evidence point.** ## GPU topology-smoke observations (not a latency experiment) The runner is a one-request topology/usage check, with a fresh server for each TP. Its E2E values are retained only to audit that a request completed; they are not steady-state samples, are not TTFT/TPOT measurements, and must not be ranked. | TP | Exact usage | Smoke E2E (audit only) | Resolved attention page/block size | GPU KV cache at `max_model_len=4096` | |---|---|---:|---:|---:| | 1 | 2048 prompt, 1 completion | 4581.977 ms | 784 tokens | 306,289 tokens | | 2 | 2048 prompt, 1 completion | 5606.756 ms | 784 tokens | 1,068,600 tokens | | 4 | 2048 prompt, 1 completion | 6252.274 ms | 784 tokens | 2,630,087 tokens | vLLM accepted the requested `--block-size 16` but then logged, at every TP, `Setting attention block size to 784 tokens to ensure that attention page size is >= mamba page size.` Thus **block size 16 is not the actual engine state for this model**. A later real/simulator comparison must freeze and expose the resolved 784-token page/state contract (or deliberately select a different legal engine mode before profiling); it may not label this stack as a block-16 comparison. First TP1 startup also compiled FlashInfer GDN kernels and performed a cold profile/warmup; these one-time costs remain outside all request-latency metrics. ## Initial read-only `dash0` inventory (historical; superseded where noted) All paths below were inspected without mutation on 2026-07-17. | Area | Observation | Consequence | |---|---|---| | Disk | `/home/admin` root filesystem: 245 GiB available. Shared CPFS mounted at `/home/admin/cpfs`: 1.1 TiB available, already 95% used. Existing `/home/admin/cpfs/wjh/models` is 1.2 TiB and `/home/admin/resource/model` is 1.7 TiB. | Stage only after an explicit capacity check; do not rely on the root filesystem or an unconstrained default cache. | | Checked model/cache locations | Before staging, the usual locations were absent. | **Superseded:** the immutable snapshot is now at `/home/admin/cpfs/wjh/models/Qwen/Qwen3.6-27B` and is bound by the inventory above. | | Main checkout | `/home/admin/cpfs/wjh/aituner/aituner` had tracked and untracked user changes. | The smoke wrote only new artifacts and did not modify its source. A future profile implementation must use an explicitly versioned source/patch root rather than treating this dirty checkout as code provenance. | | Frontier | Historical Frontier source exists at `/home/admin/cpfs/wjh/frontier-qwen30-vllm020-profile-v1/Frontier`, commit `d9cfeb6d8791fbf2f295dd9744c56a666171776e` (`pre-release-v0.2`), with a clean tracked tree. A source search found no `Qwen3.6`, `Qwen3_5`, or `qwen3_5` model-contract hit. | It is not an existing Qwen3.6 profile/model contract. A new dense adapter/spec must be versioned and accepted before it may be used. | | Existing profiler contract | The old Qwen3-30B profile launcher hard-codes `model=qwen3-a3b-30b-moe`, `--is_moe`, vLLM 0.20 API adapters, and Frontier commit `d9cf…`. | It cannot be repurposed by only changing a model path; its MoE assumptions and compatibility adapter are excluded from the dense baseline. | | Qwen3.6 artifacts | A read-only search of the checkout's `runs` subtree found no Qwen3.6-named run/profile/model path. | No dense Frontier profile or real/simulator output is admissible today. | ## Remaining gates before any latency-matrix cell 1. **Trace/tokenizer lock:** materialize the Qwen3.6 tokenized Fixed/Trace-P/ Trace-PD vectors and hash the exact tokenizer/config subset. Trace-PD is not trace-faithful without the input, output, arrival and prefix-state vector. 2. **New Frontier contract:** introduce a versioned Qwen3.6 **non-MoE hybrid** model spec that maps the exact `config.json` to simulator operators. It must represent Gated DeltaNet/linear-state, gated attention, FFN, TP collectives and KV/ state-cache capacity, including the resolved 784-token page rule. It may not import the Qwen3-30B MoE routing profile, an all-attention Qwen3.5 profile, or an unrecorded additive timing scale. Any compatibility adapter is a separate, hashed intervention. 3. **Profile manifest:** collect and freeze per-TP (1/2/4) profile rows for the above stages over shapes covering Fixed-P, Fixed-PD, Trace-P and Trace-PD. Include prefill, decode and genuine mixed batches, plus TP2/TP4 all-reduce. The manifest must bind model revision/config hash, tokenizer hash, vLLM/environment hash, Frontier commit/adapter hash, hardware, precision, profiler command, CSV hashes, and shape coverage. 4. **CPU simulator acceptance:** reject the profile root unless every required profile family has finite rows for all legal TP values and requested shapes; all identity hashes match the real-server lock; profile/model manifests name Qwen3.6; and the simulator emits request-level TTFT/E2E/TPOT where defined for a no-GPU fixture. A missing simulator operator is a **coverage blocker**, not a license to borrow a Qwen30 profile. Only after all four gates may a separately approved real/simulator 600-s latency surface begin. The GPU smoke closed the runtime/topology gate only; it does not close any of these simulator or evaluation gates. ## Evidence locations - This record: `runs/simulator-tuning-latency-matrix-v0/dense-preflight.md`. - Matrix decision/protocol: `runs/simulator-tuning-latency-matrix-v0/experiment-card.md`. - Existing (excluded) MoE-v0.20 profile launcher and its hard-coded contract: `runs/frontier-qwen30-vllm020-profile-v1/run_frontier_linear_smoke.sh` and `runs/frontier-qwen30-vllm020-profile-v1/frontier_vllm020_compat.py`. - Existing Frontier surface runner shows the required external source, ReplayServe builder, profile root, request-level metric, and hash contract: `runs/frontier-phase-factorial-v0/run_frontier_qwen30_prefill_surface.py`. ## Not completed - No Qwen3.6 Frontier model spec or profile exists. - No Qwen3.6 fixed/trace workload vector is frozen. - No real or simulator latency measurement exists, so this record cannot change the simulator-selection verdict.