Files
aituner/runs/frontier-fidelity-envelope-v1/qwen30-latency-expansion-card.md

5.1 KiB
Raw Blame History

EXP-SIMFID-Q30-LATENCY-EXPANSION: remaining Fixed/Trace × PD/P cases

Status: completed (2026-07-18). The low-load Fixed cases below are retained as sensitivity evidence; the later 4096-token, 1.125 req/s/GPU Fixed matrix is the primary comparison. Trace-PD is excluded from this card.

Question and fixed boundary

Does the same graph-aligned Frontier configuration that selected the correct winner on Trace-PD also select the real vLLM winner when either the arrival/ prefix state or the decode phase is removed? This is a selection test, not an SLO/capacity evaluation.

Case Input / output arrival and cache Applicable objectives
Fixed-PD 2048 / 128 129 uniform requests; t'=t/TP; prefix cache off mean/p90 TTFT, TPOT, E2E
Trace-P exact held-out input, arrival, session and complete block-16 prefix relation / 1 same TP-normalized 129-row trace; prefix cache on mean/p90 TTFT, E2E; TPOT=N/A
Fixed-P 2048 / 1 129 uniform requests; t'=t/TP; prefix cache off mean/p90 TTFT, E2E; TPOT=N/A

All cases use Qwen3-30B-A3B BF16, community vLLM 0.20.0, H20, TP∈{1,2,4} and MNS∈{8,16,32,64}, MBT=8192, chunked prefill, three fresh-server trials per cell. The fixed QPS is the Trace-PD base offered rate (129/600 req/s per GPU), hence the global arrival rate is multiplied by TP. It makes throughput per GPU comparable without claiming that different TP values see identical cluster-level load.

Simulator and real contracts

  • Frontier uses commit deadc4a…, piecewise, the frozen CUDA-event profile for prefill/mixed batches and KERNEL_ONLY profile for captured pure decode.
  • Trace-P reuses the already verified graph buckets/KV capacities because the server CLI is unchanged from Trace-PD. The only workload change is OSL→1.
  • Fixed-P/PD disables vLLM prefix caching on both sides. Before freezing their simulator surfaces, a no-request server-start preflight records the actual graph captures and KV-block capacities for every (TP,MNS); no request latency or winner is used as calibration.
  • Each real result verifies every request's input/output usage and row-vector digest. OSL=1 produces JSON null TPOT samples and is never converted to zero.

Decision rule and cost

For every applicable objective, compare the complete 12-cell simulator and pooled three-trial real surface: winner match, selected-config real regret, and non-tied pair direction agreement. A simulator crash or missing request metric is a coverage failure, not a high-latency cell.

The no-request Fixed runtime preflight is capped at 2 H20-GPUh. Each 36-run real surface is estimated at 13 nominal / 41 worst-case H20-GPUh, plus CPU-only Frontier replay. Launch logs record the resolved inputs, paths, and duration.

Result: exact-selection fidelity is not stable across workload phase/state

All three real surfaces completed 36 valid fresh-server trials (12 TP×MNS cells × 3 trials), each with 129/129 exact-usage requests. The corresponding 12-cell Frontier replays completed all 129 requests per cell. The following table compares the simulator argmin with the pooled real argmin; regret is the selected simulator configuration's real relative latency regret. It deliberately reports exact winner equality rather than treating a high pairwise-order agreement as success.

Case Objective Frontier winner Real winner Match Real regret
Trace-P (OSL=1) TTFT mean TP4/MNS16 TP4/MNS64 no 3.2%
Trace-P (OSL=1) TTFT p90 TP4/MNS16 TP4/MNS8 no 0.4%
Trace-P (OSL=1) E2E mean TP4/MNS16 TP4/MNS64 no 3.2%
Trace-P (OSL=1) E2E p90 TP4/MNS16 TP4/MNS8 no 0.3%
Fixed-PD TTFT mean/p90 TP4/MNS16 TP4/MNS16 yes 0.0%
Fixed-PD TPOT mean TP1/MNS16 TP4/MNS32 no 14.0%
Fixed-PD TPOT p90 TP1/MNS16 TP4/MNS8 no 13.4%
Fixed-PD E2E mean TP1/MNS16 TP4/MNS32 no 26.9%
Fixed-PD E2E p90 TP1/MNS16 TP4/MNS8 no 25.7%
Fixed-P (OSL=1) TTFT/E2E mean/p90 TP4/MNS16 TP4/MNS8 no 1.5--3.6%

The decisive failure is Fixed-PD: its TPOT and E2E pairwise orders are exactly reversed on all 48 informative pairs. Thus the graph-aligned profile can preserve coarse TP ordering yet still choose the wrong configuration for the serving objective. The prior Trace-PD success is therefore a workload-specific success, not evidence that Frontier has solved config tuning.

Superseding primary Fixed workload

The subsequent pressure sweep changed the Fixed shapes to 4096->256 and 4096->1, raised the offered load to 1.125 req/s/GPU, and used 257 requests per cell. Those Fixed-PD/PO results, together with Trace-PD and Trace-PO, are reported in the consolidated Qwen30/Qwen235 selection-regret report.

Under that primary Fixed-PD workload, Frontier's selected config is 58.0% slower for real mean TTFT and 56.2% slower for real p90 TTFT than the real-hardware optimum. The low-load table above must therefore not be quoted as the final Fixed-PD result.