Files
aituner/docs/simulator-tuning/frontier-selection-regret-qwen30-qwen235-20260719.md

76 lines
3.5 KiB
Markdown

# Frontier selection regret: Qwen3-30B and Qwen3-235B
> Date: 2026-07-19
> Scope: H20, community vLLM 0.20, Frontier piecewise simulation, no SLO gate
## Question and metric
For each workload and latency objective, Frontier selects the configuration
with the lowest simulated latency. We then look up that configuration on the
complete real-hardware surface and compare it with the real-hardware optimum.
```text
selection regret = real_latency(Frontier winner) / real_latency(real winner) - 1
```
Lower is better. `0%` means Frontier selected the real winner. Positive values
mean that following Frontier produces slower real serving. Each objective is
selected independently; this table does not combine TTFT, TPOT, and E2E into a
single score.
## Qwen3-30B-A3B
Configuration surface: `TP in {1,2,4} x MNS in {8,16,32,64}`, with
`MBT=8192`. Each real cell uses three fresh-server trials.
| Workload | TTFT mean | TTFT p90 | TPOT mean | TPOT p90 | E2E mean | E2E p90 |
|---|---:|---:|---:|---:|---:|---:|
| Trace-PD | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% |
| Fixed-PD, 4096->256, 1.125 req/s/GPU | **58.0%** | **56.2%** | 0.0% | 0.0% | 1.7% | 5.5% |
| Trace-PO, OSL=1 | 3.2% | 0.4% | N/A | N/A | 3.2% | 0.3% |
| Fixed-PO, 4096->1, 1.125 req/s/GPU | 0.3% | 0.5% | N/A | N/A | 0.3% | 0.5% |
Interpretation: Frontier is near-optimal for Trace-PD and both prefill-only
cases, but the high-pressure Fixed-PD TTFT choice is materially wrong: its
selected configuration is 56--58% slower than the real TTFT optimum.
## Qwen3-235B-A22B-FP8
Configuration surface: `{TP4/EP1, TP8/EP8} x MNS in {64,128}`, with
`MBT=8192`. Each workload has 129 requests per cell and each real cell uses
three fresh-server trials.
| Workload | TTFT mean | TTFT p90 | TPOT mean | TPOT p90 | E2E mean | E2E p90 |
|---|---:|---:|---:|---:|---:|---:|
| Trace-PD | 0.0% | 0.0% | 0.0% | 0.0% | 0.6% | 6.2% |
| Fixed-PD, 4096->256, 0.2 req/s/GPU | 4.2% | 0.2% | **33.0%** | **37.2%** | **30.7%** | **34.6%** |
| Trace-PO, OSL=1 | 7.0% | **21.2%** | N/A | N/A | 7.0% | **21.2%** |
| Fixed-PO, 4096->1, 0.2 req/s/GPU | 5.9% | 1.7% | N/A | N/A | 5.9% | 1.7% |
Interpretation: Trace-PD is mostly near-optimal. Fixed-PD reverses the real
decode/E2E preference between the tested parallel configurations and incurs
31--37% regret. Trace-PO also has a material p90 failure of 21.2%.
## Decision
The tested Frontier stack has **not** solved serving configuration tuning.
Its selected configuration can be near-optimal for one workload and materially
wrong for another on the same model and hardware. The strongest current
counterexamples are Qwen3-30B Fixed-PD TTFT and Qwen3-235B Fixed-PD TPOT/E2E.
This statement is limited to the two tested MoE models and Frontier. It is not
yet evidence about dense models, Vidur/APEX as separately reproduced systems,
other hardware, or SLO-constrained tuning.
## Provenance
Primary immutable analysis artifacts on `dash0`:
- Qwen3-30B Trace-PD: `/home/admin/cpfs/wjh/aituner/graph-piecewise-qwen30-20260717/simulator-piecewise-surface-v2/analysis/comparison.json`
- Qwen3-30B Fixed-PD/PO: `/home/admin/cpfs/wjh/aituner/qwen30-fixed-pressure-surface-20260719-r1/analysis/`
- Qwen3-30B Trace-PO: `/home/admin/cpfs/wjh/aituner/qwen30-latency-expansion-20260718-r2/analysis-r6/trace-po-comparison.json`
- Qwen3-235B four-case matrix: `/home/admin/cpfs/wjh/aituner/qwen235-v020-fourcase-20260719-r1/analysis/comparison.json`
The Qwen3-235B artifact root includes `provenance/artifacts.sha256`; the final
matrix contains 48/48 valid real trials and 16/16 complete simulator cells.