76 lines
3.5 KiB
Markdown
76 lines
3.5 KiB
Markdown
# Frontier selection regret: Qwen3-30B and Qwen3-235B
|
|
|
|
> Date: 2026-07-19
|
|
> Scope: H20, community vLLM 0.20, Frontier piecewise simulation, no SLO gate
|
|
|
|
## Question and metric
|
|
|
|
For each workload and latency objective, Frontier selects the configuration
|
|
with the lowest simulated latency. We then look up that configuration on the
|
|
complete real-hardware surface and compare it with the real-hardware optimum.
|
|
|
|
```text
|
|
selection regret = real_latency(Frontier winner) / real_latency(real winner) - 1
|
|
```
|
|
|
|
Lower is better. `0%` means Frontier selected the real winner. Positive values
|
|
mean that following Frontier produces slower real serving. Each objective is
|
|
selected independently; this table does not combine TTFT, TPOT, and E2E into a
|
|
single score.
|
|
|
|
## Qwen3-30B-A3B
|
|
|
|
Configuration surface: `TP in {1,2,4} x MNS in {8,16,32,64}`, with
|
|
`MBT=8192`. Each real cell uses three fresh-server trials.
|
|
|
|
| Workload | TTFT mean | TTFT p90 | TPOT mean | TPOT p90 | E2E mean | E2E p90 |
|
|
|---|---:|---:|---:|---:|---:|---:|
|
|
| Trace-PD | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% |
|
|
| Fixed-PD, 4096->256, 1.125 req/s/GPU | **58.0%** | **56.2%** | 0.0% | 0.0% | 1.7% | 5.5% |
|
|
| Trace-PO, OSL=1 | 3.2% | 0.4% | N/A | N/A | 3.2% | 0.3% |
|
|
| Fixed-PO, 4096->1, 1.125 req/s/GPU | 0.3% | 0.5% | N/A | N/A | 0.3% | 0.5% |
|
|
|
|
Interpretation: Frontier is near-optimal for Trace-PD and both prefill-only
|
|
cases, but the high-pressure Fixed-PD TTFT choice is materially wrong: its
|
|
selected configuration is 56--58% slower than the real TTFT optimum.
|
|
|
|
## Qwen3-235B-A22B-FP8
|
|
|
|
Configuration surface: `{TP4/EP1, TP8/EP8} x MNS in {64,128}`, with
|
|
`MBT=8192`. Each workload has 129 requests per cell and each real cell uses
|
|
three fresh-server trials.
|
|
|
|
| Workload | TTFT mean | TTFT p90 | TPOT mean | TPOT p90 | E2E mean | E2E p90 |
|
|
|---|---:|---:|---:|---:|---:|---:|
|
|
| Trace-PD | 0.0% | 0.0% | 0.0% | 0.0% | 0.6% | 6.2% |
|
|
| Fixed-PD, 4096->256, 0.2 req/s/GPU | 4.2% | 0.2% | **33.0%** | **37.2%** | **30.7%** | **34.6%** |
|
|
| Trace-PO, OSL=1 | 7.0% | **21.2%** | N/A | N/A | 7.0% | **21.2%** |
|
|
| Fixed-PO, 4096->1, 0.2 req/s/GPU | 5.9% | 1.7% | N/A | N/A | 5.9% | 1.7% |
|
|
|
|
Interpretation: Trace-PD is mostly near-optimal. Fixed-PD reverses the real
|
|
decode/E2E preference between the tested parallel configurations and incurs
|
|
31--37% regret. Trace-PO also has a material p90 failure of 21.2%.
|
|
|
|
## Decision
|
|
|
|
The tested Frontier stack has **not** solved serving configuration tuning.
|
|
Its selected configuration can be near-optimal for one workload and materially
|
|
wrong for another on the same model and hardware. The strongest current
|
|
counterexamples are Qwen3-30B Fixed-PD TTFT and Qwen3-235B Fixed-PD TPOT/E2E.
|
|
|
|
This statement is limited to the two tested MoE models and Frontier. It is not
|
|
yet evidence about dense models, Vidur/APEX as separately reproduced systems,
|
|
other hardware, or SLO-constrained tuning.
|
|
|
|
## Provenance
|
|
|
|
Primary immutable analysis artifacts on `dash0`:
|
|
|
|
- Qwen3-30B Trace-PD: `/home/admin/cpfs/wjh/aituner/graph-piecewise-qwen30-20260717/simulator-piecewise-surface-v2/analysis/comparison.json`
|
|
- Qwen3-30B Fixed-PD/PO: `/home/admin/cpfs/wjh/aituner/qwen30-fixed-pressure-surface-20260719-r1/analysis/`
|
|
- Qwen3-30B Trace-PO: `/home/admin/cpfs/wjh/aituner/qwen30-latency-expansion-20260718-r2/analysis-r6/trace-po-comparison.json`
|
|
- Qwen3-235B four-case matrix: `/home/admin/cpfs/wjh/aituner/qwen235-v020-fourcase-20260719-r1/analysis/comparison.json`
|
|
|
|
The Qwen3-235B artifact root includes `provenance/artifacts.sha256`; the final
|
|
matrix contains 48/48 valid real trials and 16/16 complete simulator cells.
|