Files
aituner/runs/frontier-code-trace-v0

Frontier code-trace campaign handoff

Phase A code prefill+decode 已进入 61min real matrix。data/profile、 max-length、Frontier rho calibration 和 TP2/TP4 paired canary 均已完成; 第一批 TP4 三个 load 与 TP2 low-rho diagnostic 正在 dash1dash4 并行 运行。code prefill-only 的独立 sim calibration 也已完成。

完整设计与 gate 见 experiment-card.md

当前资产与下一步

  • development window0513 [3480,7140)61min
  • held-out window0529 [2640,6240),只在 development 判据冻结后使用;
  • profileprofiles/profile-v6-code-longctx/,覆盖 TP1/2/4 和 131072 KV context
  • full paired inputsCPFS runs/frontier-code-trace-v0/inputs/full-r0p{0002,0004,0008,0016}-v1/
  • compact provenanceresults/calibration-summary.jsonresults/prefill-only-calibration-summary.jsonresults/canary-analysis-tp{2,4}-v*.jsonresults/paired-input-manifests/
  • 当前 A4 wave 1TP4 rho={0.0002,0.0008,0.0016} trial 1以及 TP2 rho=0.0002 trial 1 diagnostic
  • TP4 canary 的 TTFT/E2E、prefix hit 与 decode batch 通过TP2 TTFT p90 低估 32.1%,因此 TP2 其余 cell 暂不扩展;
  • prefill-onlyTP2 已冻结 rho={0.0004,0.0008,0.0016} TP4 到 0.0032 仍亚临界,需追加更高 rho 后冻结 near-knee。

source trace 的远端位置是:

/home/admin/cpfs/wjh/ali-trace/trace-glm5.1-formatted/

以下命令保留为从 source 重新构建时的复现入口。

1. 审计所有 1h+ code source

在持有 trace 的机器、repo 根目录执行:

python3 runs/frontier-code-trace-v0/audit_code_trace.py \
  --trace-root ~/ali-trace/trace-glm5.1-formatted \
  --output runs/frontier-code-trace-v0/inputs/code-audit.json

如果目录里混有非 request JSONL先只读列举文件再用多个 --source 显式指定。审计输出必须满足:

data_gate = PASS
selected.hash_contract.exact_source_block_size != null
max_model_len_recommendation != null
selected.selected_window_stats.max_model_len_coverage[推荐值].coverage = 1.0

全量审计已确认 source block size=512development source window 若 100% 覆盖需要 262144但正式 server cap 以 session-sampled paired cell 的实际 ISL+OSL max 向上对齐,不能把 full-window 262144 无条件套到低 rho cell。 审计会单独记录并排除 input_length<=0output_length<=0 的 source 行;这些行只有在 raw trace 同样显示 zero usage/empty response 时才按 “未发生模型执行”处理,不能无记录过滤。

2. 物化稳定窗口

python3 runs/frontier-code-trace-v0/prepare_code_window.py \
  --audit runs/frontier-code-trace-v0/inputs/code-audit.json \
  --output-root runs/frontier-code-trace-v0/inputs/code-window

输出是 6075min code-raw-window.jsonl 和 manifest。source 文件不修改。

3. 生成 P+D paired trace

若没有 prompt sidecar先生成 shape/prefix-faithful synthetic prompts

python3 runs/frontier-s3-real-v0/remap_hash_blocks.py \
  --input runs/frontier-code-trace-v0/inputs/code-window/code-raw-window.jsonl \
  --output-root runs/frontier-code-trace-v0/inputs/code-pd-rho-max \
  --source-block-size 512 \
  --workload-mode prefill_decode \
  --rho 1.0 \
  --max-total-tokens 131072 \
  --validate-parents

命令中的 512131072 必须替换为 audit manifest 值。若存在对齐 prompt sidecar--prompt--tokenizer,并要求 synthetic fallback 为 0。

正式 rho 不能直接用 1.0;先从最大 remap cache 按 session-coherent sampling_u 过滤,分别标定 low/mid/near-knee。

4. 生成 prefill-only paired trace

对 chat/code 使用同一个转换接口:

python3 runs/frontier-s3-real-v0/remap_hash_blocks.py \
  --input INPUT_WINDOW.jsonl \
  --output-root OUTPUT_ROOT \
  --source-block-size SOURCE_BLOCK_SIZE \
  --workload-mode prefill_only \
  --rho RHO \
  --max-total-tokens MAX_MODEL_LEN \
  --validate-parents

该模式会同时把 Frontier num_decode_tokens、real request min/max_tokens 和 remapped row 的 output_length 固定为 1。

5. max-model-len 真机 gate

现有 real runner 新增了三个显式环境变量chat 默认行为不变:

MAX_MODEL_LEN=ACTUAL_CELL_MAX_ROUNDED_UP \
TRACE_INPUT_ROOT=/absolute/path/to/materialized/code-cell \
ALLOW_SYNTHETIC_PROMPTS=true \
OUTPUT_ROOT=/absolute/path/to/new/output \
bash runs/frontier-s3-real-v0/run_full_real.sh RHO_LABEL tp4_mns16 1 PORT
  • MAX_MODEL_LEN 必须覆盖 manifest 中该 paired cell 的实际最大请求;
  • TRACE_INPUT_ROOT 内必须有 real_requests.jsonlmanifest.json
  • synthetic prompt 默认拒绝,只有在 experiment card 明确降级 claim 后才设为 true
  • runner 会在启动前扫描 paired requests若任何 ISL+OSL 超 cap 立即失败。

长上下文 server 必须同时设置 VLLM_ALLOW_LONG_MAX_MODEL_LEN=1--hf-overrides '{"max_position_embeddings":MAX_MODEL_LEN}'runner 已在 ALLOW_LONG_CONTEXT_SERVER=true 时自动处理。长上下文默认使用 host-local vLLM compile cache并按 topology 复用 FlashInfer workspace启动 compile 不进入 workload latency。

6. decode-only

当前 materializer 故意不提供 decode_only 选项。已安装 vLLM 0.20.0 包含 DecodeBenchConnector,但它在首次 admission 后同步填 dummy KV fill time 必须与 KV-ready arrival 分离。Frontier Request 支持 num_processed_tokens,当前 trace generator 尚未从 CSV 注入该值。 只有 real 首步无 prefill、sim ledger 首步为 decode 的 C0 gate 通过后, 才创建 strict decode-only jobs。

本地验证

python3 -m unittest -v \
  runs/frontier-code-trace-v0/test_code_trace_preflight.py \
  runs/frontier-s3-real-v0/test_remap_hash_blocks.py \
  runs/frontier-s3-real-v0/test_select_chat_window.py
python3 -m py_compile \
  runs/frontier-code-trace-v0/*.py \
  runs/frontier-s3-real-v0/*.py
bash -n runs/frontier-s3-real-v0/run_full_real.sh