Files
aituner/runs/frontier-code-trace-v0/README.md

150 lines
5.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Frontier code-trace campaign handoff
Phase A code prefill+decode 已开始执行。data/profile/max-length gate 和
Frontier rho calibration 已完成TP2/TP4 paired 10min real canary 正在
运行。61min full paired inputs 与 code prefill-only calibration cache
均已物化,尚未越过 canary gate 启动正式 1h real matrix。
完整设计与 gate 见 [`experiment-card.md`](experiment-card.md)。
## 当前资产与下一步
- development window0513 `[3480,7140)`61min
- held-out window0529 `[2640,6240)`,只在 development 判据冻结后使用;
- profile`profiles/profile-v6-code-longctx/`,覆盖 TP1/2/4 和 131072
KV context
- full paired inputsCPFS
`runs/frontier-code-trace-v0/inputs/full-r0p{0002,0004,0008,0016}-v1/`
- compact provenance`results/calibration-summary.json`
`results/paired-input-manifests/`
- 下一步:完成 paired canary 分析;若 real 零失败、digest/hit ratio
对齐且 queue 不发散,启动 TP4
`rho={0.0002,0.0008,0.0016}` 与 TP2
`rho={0.0002,0.0004,0.0008}` 的 1h jobs。
source trace 的远端位置是:
```text
/home/admin/cpfs/wjh/ali-trace/trace-glm5.1-formatted/
```
以下命令保留为从 source 重新构建时的复现入口。
## 1. 审计所有 1h+ code source
在持有 trace 的机器、repo 根目录执行:
```bash
python3 runs/frontier-code-trace-v0/audit_code_trace.py \
--trace-root ~/ali-trace/trace-glm5.1-formatted \
--output runs/frontier-code-trace-v0/inputs/code-audit.json
```
如果目录里混有非 request JSONL先只读列举文件再用多个 `--source` 显式指定。审计输出必须满足:
```text
data_gate = PASS
selected.hash_contract.exact_source_block_size != null
max_model_len_recommendation != null
selected.selected_window_stats.max_model_len_coverage[推荐值].coverage = 1.0
```
全量审计已确认 source block size=512development source window 若 100%
覆盖需要 262144但正式 server cap 以 session-sampled paired cell 的实际
`ISL+OSL max` 向上对齐,不能把 full-window 262144 无条件套到低 rho cell。
审计会单独记录并排除 `input_length<=0``output_length<=0` 的 source
行;这些行只有在 raw trace 同样显示 zero usage/empty response 时才按
“未发生模型执行”处理,不能无记录过滤。
## 2. 物化稳定窗口
```bash
python3 runs/frontier-code-trace-v0/prepare_code_window.py \
--audit runs/frontier-code-trace-v0/inputs/code-audit.json \
--output-root runs/frontier-code-trace-v0/inputs/code-window
```
输出是 6075min `code-raw-window.jsonl` 和 manifest。source 文件不修改。
## 3. 生成 P+D paired trace
若没有 prompt sidecar先生成 shape/prefix-faithful synthetic prompts
```bash
python3 runs/frontier-s3-real-v0/remap_hash_blocks.py \
--input runs/frontier-code-trace-v0/inputs/code-window/code-raw-window.jsonl \
--output-root runs/frontier-code-trace-v0/inputs/code-pd-rho-max \
--source-block-size 512 \
--workload-mode prefill_decode \
--rho 1.0 \
--max-total-tokens 131072 \
--validate-parents
```
命令中的 `512``131072` 必须替换为 audit manifest 值。若存在对齐 prompt sidecar`--prompt``--tokenizer`,并要求 synthetic fallback 为 0。
正式 rho 不能直接用 1.0;先从最大 remap cache 按 session-coherent `sampling_u` 过滤,分别标定 low/mid/near-knee。
## 4. 生成 prefill-only paired trace
对 chat/code 使用同一个转换接口:
```bash
python3 runs/frontier-s3-real-v0/remap_hash_blocks.py \
--input INPUT_WINDOW.jsonl \
--output-root OUTPUT_ROOT \
--source-block-size SOURCE_BLOCK_SIZE \
--workload-mode prefill_only \
--rho RHO \
--max-total-tokens MAX_MODEL_LEN \
--validate-parents
```
该模式会同时把 Frontier `num_decode_tokens`、real request `min/max_tokens` 和 remapped row 的 `output_length` 固定为 1。
## 5. max-model-len 真机 gate
现有 real runner 新增了三个显式环境变量chat 默认行为不变:
```bash
MAX_MODEL_LEN=ACTUAL_CELL_MAX_ROUNDED_UP \
TRACE_INPUT_ROOT=/absolute/path/to/materialized/code-cell \
ALLOW_SYNTHETIC_PROMPTS=true \
OUTPUT_ROOT=/absolute/path/to/new/output \
bash runs/frontier-s3-real-v0/run_full_real.sh RHO_LABEL tp4_mns16 1 PORT
```
- `MAX_MODEL_LEN` 必须覆盖 manifest 中该 paired cell 的实际最大请求;
- `TRACE_INPUT_ROOT` 内必须有 `real_requests.jsonl``manifest.json`
- synthetic prompt 默认拒绝,只有在 experiment card 明确降级 claim 后才设为 `true`
- runner 会在启动前扫描 paired requests若任何 `ISL+OSL` 超 cap 立即失败。
长上下文 server 必须同时设置
`VLLM_ALLOW_LONG_MAX_MODEL_LEN=1`
`--hf-overrides '{"max_position_embeddings":MAX_MODEL_LEN}'`runner 已在
`ALLOW_LONG_CONTEXT_SERVER=true` 时自动处理。长上下文默认使用 host-local
vLLM compile cache并按 topology 复用 FlashInfer workspace启动 compile
不进入 workload latency。
## 6. decode-only
当前 materializer 故意不提供 `decode_only` 选项。已安装 vLLM 0.20.0
包含 `DecodeBenchConnector`,但它在首次 admission 后同步填 dummy KV
fill time 必须与 KV-ready arrival 分离。Frontier `Request` 支持
`num_processed_tokens`,当前 trace generator 尚未从 CSV 注入该值。
只有 real 首步无 prefill、sim ledger 首步为 decode 的 C0 gate 通过后,
才创建 strict decode-only jobs。
## 本地验证
```bash
python3 -m unittest -v \
runs/frontier-code-trace-v0/test_code_trace_preflight.py \
runs/frontier-s3-real-v0/test_remap_hash_blocks.py \
runs/frontier-s3-real-v0/test_select_chat_window.py
python3 -m py_compile \
runs/frontier-code-trace-v0/*.py \
runs/frontier-s3-real-v0/*.py
bash -n runs/frontier-s3-real-v0/run_full_real.sh
```