xserv/docs at 80157e614aeaaecc9031550dc8a0e9d4d4c726c8 - xserv - Local Gitea

gahow/xserv

Files

History

Gahow Wang 80157e614a docs: update llama.cpp comparison with 8192 results (OOM fixed)

Re-ran the full comparison at --max-seq-len 8192 now that xserv handles it:
- OOM finding resolved — pool sized to available VRAM + vLLM-style host swap;
  8192 runs with 0 swap events (swap is the overload safety net).
- Quality at parity with equal context: AIME 20.0% vs 20.0%, GSM8K 98% vs 96%.
- Speed unchanged relative to llama.cpp (~0.42-0.60x); TPOT is bandwidth-bound.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

2026-05-28 21:32:14 +08:00

..

docs: update llama.cpp comparison with 8192 results (OOM fixed)

2026-05-28 21:32:14 +08:00

00-roadmap.md

fix: comprehensive review + 14 bug fixes + Phase 12/14 overhaul

2026-05-22 17:53:28 +08:00

01-cuda-ffi.md

docs: add design docs + takeaways for Phase 2 and Phase 3

2026-05-21 20:59:45 +08:00

02-tensor.md

docs: add design docs + takeaways for Phase 2 and Phase 3

2026-05-21 20:59:45 +08:00

03-gemm.md

docs: add design docs + takeaways for Phase 2 and Phase 3

2026-05-21 20:59:45 +08:00

04-transformer-kernels.md

phase 4: transformer core kernels

2026-05-21 21:07:24 +08:00

05-attention.md

phase 5: naive multi-head attention

2026-05-21 21:17:23 +08:00

06-model-loading.md

phase 6+7+8: model loading, BPE tokenizer, GPT-2 inference (Milestone ①)

2026-05-21 22:04:00 +08:00

07-tokenizer.md

phase 6+7+8: model loading, BPE tokenizer, GPT-2 inference (Milestone ①)

2026-05-21 22:04:00 +08:00

08-gpt2.md

phase 6+7+8: model loading, BPE tokenizer, GPT-2 inference (Milestone ①)

2026-05-21 22:04:00 +08:00

09-kv-cache.md

phase 9: KV cache + autoregressive generation

2026-05-21 23:39:41 +08:00

10-qwen3.md

fix: comprehensive review + 14 bug fixes + Phase 12/14 overhaul

2026-05-22 17:53:28 +08:00

11-paged-attention.md

docs: Phase 14 design doc + benchmark, fix Phase 11/12 honesty

2026-05-22 18:51:29 +08:00

12-continuous-batching.md

docs: Phase 14 design doc + benchmark, fix Phase 11/12 honesty

2026-05-22 18:51:29 +08:00

13-http-api.md

docs: split Phase 12 and Phase 13 into separate design documents

2026-05-22 13:15:27 +08:00

14-flash-attention.md

docs: Phase 14 design doc + benchmark, fix Phase 11/12 honesty

2026-05-22 18:51:29 +08:00

15-performance.md

docs: Phase 15 design doc + benchmark report

2026-05-23 00:39:27 +08:00

16-llama-cpp-comparison.md

docs: update llama.cpp comparison with 8192 results (OOM fixed)

2026-05-28 21:32:14 +08:00

TO-BE-FIXED.md

fix: 12 bug fixes from comprehensive review — 51 tok/s verified on RTX 5090

2026-05-23 14:13:43 +08:00