AI Muninn (coolthor)
local-llmdev-tools
Original source

Chat Template + MTP 深度調整:免費提速本地 LLM

Summary

兩個零成本改動讓 RTX 2080 Ti 上的 Qwen3.8-27B 從 41.6 tok/s 提升到 47.6 tok/s,更重要的是 64K 長對話從每輪等 163 秒變成秒回。模型不動、卡不動、量化不動。

兩個改動

1. 換社群修復版 Chat Template

  • GGUF 內嵌的 template 是轉檔者當時抄的,不會跟上游更新
  • 舊版三個病:reasoning_effort 預設 xhigh(狂燒 token)、effort 名稱白名單不認 “high” 就炸、缺歷史 處理
  • 缺 think 處理 → 每輪渲染前綴不一致 → KV prefix cache 全部失效 → 整段重算
  • 修復版:froggeric/Qwen-Fixed-Chat-Templates (Apache-2.0)
  • 效果:KV cache 命中 94.5%(理論滿分)

2. MTP 投機解碼深度從 2 開到 4

  • Qwen3.8 MTP head 訓練深度是 7,預設只用 2-3 格
  • 深度 2:接受率 1.00,每步產出 3.0(考卷太簡單)
  • 深度 4(甜點):接受率 0.99,每步產出 4.95
  • 深度 5:接受率 0.95,邊際遞減
  • 深度 6+:MTP buffer 需多 260MB VRAM,開不了機

診斷方法

curl -s http://localhost:8082/props | python3 -c "import json,sys; print(json.load(sys.stdin)['chat_template'])" > current-template.jinja
grep -n "xhigh" current-template.jinja         # 指紋一
grep -n "raise_exception" current-template.jinja # 指紋二
grep -c "think" current-template.jinja          # 指紋三(<10 = 舊版)

啟動參數

llama-server -m model.gguf --jinja \
  --chat-template-file ./chat_template.jinja \
  --reasoning-format deepseek \
  --spec-type draft-mtp --spec-draft-n-max 4

注:--reasoning-format deepseek 選的是 標籤格式(DeepSeek R1 帶起的寫法),Qwen 用同一套標籤。

踩坑筆記

  • 深度 6 不是「變慢」是「開不了機」(VRAM OOM 在啟動時)
  • crash-loop 中的殭屍進程能回 health check 但不能生成
  • 存活檢查要驗「能出字」,不是「能回應」
  • 探針數字 vs 真實流量:實戰接受率約 0.67-0.78,打八折估算