WeChat (AutoML机器学习)
asrai-engineering
Original source

AmphionASR | 热词、目标说话人、耳语——四种条件化 ASR 塞进一个 1.7B SpeechLLM

Summary

Deep technical analysis of AmphionASR (2026.07.28 tech report) — a 1.7B SpeechLLM (2.0B total with audio encoder) that unifies four conditioned ASR tasks into one model via prompt templates:

The Four Conditioned Tasks

  • Hotword ASR — user provides a word list (names, terms), model biases toward them
  • Target Speaker ASR (TS-ASR) — given 3-5s enrollment audio, only transcribe that speaker
  • Degradation-robust ASR — handles far-field, noise, reverb, distortion, packet loss
  • Whispered-speech ASR — handles whispered speech (no vocal cord vibration, different spectrum)

Architecture

  • Base: Qwen3-ASR-1.7B (audio encoder 300M + LLM 1.7B)
  • RAG retrieval branch for hotwords: only 2.6M additional parameters (two MLP adapters)
  • All conditions encoded as prompt templates — no architecture changes needed
  • Key innovation: max-over-time scoring for hotword retrieval (instead of utterance-level pooling), preserving sensitivity to short keyword spans

Training

  • 3-stage LoRA SFT: Stage 1 (degradation+whisper, encoder unfrozen), Stage 2 (hotword+TS-ASR, encoder frozen), Stage 3 (replay for anti-forgetting)
  • GRPO reinforcement learning on top
  • TS-ASR training data is entirely synthetic with anti-hallucination negatives (teach model to output nothing when target speaker absent)
  • Hotword lists “dirtied” during training: random dropout, hard negatives, random distractors

Key Results

  • Whispered ASR: Chinese CER 0.58%, English WER 6.11% — first in both languages, clear win
  • Hotword ASR: Entity error rate reduced ~44-55% with RAG retrieval; BUT without hotwords, worse than base model (23.18% vs 18.63% Chinese entity error)
  • TS-ASR: WER 13.02% vs base’s 78.99%; false alarm dropped from 100% to 6.10% — capability created from scratch
  • Degradation-robust: Won 8/16 subsets on Voice-in-the-Wild-Bench
  • General ASR: 6/10 test sets show slight degradation (0.15-0.85 points)

Author’s Key Takeaways

  1. The real value of unification is engineering cost reduction (one model vs four pipelines), not per-task SOTA
  2. The w/o RAG regression is the most honest and informative number — multi-task SFT has a cost
  3. max-over-time scoring is broadly applicable beyond ASR (any retrieval where query matches a local segment)
  4. “Output nothing” must be explicitly trained — Qwen3-ASR had 100% false alarm on silence because it never saw empty labels
  5. Frame-level retriever was tested but NOT deployed to production — always read the limitations section

Open Questions

  • Cross-condition behavior untested (hotword + target speaker simultaneously)
  • Retrieval recall scaling beyond 10K vocabulary unknown
  • GRPO contribution not isolated (no SFT-only ablation)