AmphionASR | 热词、目标说话人、耳语——四种条件化 ASR 塞进一个 1.7B SpeechLLM
- URL: https://mp.weixin.qq.com/s/3JfnRFsG5BxrhJdBjDigAg
- Date Saved: 2026-09-01
- Source: WeChat (AutoML机器学习)
- Tags: asr, ai-engineering
- Repo: https://github.com/open-mmlab/Amphion (10.3k stars)
Summary
Deep technical analysis of AmphionASR (2026.07.28 tech report) — a 1.7B SpeechLLM (2.0B total with audio encoder) that unifies four conditioned ASR tasks into one model via prompt templates:
The Four Conditioned Tasks
- Hotword ASR — user provides a word list (names, terms), model biases toward them
- Target Speaker ASR (TS-ASR) — given 3-5s enrollment audio, only transcribe that speaker
- Degradation-robust ASR — handles far-field, noise, reverb, distortion, packet loss
- Whispered-speech ASR — handles whispered speech (no vocal cord vibration, different spectrum)
Architecture
- Base: Qwen3-ASR-1.7B (audio encoder 300M + LLM 1.7B)
- RAG retrieval branch for hotwords: only 2.6M additional parameters (two MLP adapters)
- All conditions encoded as prompt templates — no architecture changes needed
- Key innovation: max-over-time scoring for hotword retrieval (instead of utterance-level pooling), preserving sensitivity to short keyword spans
Training
- 3-stage LoRA SFT: Stage 1 (degradation+whisper, encoder unfrozen), Stage 2 (hotword+TS-ASR, encoder frozen), Stage 3 (replay for anti-forgetting)
- GRPO reinforcement learning on top
- TS-ASR training data is entirely synthetic with anti-hallucination negatives (teach model to output nothing when target speaker absent)
- Hotword lists “dirtied” during training: random dropout, hard negatives, random distractors
Key Results
- Whispered ASR: Chinese CER 0.58%, English WER 6.11% — first in both languages, clear win
- Hotword ASR: Entity error rate reduced ~44-55% with RAG retrieval; BUT without hotwords, worse than base model (23.18% vs 18.63% Chinese entity error)
- TS-ASR: WER 13.02% vs base’s 78.99%; false alarm dropped from 100% to 6.10% — capability created from scratch
- Degradation-robust: Won 8/16 subsets on Voice-in-the-Wild-Bench
- General ASR: 6/10 test sets show slight degradation (0.15-0.85 points)
Author’s Key Takeaways
- The real value of unification is engineering cost reduction (one model vs four pipelines), not per-task SOTA
- The w/o RAG regression is the most honest and informative number — multi-task SFT has a cost
- max-over-time scoring is broadly applicable beyond ASR (any retrieval where query matches a local segment)
- “Output nothing” must be explicitly trained — Qwen3-ASR had 100% false alarm on silence because it never saw empty labels
- Frame-level retriever was tested but NOT deployed to production — always read the limitations section
Open Questions
- Cross-condition behavior untested (hotword + target speaker simultaneously)
- Retrieval recall scaling beyond 10K vocabulary unknown
- GRPO contribution not isolated (no SFT-only ablation)