GitHub
asrai-toolsdev-tools
Original source

FunASR — Industrial Speech Recognition Toolkit

Summary

Open-source speech recognition toolkit from ModelScope/Alibaba DAMO Academy. Not a single model but a full toolkit — pick the right model per job, compose pipelines (ASR + VAD + punctuation + speaker diarization), and deploy via OpenAI-compatible API or edge binary.

Key Models

  • Fun-ASR-Nano (800M) — flagship LLM-ASR for zh/en/ja + Chinese dialects, GPU required
  • Fun-ASR-MLT-Nano (800M) — 31 languages
  • SenseVoiceSmall (234M) — 5-language ASR + emotion recognition + audio events, CPU-viable (17x realtime)
  • Paraformer-zh (220M) — low-latency streaming ASR for zh/en
  • Qwen3-ASR (1.7B) — 52 languages (the base model that AmphionASR builds on)
  • Also wraps Whisper-large-v3, GLM-ASR-Nano, emotion2vec

Why It Matters

  • 340x realtime with Fun-ASR-Nano + vLLM batch processing
  • Full pipeline composition: VAD (fsmn-vad) + ASR + punctuation (ct-punc) + speaker diarization (CAM++) in one AutoModel call
  • CPU/edge deployment via GGUF format (like whisper.cpp but ~3x lower CER on Chinese)
  • OpenAI-compatible API server: funasr-server --device cuda
  • MCP server integration for AI agents (Claude/Cursor)
  • Streaming WebSocket support via Paraformer
  • CLI tool: funasr audio.wav --spk --timestamps -f json

Deployment Options

  • Python pip package (pip install funasr)
  • vLLM acceleration for batch processing
  • Docker containers for streaming service
  • GGUF binary for CPU/edge (no Python needed, supports Vulkan GPU)
  • OpenAI-compatible REST API

Context

  • Very actively maintained (v1.4.11 released 2026-08-30, commits daily)
  • This is the FunASR-Realtime that came within 0.52 points of AmphionASR on Chinese entity recognition — without any hotword input
  • Practical choice for self-hosted ASR: mature toolkit, multiple model options, production-ready serving