Skip to content

Commit a078fe3

Browse files
authored
Merge PR #3687: native Transformers adoption documentation
Signed-off-by: zhifu gao <zhifu.gzf@alibaba-inc.com>
2 parents 3ea977f + 3a8a2d0 commit a078fe3

38 files changed

Lines changed: 413 additions & 326 deletions

‎README.md‎

Lines changed: 10 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -29,6 +29,14 @@
2929

3030
## Quick Start
3131

32+
### Native Transformers
33+
34+
For Fun-ASR-Nano transcription with the Hugging Face API, start with the [Transformers 5.17.0 CPU quickstart](./docs/transformers_native.md). No FunASR toolkit or remote Python code is needed.
35+
36+
[Space](https://huggingface.co/spaces/FunAudioLLM/Fun-ASR-Nano) · [Notebook](https://colab.research.google.com/github/QwenAudio/Fun-ASR/blob/main/examples/colab/fun_asr_nano_transformers.ipynb) · [Python / batch examples](https://github.com/QwenAudio/Fun-ASR/tree/main/examples/transformers)
37+
38+
### FunASR toolkit and pipelines
39+
3240
[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/modelscope/FunASR/blob/main/examples/colab/funasr_quickstart.ipynb)
3341

3442
No local setup? Open the [Colab quickstart](./examples/colab/) to transcribe a public sample or upload your own audio in a browser.
@@ -55,7 +63,7 @@ PY
5563
Only use `device="cuda"` when this prints `True`; otherwise use `device="cpu"`
5664
or reinstall PyTorch with the correct CUDA wheel.
5765

58-
**Flagship model — Fun-ASR-Nano** (LLM-ASR for Chinese, English, and Japanese, plus Chinese dialect groups and regional accents; needs a GPU):
66+
**FunASR toolkit GPU example: Fun-ASR-Nano** (Chinese, English, Japanese, and Chinese dialect groups and regional accents; the separate native Transformers CPU path is linked above):
5967

6068
```python
6169
from funasr import AutoModel
@@ -357,7 +365,7 @@ recordings with the same evaluation scope.
357365

358366
- **MOSS-Transcribe-Diarize** brings long-form ASR, timestamps, and anonymous speaker labels to FunASR services, Docker, Kubernetes, vLLM/SGLang workflows, and FunClip. [Deploy MOSS ->](./docs/moss_transcribe_diarize.md)
359367
- **FunASR 1.4.15** adds tested NumPy 2 compatibility and fixes streaming KWS/VAD boundaries and checkpoint ranking. Install with `python -m pip install -U "funasr==1.4.15"`. [Release and verification scope ->](https://github.com/modelscope/FunASR/releases/tag/v1.4.15)
360-
- **Native Transformers:** Fun-ASR-Nano [merged upstream](https://github.com/huggingface/transformers/pull/46180). Use the official `-hf` checkpoint and pinned source; stable 5.16.1 does not include it. [Installation and inference ->](./docs/transformers_native.md)
368+
- **Native Transformers:** Released **5.17.0** supports Fun-ASR-Nano with the official `-hf` checkpoint, CPU examples and a notebook. [Get started ->](./docs/transformers_native.md)
361369

362370
> See [GitHub Releases](https://github.com/modelscope/FunASR/releases) for the complete changelog and downloadable assets.
363371

‎README_ja.md‎

Lines changed: 9 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -30,6 +30,14 @@
3030

3131
## クイックスタート
3232

33+
### ネイティブ Transformers
34+
35+
Hugging Face API で Fun-ASR-Nano を使う場合は [Transformers 5.17.0 CPU ガイド(英語)](./docs/transformers_native.md) から開始できます。FunASR toolkit とリモート Python コードは不要です。
36+
37+
[Space](https://huggingface.co/spaces/FunAudioLLM/Fun-ASR-Nano) · [Notebook](https://colab.research.google.com/github/QwenAudio/Fun-ASR/blob/main/examples/colab/fun_asr_nano_transformers.ipynb) · [Python / batch examples](https://github.com/QwenAudio/Fun-ASR/tree/main/examples/transformers)
38+
39+
### FunASR toolkit とパイプライン
40+
3341
```bash
3442
python -m pip install torch torchaudio
3543
python -m pip install funasr
@@ -105,7 +113,7 @@ CER/WER をそろえて比較してください。オフラインのスループ
105113

106114
- **MOSS-Transcribe-Diarize** を FunASR service、Docker、Kubernetes、vLLM/SGLang workflow、FunClip に統合し、長時間 ASR、timestamp、匿名 speaker label を一度に処理できます。[MOSS をデプロイ ->](./docs/moss_transcribe_diarize.md)
107115
- **FunASR 1.4.15** はテスト済みの NumPy 2 互換性を追加し、ストリーミング KWS/VAD の境界処理と checkpoint の順位付けを修正します。`python -m pip install -U "funasr==1.4.15"`。[リリースと検証範囲 ->](https://github.com/modelscope/FunASR/releases/tag/v1.4.15)
108-
- **ネイティブ Transformers:** Fun-ASR-Nano が[上流にマージ](https://github.com/huggingface/transformers/pull/46180)されました。公式 `-hf` checkpoint と固定ソースを使用します。安定版 5.16.1 には未収録です。[導入ガイド(英語) ->](./docs/transformers_native.md)
116+
- **ネイティブ Transformers:** 正式版 **5.17.0** が Fun-ASR-Nano に対応。公式 `-hf` checkpoint、CPU サンプル、Notebook:[導入ガイド(英語) ->](./docs/transformers_native.md)
109117

110118
> 完全な変更履歴と download asset は [GitHub Releases](https://github.com/modelscope/FunASR/releases) を参照してください。
111119

‎README_ko.md‎

Lines changed: 9 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -30,6 +30,14 @@
3030

3131
## 빠른 시작
3232

33+
### 네이티브 Transformers
34+
35+
Hugging Face API로 Fun-ASR-Nano를 사용하려면 [Transformers 5.17.0 CPU 가이드(영문)](./docs/transformers_native.md)에서 시작하세요. FunASR toolkit이나 원격 Python 코드가 필요 없습니다.
36+
37+
[Space](https://huggingface.co/spaces/FunAudioLLM/Fun-ASR-Nano) · [Notebook](https://colab.research.google.com/github/QwenAudio/Fun-ASR/blob/main/examples/colab/fun_asr_nano_transformers.ipynb) · [Python / batch examples](https://github.com/QwenAudio/Fun-ASR/tree/main/examples/transformers)
38+
39+
### FunASR toolkit 및 파이프라인
40+
3341
```bash
3442
python -m pip install torch torchaudio
3543
python -m pip install funasr
@@ -104,7 +112,7 @@ checkpoint/revision, 오디오 집합, 하드웨어, 배치, 워밍업, 측정
104112

105113
- **MOSS-Transcribe-Diarize**를 FunASR service, Docker, Kubernetes, vLLM/SGLang workflow, FunClip에 통합해 긴 오디오 ASR, timestamp, 익명 speaker label을 한 번에 처리합니다. [MOSS 배포 ->](./docs/moss_transcribe_diarize.md)
106114
- **FunASR 1.4.15**는 테스트를 거친 NumPy 2 호환성을 추가하고 스트리밍 KWS/VAD 경계 처리와 checkpoint 순위 산정을 수정합니다. `python -m pip install -U "funasr==1.4.15"`. [릴리스 및 검증 범위 ->](https://github.com/modelscope/FunASR/releases/tag/v1.4.15)
107-
- **네이티브 Transformers:** Fun-ASR-Nano가 [업스트림에 병합](https://github.com/huggingface/transformers/pull/46180)되었습니다. 공식 `-hf` checkpoint와 고정 소스를 사용하세요. 안정 버전 5.16.1에는 아직 포함되지 않았습니다. [설치 가이드(영문) ->](./docs/transformers_native.md)
115+
- **네이티브 Transformers:** 정식 버전 **5.17.0**이 Fun-ASR-Nano를 지원합니다. 공식 `-hf` checkpoint, CPU 예제, Notebook: [설치 가이드(영문) ->](./docs/transformers_native.md)
108116

109117
> 전체 변경 기록과 download asset은 [GitHub Releases](https://github.com/modelscope/FunASR/releases)에서 확인할 수 있습니다.
110118

‎README_zh.md‎

Lines changed: 9 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -28,6 +28,14 @@
2828

2929
## 快速开始
3030

31+
### 原生 Transformers
32+
33+
使用 Hugging Face API 转写 Fun-ASR-Nano,先看 [Transformers 5.17.0 CPU 快速开始](./docs/transformers_native_zh.md),不需要安装 FunASR 工具库或执行远程 Python 代码。
34+
35+
[Space](https://huggingface.co/spaces/FunAudioLLM/Fun-ASR-Nano) · [Notebook](https://colab.research.google.com/github/QwenAudio/Fun-ASR/blob/main/examples/colab/fun_asr_nano_transformers.ipynb) · [Python / batch examples](https://github.com/QwenAudio/Fun-ASR/tree/main/examples/transformers)
36+
37+
### FunASR 工具库与流水线
38+
3139
[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/modelscope/FunASR/blob/main/examples/colab/funasr_quickstart.ipynb)
3240

3341
不想先配置本地环境?可以打开 [Colab 快速体验](./examples/colab/README_zh.md) 在浏览器里转写公开样例或上传自己的音频。
@@ -150,7 +158,7 @@ checkpoint/revision、音频集、硬件、批量大小、预热、计时范围
150158

151159
- **MOSS-Transcribe-Diarize** 已接入 FunASR 服务、Docker、Kubernetes、vLLM/SGLang 工作流和 FunClip,一次完成长音频转写、时间戳与匿名说话人标注。[部署 MOSS ->](./docs/moss_transcribe_diarize_zh.md)
152160
- **FunASR 1.4.15** 新增经过测试的 NumPy 2 兼容支持,修复流式 KWS/VAD 边界处理和 checkpoint 排序。升级命令:`python -m pip install -U "funasr==1.4.15"`。[发布说明与验证范围 ->](https://github.com/modelscope/FunASR/releases/tag/v1.4.15)
153-
- **原生 Transformers:** Fun-ASR-Nano [已合入上游](https://github.com/huggingface/transformers/pull/46180)。使用官方 `-hf` checkpoint 与固定源码;稳定版 5.16.1 尚未包含。[安装与推理 ->](./docs/transformers_native_zh.md)
161+
- **原生 Transformers:** 正式版 **5.17.0** 已支持 Fun-ASR-Nano。官方 `-hf` 权重、CPU 示例与 Notebook:[安装与推理 ->](./docs/transformers_native_zh.md)
154162

155163
> 完整改动记录和可下载资产请查看 [GitHub Releases](https://github.com/modelscope/FunASR/releases)。
156164

‎docs/deployment_matrix.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -6,7 +6,7 @@ Use this page to choose the shortest deployment path for a product, demo, benchm
66

77
| Path | Best for | Start here | Operational notes |
88
|---|---|---|---|
9-
| Native Transformers | Python applications using Hugging Face processors and generation | [Native Fun-ASR-Nano guide](./transformers_native.md) | Official `-hf` checkpoint and pinned source; stable 5.16.1 lacks support at the 2026-09-09 check. Python inference, not an HTTP or realtime server. |
9+
| Native Transformers | First Nano transcript, notebooks and Hugging Face Python applications | [Native Fun-ASR-Nano guide](./transformers_native.md) | Released **5.17.0**, official `-hf` checkpoint, CPU examples and batching. Python inference, not an HTTP or realtime server. |
1010
| Colab notebook | Browser smoke tests, first evaluation, shareable demos | [Colab quickstart](../examples/colab/) | No local setup; first run downloads model files, GPU runtime is faster. |
1111
| Python API | Notebooks, offline jobs, first model evaluation | [README quick start](../README.md#quick-start) | Lowest ceremony; caller owns batching, retries, and files. |
1212
| OpenAI-compatible API | Private speech API, agents, Dify/LangChain/AutoGen-style clients | [OpenAI API example](../examples/openai_api/) | Easiest integration for apps that already support OpenAI audio APIs. |

‎docs/deployment_matrix_ja.md‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,9 @@
11
# FunASR デプロイ選択マトリクス
22

3+
## Transformers で Nano を試す
4+
5+
中国語・英語・日本語の文字起こしには [Transformers 5.17.0 ガイド(英語)](./transformers_native.md) と公式 `FunAudioLLM/Fun-ASR-Nano-2512-hf` checkpoint を使えます。CPU サンプルがあり、toolkit とサービスの経路は別です。ネイティブ出力はテキストで、タイムスタンプ・話者・HTTP サーバーを追加しません。
6+
37
プロダクト、デモ、ベンチマーク、社内ワークフローに合わせて最短のデプロイ経路を選ぶためのガイドです。まずは要件を満たす最小構成から始め、throughput、latency、integration 要件が明確になったら重い runtime に移行してください。
48

59
## クイック判断表

‎docs/deployment_matrix_ko.md‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,9 @@
11
# FunASR 배포 선택 매트릭스
22

3+
## Transformers로 Nano 시작하기
4+
5+
중국어·영어·일본어 전사는 [Transformers 5.17.0 가이드(영문)](./transformers_native.md)와 공식 `FunAudioLLM/Fun-ASR-Nano-2512-hf` checkpoint로 시작할 수 있습니다. CPU 예제가 있으며 toolkit과 서비스 경로는 별도입니다. 네이티브 출력은 텍스트이며 타임스탬프·화자·HTTP 서버를 추가하지 않습니다.
6+
37
제품, 데모, 벤치마크, 내부 워크플로에 맞는 가장 짧은 배포 경로를 고르기 위한 가이드입니다. 먼저 요구를 만족하는 최소 구성에서 시작하고, throughput, latency, integration 요구가 명확해질 때 더 무거운 runtime으로 이동하세요.
48

59
## 빠른 결정 표

‎docs/deployment_matrix_zh.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -6,7 +6,7 @@
66

77
| 路径 | 适合场景 | 从这里开始 | 运维提示 |
88
|---|---|---|---|
9-
| 原生 Transformers | 已使用 Hugging Face 处理器与生成接口的 Python 应用 | [原生 Fun-ASR-Nano 指南](./transformers_native_zh.md) | 官方 `-hf` checkpoint 与固定源码;2026-09-09 核验的稳定版 5.16.1 尚未包含。是 Python 推理入口,不是 HTTP 或实时服务器。 |
9+
| 原生 Transformers | Nano 首次转写、Notebook 与 Hugging Face Python 应用 | [原生 Fun-ASR-Nano 指南](./transformers_native_zh.md) | 正式版 **5.17.0**,官方 `-hf` checkpoint,CPU 与批处理示例。是 Python 推理入口,不是 HTTP 或实时服务器。 |
1010
| Colab Notebook | 浏览器 smoke test、首次评估、可分享 demo | [Colab 快速体验](../examples/colab/README_zh.md) | 不需要本地环境;首次运行会下载模型,GPU runtime 更快。 |
1111
| Python API | Notebook、离线任务、首次模型评测 | [README 快速开始](../README_zh.md#快速开始) | 最简单;调用方自己负责批处理、重试和文件管理。 |
1212
| OpenAI 兼容 API | 私有语音 API、Agent、Dify/LangChain/AutoGen 风格客户端 | [OpenAI API 示例](../examples/openai_api/README_zh.md) | 已支持 OpenAI audio API 的应用最容易接入。 |

‎docs/index.rst‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -12,6 +12,8 @@ entry to these guides. Source Markdown remains in this repository.
1212
:maxdepth: 1
1313
:caption: Get Started
1414

15+
transformers_native
16+
transformers_native_zh
1517
installation/installation
1618
installation/installation_zh
1719
installation/docker

‎docs/installation/installation.md‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -2,6 +2,8 @@
22

33
# Install the Python SDK
44

5+
> **Only need native Fun-ASR-Nano transcription?** Start with [Transformers 5.17.0](../transformers_native.md). It loads the separate `-hf` checkpoint without the FunASR toolkit. This page covers the `funasr.AutoModel` toolkit path; do not mix dependencies, parameters or output contracts.
6+
57
Use this guide for `from funasr import AutoModel`. For a packaged C++ service, start with [Docker and runtime images](./docker.md). After installation, continue to the [SDK tutorial](../tutorial/README.md).
68

79
## 1. Create an isolated environment

0 commit comments

Comments
 (0)