Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,208 @@
"""Client Zonos v0.1-transformer pour Phase A0 #17586 bakeoff_large.

Modèle : Zyphra/Zonos-v0.1-transformer (Apache-2.0, FR par code eSpeak
`fr-fr`, clonage de voix par embedding de référence).

Installation (env propre, règle F) — ATTENTION : le paquet PyPI `zonos` est un
placeholder squatté (0.1.0.dev0, wheel vide, sans module) ; le paquet officiel
s'installe depuis le repo GitHub Zyphra/Zonos :
py -3.10 -m venv _runtime/venv-zonos
_runtime/venv-zonos/Scripts/python.exe -m pip install --upgrade pip
_runtime/venv-zonos/Scripts/python.exe -m pip install torch==2.11.0 torchaudio==2.11.0 --index-url https://download.pytorch.org/whl/cu126
git clone --depth 1 https://github.com/Zyphra/Zonos.git _runtime/Zonos
_runtime/venv-zonos/Scripts/python.exe -m pip install -e ./_runtime/Zonos soundfile faster-whisper

Variante TRANSFORMER (pas hybride) : la variante hybride
(`Zyphra/Zonos-v0.1-hybrid`) exige mamba-ssm + causal-conv1d (extras `compile`,
noyaux CUDA compilés, sans wheels Windows) ; la variante pure transformer est
un checkpoint officiel du même repo, sans dépendance CUDA compilée — choix de
variante documenté, pas un contournement.

Interface uniforme (cf. __init__.py) :
load_model(device, dtype) -> model
synth(text, out_wav, model, language='French', **kw) -> dict

- language : 'French' (mappé vers le code eSpeak 'fr-fr' de Zonos ; les
autres langues Zonos passent aussi par leur nom anglais)
- speaker/instruct : non applicables à Zonos v0.1 (ignorés) — le timbre
vient du wav de référence (clonage), la langue du code ISO.

Voix de référence : asset zh `zero_shot_prompt.wav` du repo CosyVoice du
_runtime (même convention que le client cosyvoice3 — timbre identique entre
clients du bakeoff, clonage cross-lingual zh->FR, capacité sous test).

CLI :
python clients/zonos.py --text "..." --out <wav> [--language French] [--device cuda]
"""
from __future__ import annotations

import argparse
import json
import time
from pathlib import Path

import torch
import torchaudio

import soundfile as sf


DEFAULT_MODEL_ID = "Zyphra/Zonos-v0.1-transformer"
DEFAULT_LANGUAGE = "French"

# Nom bench -> code eSpeak conditionné par Zonos (make_cond_dict language=)
_LANGUAGE_CODES = {
"english": "en-us",
"french": "fr-fr",
"german": "de",
"spanish": "es",
"italian": "it",
"portuguese": "pt",
"polish": "pl",
"dutch": "nl",
"russian": "ru",
}


def _reference_wav() -> Path:
"""Asset de référence (voix zh) : runtime CosyVoice du worktree voisin,
sinon _runtime local. Identique entre clients du bakeoff."""
here = Path(__file__).resolve()
for _ in range(10):
if (here / "_runtime").exists():
break
here = here.parent
candidates = [
here / "_runtime" / "CosyVoice" / "asset" / "zero_shot_prompt.wav",
here.parent / "CoursIA-17586-cosyvoice3" / "_runtime" / "CosyVoice" / "asset" / "zero_shot_prompt.wav",
here.parent.parent / "CoursIA-17586-cosyvoice3" / "_runtime" / "CosyVoice" / "asset" / "zero_shot_prompt.wav",
]
for c in candidates:
if c.exists():
return c
return candidates[0]


def get_supported_languages() -> list[str]:
return [k.capitalize() for k in _LANGUAGE_CODES]


def load_model(model_id: str = DEFAULT_MODEL_ID, device: str = "cuda", dtype: str = "bf16"):
"""Charge Zonos transformer avec mesure VRAM pic.

Le paramètre `dtype` du banc est non applicable : `from_pretrained` force
le backbone en bfloat16 (`.to(device, torch.bfloat16)` dans le repo) —
accepté ici seulement pour l'interface uniforme.
"""
from zonos.model import Zonos

print(f" [Zonos] Loading {model_id} on {device} (bf16, forced by loader)...")
if torch.cuda.is_available():
torch.cuda.reset_peak_memory_stats()
t0 = time.time()
model = Zonos.from_pretrained(model_id, device=device)
dt = time.time() - t0
vram_peak_gb = torch.cuda.max_memory_allocated() / 1024**3 if torch.cuda.is_available() else 0.0
print(f" [Zonos] Loaded in {dt:.1f}s, VRAM peak {vram_peak_gb:.2f} GB")
return model


def synth(
text: str,
out_wav: str,
model,
language: str = DEFAULT_LANGUAGE,
speaker: str | None = None, # non applicable Zonos (clonage par référence)
instruct: str | None = None, # non applicable Zonos v0.1
**kwargs,
) -> dict:
"""Synthèse text->wav par clonage cross-lingual de l'asset de référence."""
from zonos.conditioning import make_cond_dict

lang_code = _LANGUAGE_CODES.get(language.strip().lower())
if lang_code is None:
raise ValueError(
f"Langue {language!r} non supportée par le client zonos "
f"(supportées : {', '.join(sorted(set(_LANGUAGE_CODES)))})"
)

ref_path = _reference_wav()
if not ref_path.exists():
raise FileNotFoundError(
f"Voix de référence absente : {ref_path} — cloner le repo CosyVoice "
"(git clone --recursive https://github.com/FunAudioLLM/CosyVoice.git) "
"ou poser un wav de référence à cet emplacement."
)
# torchaudio 2.11 route load() sur torchcodec (non installe) — soundfile,
# deja dans le venv, lit le wav sans dependance supplementaire.
data, sr = sf.read(str(ref_path), dtype="float32", always_2d=True)
spkref = torch.from_numpy(data.T) # (channels, samples)
# Speaker = embedding calculé par le modèle (make_speaker_embedding),
# pas le wav brut — cf sample.py du repo officiel.
# Bug upstream : SpeakerEmbeddingLDA init sous torch.device(cuda) mais
# torch.load(map_location="cpu") laisse le ResNet sur CPU ; le mel fbank
# construit dans le contexte cuda finit CPU aussi -> RuntimeError device
# mismatch. On instancie le LDA sur CPU et on l'injecte : l'embedding est
# ensuite transfere sur cuda dans make_cond_dict (via prepare_conditioning).
from zonos.speaker_cloning import SpeakerEmbeddingLDA

if model.spk_clone_model is None:
model.spk_clone_model = SpeakerEmbeddingLDA(device="cpu")
spk_emb = model.make_speaker_embedding(spkref, sr)

if torch.cuda.is_available():
torch.cuda.reset_peak_memory_stats()
t0 = time.time()
cond_dict = make_cond_dict(
text=text,
language=lang_code,
speaker=spk_emb,
)
conditioning = model.prepare_conditioning(cond_dict)
# disable_torch_compile=True : MSVC cl.exe absent sur po-2023 (torch.compile
# leve InductorError "Compiler: cl is not found"). Le mode eager est plus
# lent mais fonctionne partout.
codes = model.generate(conditioning, disable_torch_compile=True)
wavs = model.autoencoder.decode(codes).cpu()
dt = time.time() - t0

wav = wavs[0]
out_sr = int(model.autoencoder.sampling_rate)
# torchaudio 2.11 route save() sur torchcodec (non installe) — soundfile,
# deja dans le venv, ecrit le wav sans dependance supplementaire.
sf.write(out_wav, wav.squeeze(0).numpy(), out_sr)
duration_s = float(wav.shape[-1] / out_sr)
rtf = dt / duration_s if duration_s > 0 else float("inf")
vram_peak_gb = torch.cuda.max_memory_allocated() / 1024**3 if torch.cuda.is_available() else 0.0
return {
"model": DEFAULT_MODEL_ID,
"language": language,
"language_code": lang_code,
"reference_wav": str(ref_path),
"out_wav": str(out_wav),
"sample_rate": out_sr,
"duration_s": duration_s,
"wallclock_s": float(dt),
"rtf": float(rtf),
"vram_peak_gb": float(vram_peak_gb),
"n_samples": int(wav.shape[-1]),
}


def main():
p = argparse.ArgumentParser(description=__doc__.splitlines()[0])
p.add_argument("--text", required=True, help="Texte à synthétiser")
p.add_argument("--out", required=True, help="Chemin .wav de sortie")
p.add_argument("--language", default=DEFAULT_LANGUAGE, choices=get_supported_languages())
p.add_argument("--model-id", default=DEFAULT_MODEL_ID)
p.add_argument("--device", default="cuda" if torch.cuda.is_available() else "cpu")
p.add_argument("--dtype", default="bf16", choices=["bf16", "fp16", "fp32"])
args = p.parse_args()

model = load_model(args.model_id, device=args.device, dtype=args.dtype)
result = synth(text=args.text, out_wav=args.out, model=model, language=args.language)
print(json.dumps(result, indent=2, ensure_ascii=False))


if __name__ == "__main__":
main()
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ Tell c.c.c.d.F strict fondateur (règle globale) : **RÉPARER, ne JAMAIS contour

```bash
mkdir -p /c/ProgramData/sox-portable
curl -L -o /c/ProgramData/sox-portable/sox.zip https://sourceforge.net/projects/sox/files/sox/14.4.2/sox-14.4.2-win32.zip/download
curl -L -o /c/ProgramData/sox-portable/sox.zip "https://sourceforge.net/projects/sox/files/sox/14.4.2/sox-14.4.2-win32.zip/download"
cd /c/ProgramData/sox-portable && powershell -NoProfile -Command "Expand-Archive -Path sox.zip -DestinationPath . -Force"
# PATH local pour cette session :
export PATH="/c/ProgramData/sox-portable/sox-14.4.2:$PATH"
Expand All @@ -21,6 +21,36 @@ sox --version # SoX v14.4.2

Pour rendre sox permanent : ajouter `C:\ProgramData\sox-portable\sox-14.4.2` au PATH système via `sysdm.cpl` → Environment Variables → Path (action user one-time, RECOVERABLE-USER-HAND).

### espeak-ng 1.52.0 (Zonos)

`zonos.conditioning.phonemize` exige le binaire `espeak-ng` au runtime.
**Scoop (userspace, sans admin)** — la voie retenue sur po-2023 :

```bash
# une fois par machine (userspace, pas d'elevation) :
powershell -NoProfile -Command "irm get.scoop.sh -OutFile \"$env:TEMP\install-scoop.ps1\"; & \"$env:TEMP\install-scoop.ps1\" -ScoopDir 'C:\ProgramData\scoop-user'"
C:\ProgramData\scoop-user\shims\scoop.cmd install espeak-ng
```

Variables d'env requises pour le client Zonos (session ou shell) :

```bash
export PATH="/c/ProgramData/scoop-user/apps/espeak-ng/current:$PATH"
export PHONEMIZER_ESPEAK_LIBRARY="C:\\ProgramData\\scoop-user\\apps\\espeak-ng\\current\\libespeak-ng.dll"
export PHONEMIZER_ESPEAK_DATA_PATH="C:\\ProgramData\\scoop-user\\apps\\espeak-ng\\current\\espeak-ng-data"
```

Vérif :

```bash
"/c/ProgramData/scoop-user/apps/espeak-ng/current/espeak-ng.exe" --version
# eSpeak NG text-to-speech: 1.52.0 Data at: C:\ProgramData\scoop-user\apps\espeak-ng\current\espeak-ng-data
```

PIE : le MSI officiel exige l'admin (erreur 1925 en non-élevé) et l'extraction
CAB manuelle produit des DLLs corrompues (access violation). Scoop fait les
deux correctement en userspace.

## Création du venv Python 3.12

```bash
Expand Down Expand Up @@ -66,3 +96,55 @@ _runtime/venv-qwen3tts/Scripts/python.exe -c "import qwen_tts; print('OK', qwen_
- RTX 3080 Ti Laptop GPU (16 GB) — fallback si RTX 3090 occupée

Tell c.c.c.d.767-L1 strict fondateur : **zero-dep-manifeste ≠ zero-dep-réel**. Le test ci-dessus (import + GPU dispo) est OBLIGATOIRE avant tout client.py.

## Zonos (client `zonos.py`, venv `venv-zonos`)

Slot 3 shortlist — Zyphra/Zonos-v0.1-transformer (Apache-2.0, FR par code eSpeak `fr-fr`).

**PIE — PyPI `zonos` est un placeholder squatté** (0.1.0.dev0, wheel vide,
Home-page `github.com/yourusername/zonos`, Auteur "Your Name") : `pip install zonos`
rend rc=0 sans livrer aucun module `zonos` (mesuré 29/09). Le paquet officiel
vient du repo GitHub **Zyphra/Zonos** (install éditable — règle F + Prong A).

**Variante transformer, pas hybride** : `Zyphra/Zonos-v0.1-hybrid` exige
`mamba-ssm` + `causal-conv1d` (extras `compile` du pyproject : noyaux CUDA
compilés, pas de wheels Windows) ; `Zyphra/Zonos-v0.1-transformer` est un
checkpoint officiel du même repo sans dépendance compilée — choix de variante
documenté, pas un contournement.

```bash
# 1. venv Python 3.10
cd D:/Dev/CoursIA-17586-zonos
py -3.10 -m venv _runtime/venv-zonos
_runtime/venv-zonos/Scripts/python.exe -m pip install --upgrade pip

# 2. torch cu126 — PIE 2.11.0 : l'index cu126 ne porte torchaudio que jusqu'a
# 2.11.0 (torchaudio==2.14.0 introuvable, mesure 29/09) ; 2.11.0 >> exigence Zonos (>=2.5.1)
_runtime/venv-zonos/Scripts/python.exe -m pip install torch==2.11.0 torchaudio==2.11.0 --index-url https://download.pytorch.org/whl/cu126

# 3. zonos officiel (repo, PAS PyPI) + jambe WER du banc
git clone --depth 1 https://github.com/Zyphra/Zonos.git _runtime/Zonos
_runtime/venv-zonos/Scripts/python.exe -m pip install -e ./_runtime/Zonos soundfile faster-whisper
```

Verif pre-banc (767-L1 zero-dep-manifeste != zero-dep-reel) :

```bash
_runtime/venv-zonos/Scripts/python.exe -c "import torch; from zonos.model import Zonos; from zonos.conditioning import make_cond_dict; print('OK', torch.__version__, torch.cuda.is_available())"
```

Banc :

```bash
_runtime/venv-zonos/Scripts/python.exe bakeoff_large/banc_phase_a0.py \
--client zonos --language French \
--out-root <GDrive>/run-<id>/A0-bakeoff/zonos
```

**Voix de reference (clonage)** : asset zh `zero_shot_prompt.wav` du repo
CosyVoice du worktree voisin `CoursIA-17586-cosyvoice3/_runtime/CosyVoice/asset/`
— meme timbre que le client cosyvoice3 (comparabilite bakeoff, cross-lingual
zh->FR). Le client le resolve par candidats (`clients/zonos.py::_reference_wav`) ;
sans lui : `git clone --recursive https://github.com/FunAudioLLM/CosyVoice.git`
dans `_runtime/`. Zonos n'a NI speaker nomme NI instruct NL (v0.1) — les kwargs
`speaker`/`instruct` du banc sont ignores.
Loading