Repository navigation
docs(genai,#14549): Wan 2.2 5B TI2V — VAE Kijai Wan 2.1 + --offload-cpu debloquent le VAE decode, verdict SOTA-OK - #17356
Merged
Conversation
… --offload-cpu debloque le VAE decode, verdict SOTA-OK (2 MP4 QA-valides) Run B1 (9f, offload, VAE Wan2.2 convertie): OOM 23372 MiB — alloc monolithique quasi-constante 23,4-26,5 GB sur 9/17/33 frames, insensible au nombre de frames. Run B2 (variable unique: VAE Kijai Wan 2.1 bf16): SUCCES — vae-decode 3337 ms, MP4 640x352 9f en 62,8 s, pic VRAM 7719 MiB. QA pixel objective: objet rouge 9/9 frames, oscillation verticale du centroide (rebond). Run B3 (B2 + --seed 42): second succes (61,3 s) + finding: la seed video n'est pas gouvernee par --seed (1246342748 vs 42) — non reproductible, documente. Flag reel verifie --help firsthand: --offload-cpu (le --stream-weights du plan voie (d) n'existe pas — lecon phantom-tool-name). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This was referenced Sep 22, 2026
Owner
Author
|
[ADJOINT PREFLIGHT] |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Grain: DEEP/genai — lane myia-po-2023:CoursIA — prev: MED/refactor #17347
Wan 2.2 5B TI2V Run B —
--offload-cpu+ VAE Kijai Wan 2.1 : le VAE decode passe, verdictSOTA-OK(voie (d) c.446-bis, #14549)Diagnostic de départ (mesuré c.446-bis, repris tel quel)
Wan 2.2 5B TI2V était bloqué
RECOVERABLE-LOCALsur un seul goulot : le VAE decode monolithique deWanVae.DecodeNative:229tente uncudaMallocunique — 26 552 MiB @ 33 frames, 23 897 MiB @ 17 frames — sur une RTX 3090 de 24 576 MiB avec DiT (3,27 GB) + UMT5 (3,65 GB) résidents. DiT charge, denoise passe 100 %, seul le decode final OOM.Correction de nom (leçon phantom-tool-name, appliquée firsthand)
Le plan voie (d) de c.446-bis citait
--stream-weights: ce flag n'existe pas dans le CLI v3.3.0.0. Le flag réel, vérifié--helpavant lancement :--offload-cpu(« Stream the DiT weights from RAM instead of holding them resident in VRAM »).Trois runs mesurés firsthand (2026-09-22, RTX 3090 = Device 0, bannière ggml_cuda_init citée dans le ledger)
Contribution nouvelle B1 : l'alloc monolithique est quasi-constante ~23,4-26,5 GB de 9 à 33 frames (−12 % pour ÷4 frames) — réduire les frames n'est pas une voie ; la taille de la VAE l'est. B2 change UNE variable (le VAE) et le decode passe.
QA objective — pattern c.268 (reproductible, hors vision)
Extraction ffmpeg + analyse pixels rouges (PIL/numpy, seuils r>150, r−g>60, r−b>60) :
La VAE Wan 2.1 décode sémantiquement les latents Wan 2.2 — ce n'est pas du bruit de decode. QA vision œil-sur-artefact peut être contre-signée par lane vision en complément, non en substitution.
Verdict axe Vidéo Wan 2.2 5B TI2V :
SOTA-OKLe vrai outil (TensorSharp CLI v3.3.0.0 + DiT Wan 2.2 Q4_K_M + UMT5-XXL) proprement installé et invoqué ; 2 MP4 bout-en-bout produits et QA-validés, pic VRAM 7,7 GB / 24,5 GB. Caveat documenté : le combo gagnant exige la VAE Wan 2.1 Kijai bf16 — le chemin vendor Wan 2.2 VAE (
.pthconverti ou non) reste OOM (alloc quasi-constante, DiT offloadé ou non) ; la recovery structurelle reste la voie (c) (patch upstreamDecodeNative— la ligne B3band 1/1montre que le banding existe en interne).Table des 4 axes après ce grain : Texte
SOTA-OKbinaire (c.257) · ImageSOTA-OK(c.266-273) · Wan 2.1SOTA-OK(c.445) · Wan 2.2SOTA-OK(c.33).Findings incidents
--seedne gouverne pas la seed vidéo :--seed 42passé, seeds utilisées 978075981 (sans flag) / 1246342748 (avec flag) — aucun claim de reproductibilité seed-vidéo, documenté verbatim.--offload-cpuà 5B : pénalité denoise nulle (0,6 s/pass identique résident) — l'OOM était portée par la VAE, pas par le DiT.Acceptance #14549 — état après Run B
SOTA-OK(Wan 2.2 avec caveat combo)Périmètre strict
docs/ledgers/14549-tensorsharp-multimodal.md(section c.33 append-only, +77 lignes).C:\Users\jsboi\tensorsharp-investigation\wan_outputs\— chemins cités dans le ledger.See #14549
🤖 Generated with Claude Code