You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 8fb6d55
Browse filesBrowse the repository at this point in the historyBrowse files
@@ -39,10 +42,10 @@ World-R1 aligns text-to-video generation with 3D constraints through reinforceme
39
42
40
43
## Highlights
41
44
42
-
-RL post-training for video foundation models using Flow-GRPO, without architecture changes at inference time.
43
-
-A composite reward that combines reconstruction fidelity, trajectory alignment, meta-view semantic scoring, and general visual quality.
44
-
- A pure-text training set with about 3,000 prompts, including a dynamic subset for regularizing non-rigid motion.
45
-
-Strong 3D consistency gains reported in the paper: `+10.23 dB` PSNR over Wan2.1-T2V-1.3B and `+7.91 dB` over Wan2.1-T2V-14B.
45
+
-3D-aware reinforcement learning aligns generated videos with geometric constraints through meta-view assessment, reconstruction consistency, and trajectory alignment rewards.
46
+
-General visual quality is preserved by combining the 3D-aware reward with an aesthetic reward during Flow-GRPO-based post-training.
47
+
- A periodic dynamic-only training phase regularizes the model with dynamic-scene prompts, improving motion diversity while retaining learned 3D consistency.
48
+
-Camera-aware latent initialization converts text-specified camera motion into trajectory-guided noise wrapping, enabling implicit camera conditioning without changing the base video architecture.
46
49
47
50
## Method
48
51
@@ -52,41 +55,6 @@ World-R1 aligns text-to-video generation with 3D constraints through reinforceme
52
55
53
56
World-R1 first converts camera instructions in text prompts into explicit trajectories and injects the motion prior into the initial video latents through noise wrapping. During RL fine-tuning, the model is optimized with 3D-aware feedback from reconstruction and camera-control metrics, together with a general visual reward. A periodic dynamic-only phase prevents the model from overfitting to rigid static scenes.
54
57
55
-
## Results
56
-
57
-
| Model | PSNR ↑ | SSIM ↑ | LPIPS ↓ |
58
-
| --- | ---: | ---: | ---: |
59
-
| Wan2.1-T2V-1.3B | 17.40 | 0.550 | 0.467 |
60
-
| World-R1-Small | 27.63 | 0.858 | 0.201 |
61
-
| Wan2.1-T2V-14B | 19.76 | 0.629 | 0.405 |
62
-
| World-R1-Large | 27.67 | 0.865 | 0.162 |
63
-
64
-
## Release Contents
65
-
66
-
- Training code, reward functions, and launch scripts for the World-R1 RL pipeline.
67
-
- The 3D reward server and bundled `Depth Anything 3` integration used by the release.
68
-
- Prompt-only dataset files under `dataset/final` and `dataset/enhanced`.
69
-
- Utility scripts for inference, ablations, prompt processing, and noise-wrap analysis.
70
-
- Third-party license files under `licenses/`.
71
-
72
-
This repository does not include released World-R1 checkpoints. Base video model checkpoints should be obtained separately.
73
-
74
-
## Repository Layout
75
-
76
-
```text
77
-
World-R1/
78
-
├── config/
79
-
├── dataset/
80
-
│ ├── enhanced/
81
-
│ └── final/
82
-
├── flow_grpo/
83
-
├── licenses/
84
-
├── reward_server/
85
-
├── scripts/
86
-
├── assets/
87
-
└── pyproject.toml
88
-
```
89
-
90
58
## Setup
91
59
92
60
Use a Python 3.10+ environment with CUDA and a PyTorch build that matches your driver. A practical setup flow is:
If reward servers are already running, launch training directly:
173
129
174
130
```bash
@@ -223,17 +179,17 @@ Unless a file states otherwise, the rest of this repository is covered by the ro
223
179
If you find this repository useful, please cite:
224
180
225
181
```bibtex
226
-
@misc{wang2026worldr1,
227
-
title={World-R1: Reinforcing 3D Constraints for Text-to-Video Generation},
182
+
@article{wang2026worldr1,
228
183
author={Weijie Wang and Xiaoxuan He and Youping Gu and Zeyu Zhang and Yefei He and Yanbo Ding and Xirui Hu and Donny Y. Chen and Zhiyuan He and Yuqing Yang and Yifan Yang and Bohan Zhuang},
184
+
title={World-R1: Reinforcing 3D Constraints for Text-to-Video Generation},
185
+
journal={arXiv preprint},
229
186
year={2026},
230
-
note={ICML 2026}
231
187
}
232
188
```
233
189
234
190
## Acknowledgements
235
191
236
-
World-R1 builds on top of several strong open-source projects and model ecosystems, including Wan, CogVideoX, Flow-GRPO, and Depth Anything 3. We thank the original authors and maintainers for making those foundations available.
192
+
World-R1 builds on top of several strong open-source projects and model ecosystems, including [Wan2.1](https://github.com/Wan-Video/Wan2.1), [Flow-GRPO](https://github.com/yifan123/flow_grpo), [Depth Anything 3](https://github.com/ByteDance-Seed/Depth-Anything-3), and [Qwen3-VL](https://github.com/QwenLM/Qwen3-VL). We thank the original authors and maintainers for making those foundations available.
0 commit comments