Skip to content

Commit 8fb6d55

Browse files
authored
Update README.md
1 parent 0a99859 commit 8fb6d55

1 file changed

Lines changed: 28 additions & 72 deletions

File tree

‎README.md‎

Lines changed: 28 additions & 72 deletions
Original file line numberDiff line numberDiff line change
@@ -1,18 +1,25 @@
1-
# World-R1: Reinforcing 3D Constraints for Text-to-Video Generation
1+
<h1 align="center">World-R1: Reinforcing 3D Constraints for Text-to-Video Generation</h1>
22

33
<p align="center">
4-
Weijie Wang<sup>*&#8224;1,2</sup> &nbsp;
5-
Xiaoxuan He<sup>*1</sup> &nbsp;
6-
Youping Gu<sup>*1</sup> &nbsp;
7-
Zeyu Zhang<sup>3</sup> &nbsp;
8-
Yefei He<sup>1</sup> &nbsp;
9-
Yanbo Ding<sup>2</sup> &nbsp;
10-
Xirui Hu<sup>3</sup> &nbsp;
11-
Donny Y. Chen<sup>3</sup> &nbsp;
12-
Zhiyuan He<sup>2</sup> &nbsp;
13-
Yuqing Yang<sup>2</sup> &nbsp;
14-
Yifan Yang<sup>&#8225;2</sup> &nbsp;
15-
Bohan Zhuang<sup>&#8225;1</sup>
4+
<a href="main.pdf"><img src="https://img.shields.io/badge/Paper-B31B1B?style=for-the-badge&logo=arxiv&logoColor=white" alt="Paper"></a>
5+
<a href="https://aka.ms/world-r1"><img src="https://img.shields.io/badge/Project%20Page-000000?style=for-the-badge&logo=googlechrome&logoColor=white" alt="Project Page"></a>
6+
</p>
7+
8+
<p align="center">
9+
<a href="https://lhmd.top/">Weijie Wang</a><sup>1,2,*,&#8224;</sup> &nbsp;
10+
<a href="https://github.com/Shredded-Pork">Xiaoxuan He</a><sup>1,*</sup> &nbsp;
11+
<a href="https://github.com/Tacossp">Youping Gu</a><sup>1,*</sup> &nbsp;
12+
<a href="https://www.microsoft.com/en-us/research/people/yifanyang/">Yifan Yang</a><sup>2,&#8225;</sup>
13+
<br>
14+
<a href="https://steve-zeyu-zhang.github.io/">Zeyu Zhang</a><sup>3</sup> &nbsp;
15+
<a href="https://openreview.net/profile?id=~Yefei_He1">Yefei He</a><sup>1</sup> &nbsp;
16+
<a href="https://github.com/DINGYANB">Yanbo Ding</a><sup>2</sup> &nbsp;
17+
<a href="https://openreview.net/profile?id=~Xirui_Hu1">Xirui Hu</a><sup>3</sup>
18+
<br>
19+
<a href="https://donydchen.github.io/">Donny Y. Chen</a><sup>3</sup> &nbsp;
20+
<a href="https://www.microsoft.com/en-us/research/people/zhiyuhe/">Zhiyuan He</a><sup>2</sup> &nbsp;
21+
<a href="https://www.microsoft.com/en-us/research/people/yuqyang/">Yuqing Yang</a><sup>2</sup> &nbsp;
22+
<a href="https://bohanzhuang.github.io/">Bohan Zhuang</a><sup>1,&#8225;</sup>
1623
</p>
1724

1825
<p align="center">
@@ -27,10 +34,6 @@
2734
<sup>&#8225;</sup>Corresponding authors
2835
</p>
2936

30-
<p align="center">
31-
<a href="https://aka.ms/world-r1">Project Page</a>
32-
</p>
33-
3437
<p align="center">
3538
<img src="assets/teaser.png" alt="World-R1 teaser" width="100%">
3639
</p>
@@ -39,10 +42,10 @@ World-R1 aligns text-to-video generation with 3D constraints through reinforceme
3942

4043
## Highlights
4144

42-
- RL post-training for video foundation models using Flow-GRPO, without architecture changes at inference time.
43-
- A composite reward that combines reconstruction fidelity, trajectory alignment, meta-view semantic scoring, and general visual quality.
44-
- A pure-text training set with about 3,000 prompts, including a dynamic subset for regularizing non-rigid motion.
45-
- Strong 3D consistency gains reported in the paper: `+10.23 dB` PSNR over Wan2.1-T2V-1.3B and `+7.91 dB` over Wan2.1-T2V-14B.
45+
- 3D-aware reinforcement learning aligns generated videos with geometric constraints through meta-view assessment, reconstruction consistency, and trajectory alignment rewards.
46+
- General visual quality is preserved by combining the 3D-aware reward with an aesthetic reward during Flow-GRPO-based post-training.
47+
- A periodic dynamic-only training phase regularizes the model with dynamic-scene prompts, improving motion diversity while retaining learned 3D consistency.
48+
- Camera-aware latent initialization converts text-specified camera motion into trajectory-guided noise wrapping, enabling implicit camera conditioning without changing the base video architecture.
4649

4750
## Method
4851

@@ -52,41 +55,6 @@ World-R1 aligns text-to-video generation with 3D constraints through reinforceme
5255

5356
World-R1 first converts camera instructions in text prompts into explicit trajectories and injects the motion prior into the initial video latents through noise wrapping. During RL fine-tuning, the model is optimized with 3D-aware feedback from reconstruction and camera-control metrics, together with a general visual reward. A periodic dynamic-only phase prevents the model from overfitting to rigid static scenes.
5457

55-
## Results
56-
57-
| Model | PSNR ↑ | SSIM ↑ | LPIPS ↓ |
58-
| --- | ---: | ---: | ---: |
59-
| Wan2.1-T2V-1.3B | 17.40 | 0.550 | 0.467 |
60-
| World-R1-Small | 27.63 | 0.858 | 0.201 |
61-
| Wan2.1-T2V-14B | 19.76 | 0.629 | 0.405 |
62-
| World-R1-Large | 27.67 | 0.865 | 0.162 |
63-
64-
## Release Contents
65-
66-
- Training code, reward functions, and launch scripts for the World-R1 RL pipeline.
67-
- The 3D reward server and bundled `Depth Anything 3` integration used by the release.
68-
- Prompt-only dataset files under `dataset/final` and `dataset/enhanced`.
69-
- Utility scripts for inference, ablations, prompt processing, and noise-wrap analysis.
70-
- Third-party license files under `licenses/`.
71-
72-
This repository does not include released World-R1 checkpoints. Base video model checkpoints should be obtained separately.
73-
74-
## Repository Layout
75-
76-
```text
77-
World-R1/
78-
├── config/
79-
├── dataset/
80-
│ ├── enhanced/
81-
│ └── final/
82-
├── flow_grpo/
83-
├── licenses/
84-
├── reward_server/
85-
├── scripts/
86-
├── assets/
87-
└── pyproject.toml
88-
```
89-
9058
## Setup
9159

9260
Use a Python 3.10+ environment with CUDA and a PyTorch build that matches your driver. A practical setup flow is:
@@ -157,18 +125,6 @@ NUM_PROCESSES=6 \
157125
bash scripts/run_single_node.sh
158126
```
159127

160-
CogVideoX training:
161-
162-
```bash
163-
MODEL_FAMILY=cogvideox \
164-
MODEL_PATH=THUDM/CogVideoX1.5-5B \
165-
TRAIN_CONFIG=config/world_r1.py:world_r1_cogvideox_5b \
166-
SERVER_VISIBLE_DEVICES=0,1 \
167-
TRAIN_VISIBLE_DEVICES=2,3,4,5,6,7 \
168-
NUM_PROCESSES=6 \
169-
bash scripts/run_single_node.sh
170-
```
171-
172128
If reward servers are already running, launch training directly:
173129

174130
```bash
@@ -223,17 +179,17 @@ Unless a file states otherwise, the rest of this repository is covered by the ro
223179
If you find this repository useful, please cite:
224180

225181
```bibtex
226-
@misc{wang2026worldr1,
227-
title={World-R1: Reinforcing 3D Constraints for Text-to-Video Generation},
182+
@article{wang2026worldr1,
228183
author={Weijie Wang and Xiaoxuan He and Youping Gu and Zeyu Zhang and Yefei He and Yanbo Ding and Xirui Hu and Donny Y. Chen and Zhiyuan He and Yuqing Yang and Yifan Yang and Bohan Zhuang},
184+
title={World-R1: Reinforcing 3D Constraints for Text-to-Video Generation},
185+
journal={arXiv preprint},
229186
year={2026},
230-
note={ICML 2026}
231187
}
232188
```
233189

234190
## Acknowledgements
235191

236-
World-R1 builds on top of several strong open-source projects and model ecosystems, including Wan, CogVideoX, Flow-GRPO, and Depth Anything 3. We thank the original authors and maintainers for making those foundations available.
192+
World-R1 builds on top of several strong open-source projects and model ecosystems, including [Wan2.1](https://github.com/Wan-Video/Wan2.1), [Flow-GRPO](https://github.com/yifan123/flow_grpo), [Depth Anything 3](https://github.com/ByteDance-Seed/Depth-Anything-3), and [Qwen3-VL](https://github.com/QwenLM/Qwen3-VL). We thank the original authors and maintainers for making those foundations available.
237193

238194
## Support
239195

0 commit comments

Comments
 (0)