Skip to content

Commit 0a99859

Browse files
committed
update
1 parent 1f463b7 commit 0a99859

230 files changed

Lines changed: 37325 additions & 450 deletions

File tree

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎.gitignore‎

Lines changed: 11 additions & 418 deletions
Large diffs are not rendered by default.

‎README.md‎

Lines changed: 242 additions & 13 deletions
Original file line numberDiff line numberDiff line change
@@ -1,28 +1,257 @@
1-
# Project
1+
# World-R1: Reinforcing 3D Constraints for Text-to-Video Generation
22

3-
> This repo has been populated by an initial template to help get you started. Please
4-
> make sure to update the content to build a great experience for community-building.
3+
<p align="center">
4+
Weijie Wang<sup>*&#8224;1,2</sup> &nbsp;
5+
Xiaoxuan He<sup>*1</sup> &nbsp;
6+
Youping Gu<sup>*1</sup> &nbsp;
7+
Zeyu Zhang<sup>3</sup> &nbsp;
8+
Yefei He<sup>1</sup> &nbsp;
9+
Yanbo Ding<sup>2</sup> &nbsp;
10+
Xirui Hu<sup>3</sup> &nbsp;
11+
Donny Y. Chen<sup>3</sup> &nbsp;
12+
Zhiyuan He<sup>2</sup> &nbsp;
13+
Yuqing Yang<sup>2</sup> &nbsp;
14+
Yifan Yang<sup>&#8225;2</sup> &nbsp;
15+
Bohan Zhuang<sup>&#8225;1</sup>
16+
</p>
517

6-
As the maintainer of this project, please make a few updates:
18+
<p align="center">
19+
<sup>1</sup>Zhejiang University &nbsp;&nbsp;
20+
<sup>2</sup>Microsoft Research &nbsp;&nbsp;
21+
<sup>3</sup>Independent Researcher
22+
</p>
723

8-
- Improving this README.MD file to provide a great experience
9-
- Updating SUPPORT.MD with content about this project's support experience
10-
- Understanding the security reporting process in SECURITY.MD
11-
- Remove this section from the README
24+
<p align="center">
25+
<sup>*</sup>Equal contribution &nbsp;&nbsp;
26+
<sup>&#8224;</sup>Work done during an internship at MSRA &nbsp;&nbsp;
27+
<sup>&#8225;</sup>Corresponding authors
28+
</p>
29+
30+
<p align="center">
31+
<a href="https://aka.ms/world-r1">Project Page</a>
32+
</p>
33+
34+
<p align="center">
35+
<img src="assets/teaser.png" alt="World-R1 teaser" width="100%">
36+
</p>
37+
38+
World-R1 aligns text-to-video generation with 3D constraints through reinforcement learning. Instead of changing the base video model architecture or relying on large-scale 3D supervision, it combines camera-aware latent initialization, 3D-aware rewards from pre-trained foundation models, and a periodic decoupled training strategy to improve geometric consistency while preserving visual quality and motion diversity.
39+
40+
## Highlights
41+
42+
- RL post-training for video foundation models using Flow-GRPO, without architecture changes at inference time.
43+
- A composite reward that combines reconstruction fidelity, trajectory alignment, meta-view semantic scoring, and general visual quality.
44+
- A pure-text training set with about 3,000 prompts, including a dynamic subset for regularizing non-rigid motion.
45+
- Strong 3D consistency gains reported in the paper: `+10.23 dB` PSNR over Wan2.1-T2V-1.3B and `+7.91 dB` over Wan2.1-T2V-14B.
46+
47+
## Method
48+
49+
<p align="center">
50+
<img src="assets/pipeline.png" alt="World-R1 pipeline" width="100%">
51+
</p>
52+
53+
World-R1 first converts camera instructions in text prompts into explicit trajectories and injects the motion prior into the initial video latents through noise wrapping. During RL fine-tuning, the model is optimized with 3D-aware feedback from reconstruction and camera-control metrics, together with a general visual reward. A periodic dynamic-only phase prevents the model from overfitting to rigid static scenes.
54+
55+
## Results
56+
57+
| Model | PSNR ↑ | SSIM ↑ | LPIPS ↓ |
58+
| --- | ---: | ---: | ---: |
59+
| Wan2.1-T2V-1.3B | 17.40 | 0.550 | 0.467 |
60+
| World-R1-Small | 27.63 | 0.858 | 0.201 |
61+
| Wan2.1-T2V-14B | 19.76 | 0.629 | 0.405 |
62+
| World-R1-Large | 27.67 | 0.865 | 0.162 |
63+
64+
## Release Contents
65+
66+
- Training code, reward functions, and launch scripts for the World-R1 RL pipeline.
67+
- The 3D reward server and bundled `Depth Anything 3` integration used by the release.
68+
- Prompt-only dataset files under `dataset/final` and `dataset/enhanced`.
69+
- Utility scripts for inference, ablations, prompt processing, and noise-wrap analysis.
70+
- Third-party license files under `licenses/`.
71+
72+
This repository does not include released World-R1 checkpoints. Base video model checkpoints should be obtained separately.
73+
74+
## Repository Layout
75+
76+
```text
77+
World-R1/
78+
├── config/
79+
├── dataset/
80+
│ ├── enhanced/
81+
│ └── final/
82+
├── flow_grpo/
83+
├── licenses/
84+
├── reward_server/
85+
├── scripts/
86+
├── assets/
87+
└── pyproject.toml
88+
```
89+
90+
## Setup
91+
92+
Use a Python 3.10+ environment with CUDA and a PyTorch build that matches your driver. A practical setup flow is:
93+
94+
1. Create and activate a clean environment:
95+
96+
```bash
97+
conda create -n world-r1 python=3.10 -y
98+
conda activate world-r1
99+
pip install --upgrade pip
100+
```
101+
102+
2. Install PyTorch first. Replace the wheel index below with the one that matches your CUDA version:
103+
104+
```bash
105+
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu124
106+
```
107+
108+
3. Install the repository and the core runtime packages used by training and inference:
109+
110+
```bash
111+
pip install -e .
112+
pip install accelerate diffusers transformers peft wandb absl-py ml-collections \
113+
numpy pillow imageio tqdm requests httpx flask addict omegaconf einops \
114+
ftfy sentencepiece protobuf scipy opencv-python huggingface_hub
115+
```
116+
117+
4. Install the extra packages used by the 3D reward stack and visualization utilities:
118+
119+
```bash
120+
pip install lpips trimesh plyfile moviepy pycolmap gsplat evo e3nn hpsv2 qwen-vl-utils
121+
```
122+
123+
5. Optional acceleration packages, depending on your machine and base model setup:
124+
125+
```bash
126+
pip install xformers bitsandbytes
127+
```
128+
129+
The launcher scripts default to `python3` and `torchrun` from `PATH`. If your environment uses different executables, set them explicitly:
130+
131+
```bash
132+
export WAN_PYTHON=$(which python)
133+
export WAN_TORCHRUN=$(which torchrun)
134+
```
135+
136+
For WAN-based utilities, you can either pass `--model /path/to/checkpoint` or set:
137+
138+
```bash
139+
export WORLD_R1_WAN_MODEL=/path/to/Wan-Diffusers-checkpoint
140+
```
141+
142+
A simple smoke test after setup:
143+
144+
```bash
145+
python -c "import torch, diffusers, transformers, peft, flask, lpips; print('env ok')"
146+
```
147+
148+
## Training
149+
150+
Single-node training with local reward servers:
151+
152+
```bash
153+
MODEL_PATH=/path/to/Wan2.1-T2V-14B-Diffusers \
154+
SERVER_VISIBLE_DEVICES=0,1 \
155+
TRAIN_VISIBLE_DEVICES=2,3,4,5,6,7 \
156+
NUM_PROCESSES=6 \
157+
bash scripts/run_single_node.sh
158+
```
159+
160+
CogVideoX training:
161+
162+
```bash
163+
MODEL_FAMILY=cogvideox \
164+
MODEL_PATH=THUDM/CogVideoX1.5-5B \
165+
TRAIN_CONFIG=config/world_r1.py:world_r1_cogvideox_5b \
166+
SERVER_VISIBLE_DEVICES=0,1 \
167+
TRAIN_VISIBLE_DEVICES=2,3,4,5,6,7 \
168+
NUM_PROCESSES=6 \
169+
bash scripts/run_single_node.sh
170+
```
171+
172+
If reward servers are already running, launch training directly:
173+
174+
```bash
175+
REWARD_3D_SERVER_URL=http://127.0.0.1:18089 \
176+
GENERAL_REWARD_SERVER_URL=http://127.0.0.1:18090 \
177+
MODEL_PATH=/path/to/Wan2.1-T2V-14B-Diffusers \
178+
NUM_PROCESSES=6 \
179+
bash scripts/run_training.sh
180+
```
181+
182+
For multi-node training, start external reward servers first and then provide `REWARD_3D_SERVER_URL`, `GENERAL_REWARD_SERVER_URL`, `MASTER_ADDR`, `MASTER_PORT`, `NNODES`, and `NODE_RANK` before running `scripts/run_multi_node.sh`.
183+
184+
## Reward Servers and Utilities
185+
186+
Start the two release reward services independently:
187+
188+
```bash
189+
bash scripts/run_reward_3d_server.sh
190+
bash scripts/run_general_reward_server.sh
191+
```
192+
193+
Useful utilities included in this release:
194+
195+
- `scripts/train_world_r1.py`: main RL training entry point.
196+
- `scripts/infer_wan_lora.py`: batch inference for WAN checkpoints or LoRAs.
197+
- `scripts/noise_wrap_ablation.py`: visualize latent wrapping and generation effects.
198+
- `scripts/noise_wrap_strength_sweep.py`: sweep wrap strengths over a prompt set.
199+
- `scripts/serve_reward_3d.py` and `scripts/serve_general_reward.py`: reward service backends.
200+
201+
## Dataset
202+
203+
The repository ships prompt-only data used by the release:
204+
205+
- `dataset/final/`: the base prompt split used for training, validation, and dynamic regularization.
206+
- `dataset/enhanced/`: expanded prompt variants used for richer post-training supervision.
207+
208+
The prompt-processing helpers under `scripts/` are configured to read and write these repository-local directories by default.
209+
210+
## License
211+
212+
This project is licensed under the [MIT License](LICENSE).
213+
214+
The top-level `licenses/` directory is reserved for bundled third-party source code that remains under its upstream license:
215+
216+
- `flow_grpo/` includes bundled or adapted `Flow-GRPO` code. See [`licenses/FLOW_GRPO_LICENSE`](licenses/FLOW_GRPO_LICENSE).
217+
- `reward_server/depth_anything_3/` includes bundled or modified `Depth Anything 3` code. See [`licenses/DEPTH_ANYTHING_3_LICENSE`](licenses/DEPTH_ANYTHING_3_LICENSE).
218+
219+
Unless a file states otherwise, the rest of this repository is covered by the root MIT license. Please review the third-party license files before redistributing derivative work based on the bundled upstream code.
220+
221+
## Citation
222+
223+
If you find this repository useful, please cite:
224+
225+
```bibtex
226+
@misc{wang2026worldr1,
227+
title={World-R1: Reinforcing 3D Constraints for Text-to-Video Generation},
228+
author={Weijie Wang and Xiaoxuan He and Youping Gu and Zeyu Zhang and Yefei He and Yanbo Ding and Xirui Hu and Donny Y. Chen and Zhiyuan He and Yuqing Yang and Yifan Yang and Bohan Zhuang},
229+
year={2026},
230+
note={ICML 2026}
231+
}
232+
```
233+
234+
## Acknowledgements
235+
236+
World-R1 builds on top of several strong open-source projects and model ecosystems, including Wan, CogVideoX, Flow-GRPO, and Depth Anything 3. We thank the original authors and maintainers for making those foundations available.
237+
238+
## Support
239+
240+
See [SUPPORT.md](SUPPORT.md) for usage support and [SECURITY.md](SECURITY.md) for vulnerability reporting.
12241

13242
## Contributing
14243

15-
This project welcomes contributions and suggestions. Most contributions require you to agree to a
244+
This project welcomes contributions and suggestions. Most contributions require you to agree to a
16245
Contributor License Agreement (CLA) declaring that you have the right to, and actually do, grant us
17246
the rights to use your contribution. For details, visit [Contributor License Agreements](https://cla.opensource.microsoft.com).
18247

19248
When you submit a pull request, a CLA bot will automatically determine whether you need to provide
20-
a CLA and decorate the PR appropriately (e.g., status check, comment). Simply follow the instructions
21-
provided by the bot. You will only need to do this once across all repos using our CLA.
249+
a CLA and decorate the PR appropriately. Simply follow the instructions provided by the bot. You will
250+
only need to do this once across all repos using our CLA.
22251

23252
This project has adopted the [Microsoft Open Source Code of Conduct](https://opensource.microsoft.com/codeofconduct/).
24-
For more information see the [Code of Conduct FAQ](https://opensource.microsoft.com/codeofconduct/faq/) or
25-
contact [opencode@microsoft.com](mailto:opencode@microsoft.com) with any additional questions or comments.
253+
For more information, see the [Code of Conduct FAQ](https://opensource.microsoft.com/codeofconduct/faq/) or
254+
contact [opencode@microsoft.com](mailto:opencode@microsoft.com) with additional questions or comments.
26255

27256
## Trademarks
28257

‎SUPPORT.md‎

Lines changed: 8 additions & 19 deletions
Original file line numberDiff line numberDiff line change
@@ -1,25 +1,14 @@
1-
# TODO: The maintainer of this repo has not yet edited this file
2-
3-
**REPO OWNER**: Do you want Customer Service & Support (CSS) support for this product/project?
4-
5-
- **No CSS support:** Fill out this template with information about how to file issues and get help.
6-
- **Yes CSS support:** Fill out an intake form at [aka.ms/onboardsupport](https://aka.ms/onboardsupport). CSS will work with/help you to determine next steps.
7-
- **Not sure?** Fill out an intake as though the answer were "Yes". CSS will help you decide.
8-
9-
*Then remove this first heading from this SUPPORT.MD file before publishing your repo.*
10-
111
# Support
122

13-
## How to file issues and get help
3+
## How to file issues and get help
144

15-
This project uses GitHub Issues to track bugs and feature requests. Please search the existing
16-
issues before filing new issues to avoid duplicates. For new issues, file your bug or
17-
feature request as a new Issue.
5+
This project uses GitHub Issues to track bugs and feature requests. Please search existing
6+
issues before filing a new report, and include enough detail to reproduce the problem.
187

19-
For help and questions about using this project, please **REPO MAINTAINER: INSERT INSTRUCTIONS HERE
20-
FOR HOW TO ENGAGE REPO OWNERS OR COMMUNITY FOR HELP. COULD BE A STACK OVERFLOW TAG OR OTHER
21-
CHANNEL. WHERE WILL YOU HELP PEOPLE?**.
8+
Usage questions and reproducibility discussions should also go through GitHub Issues.
9+
For security-sensitive reports, follow the process in [SECURITY.md](SECURITY.md) instead of
10+
opening a public issue.
2211

23-
## Microsoft Support Policy
12+
## Microsoft support policy
2413

25-
Support for this **PROJECT or PRODUCT** is limited to the resources listed above.
14+
Support for this project is limited to the resources listed above.

‎assets/pipeline.png‎

272 KB
Loading

‎assets/teaser.png‎

762 KB
Loading

0 commit comments

Comments
 (0)