|
1 | | -# Project |
| 1 | +# World-R1: Reinforcing 3D Constraints for Text-to-Video Generation |
2 | 2 |
|
3 | | -> This repo has been populated by an initial template to help get you started. Please |
4 | | -> make sure to update the content to build a great experience for community-building. |
| 3 | +<p align="center"> |
| 4 | + Weijie Wang<sup>*†1,2</sup> |
| 5 | + Xiaoxuan He<sup>*1</sup> |
| 6 | + Youping Gu<sup>*1</sup> |
| 7 | + Zeyu Zhang<sup>3</sup> |
| 8 | + Yefei He<sup>1</sup> |
| 9 | + Yanbo Ding<sup>2</sup> |
| 10 | + Xirui Hu<sup>3</sup> |
| 11 | + Donny Y. Chen<sup>3</sup> |
| 12 | + Zhiyuan He<sup>2</sup> |
| 13 | + Yuqing Yang<sup>2</sup> |
| 14 | + Yifan Yang<sup>‡2</sup> |
| 15 | + Bohan Zhuang<sup>‡1</sup> |
| 16 | +</p> |
5 | 17 |
|
6 | | -As the maintainer of this project, please make a few updates: |
| 18 | +<p align="center"> |
| 19 | + <sup>1</sup>Zhejiang University |
| 20 | + <sup>2</sup>Microsoft Research |
| 21 | + <sup>3</sup>Independent Researcher |
| 22 | +</p> |
7 | 23 |
|
8 | | -- Improving this README.MD file to provide a great experience |
9 | | -- Updating SUPPORT.MD with content about this project's support experience |
10 | | -- Understanding the security reporting process in SECURITY.MD |
11 | | -- Remove this section from the README |
| 24 | +<p align="center"> |
| 25 | + <sup>*</sup>Equal contribution |
| 26 | + <sup>†</sup>Work done during an internship at MSRA |
| 27 | + <sup>‡</sup>Corresponding authors |
| 28 | +</p> |
| 29 | + |
| 30 | +<p align="center"> |
| 31 | + <a href="https://aka.ms/world-r1">Project Page</a> |
| 32 | +</p> |
| 33 | + |
| 34 | +<p align="center"> |
| 35 | + <img src="assets/teaser.png" alt="World-R1 teaser" width="100%"> |
| 36 | +</p> |
| 37 | + |
| 38 | +World-R1 aligns text-to-video generation with 3D constraints through reinforcement learning. Instead of changing the base video model architecture or relying on large-scale 3D supervision, it combines camera-aware latent initialization, 3D-aware rewards from pre-trained foundation models, and a periodic decoupled training strategy to improve geometric consistency while preserving visual quality and motion diversity. |
| 39 | + |
| 40 | +## Highlights |
| 41 | + |
| 42 | +- RL post-training for video foundation models using Flow-GRPO, without architecture changes at inference time. |
| 43 | +- A composite reward that combines reconstruction fidelity, trajectory alignment, meta-view semantic scoring, and general visual quality. |
| 44 | +- A pure-text training set with about 3,000 prompts, including a dynamic subset for regularizing non-rigid motion. |
| 45 | +- Strong 3D consistency gains reported in the paper: `+10.23 dB` PSNR over Wan2.1-T2V-1.3B and `+7.91 dB` over Wan2.1-T2V-14B. |
| 46 | + |
| 47 | +## Method |
| 48 | + |
| 49 | +<p align="center"> |
| 50 | + <img src="assets/pipeline.png" alt="World-R1 pipeline" width="100%"> |
| 51 | +</p> |
| 52 | + |
| 53 | +World-R1 first converts camera instructions in text prompts into explicit trajectories and injects the motion prior into the initial video latents through noise wrapping. During RL fine-tuning, the model is optimized with 3D-aware feedback from reconstruction and camera-control metrics, together with a general visual reward. A periodic dynamic-only phase prevents the model from overfitting to rigid static scenes. |
| 54 | + |
| 55 | +## Results |
| 56 | + |
| 57 | +| Model | PSNR ↑ | SSIM ↑ | LPIPS ↓ | |
| 58 | +| --- | ---: | ---: | ---: | |
| 59 | +| Wan2.1-T2V-1.3B | 17.40 | 0.550 | 0.467 | |
| 60 | +| World-R1-Small | 27.63 | 0.858 | 0.201 | |
| 61 | +| Wan2.1-T2V-14B | 19.76 | 0.629 | 0.405 | |
| 62 | +| World-R1-Large | 27.67 | 0.865 | 0.162 | |
| 63 | + |
| 64 | +## Release Contents |
| 65 | + |
| 66 | +- Training code, reward functions, and launch scripts for the World-R1 RL pipeline. |
| 67 | +- The 3D reward server and bundled `Depth Anything 3` integration used by the release. |
| 68 | +- Prompt-only dataset files under `dataset/final` and `dataset/enhanced`. |
| 69 | +- Utility scripts for inference, ablations, prompt processing, and noise-wrap analysis. |
| 70 | +- Third-party license files under `licenses/`. |
| 71 | + |
| 72 | +This repository does not include released World-R1 checkpoints. Base video model checkpoints should be obtained separately. |
| 73 | + |
| 74 | +## Repository Layout |
| 75 | + |
| 76 | +```text |
| 77 | +World-R1/ |
| 78 | +├── config/ |
| 79 | +├── dataset/ |
| 80 | +│ ├── enhanced/ |
| 81 | +│ └── final/ |
| 82 | +├── flow_grpo/ |
| 83 | +├── licenses/ |
| 84 | +├── reward_server/ |
| 85 | +├── scripts/ |
| 86 | +├── assets/ |
| 87 | +└── pyproject.toml |
| 88 | +``` |
| 89 | + |
| 90 | +## Setup |
| 91 | + |
| 92 | +Use a Python 3.10+ environment with CUDA and a PyTorch build that matches your driver. A practical setup flow is: |
| 93 | + |
| 94 | +1. Create and activate a clean environment: |
| 95 | + |
| 96 | +```bash |
| 97 | +conda create -n world-r1 python=3.10 -y |
| 98 | +conda activate world-r1 |
| 99 | +pip install --upgrade pip |
| 100 | +``` |
| 101 | + |
| 102 | +2. Install PyTorch first. Replace the wheel index below with the one that matches your CUDA version: |
| 103 | + |
| 104 | +```bash |
| 105 | +pip install torch torchvision --index-url https://download.pytorch.org/whl/cu124 |
| 106 | +``` |
| 107 | + |
| 108 | +3. Install the repository and the core runtime packages used by training and inference: |
| 109 | + |
| 110 | +```bash |
| 111 | +pip install -e . |
| 112 | +pip install accelerate diffusers transformers peft wandb absl-py ml-collections \ |
| 113 | + numpy pillow imageio tqdm requests httpx flask addict omegaconf einops \ |
| 114 | + ftfy sentencepiece protobuf scipy opencv-python huggingface_hub |
| 115 | +``` |
| 116 | + |
| 117 | +4. Install the extra packages used by the 3D reward stack and visualization utilities: |
| 118 | + |
| 119 | +```bash |
| 120 | +pip install lpips trimesh plyfile moviepy pycolmap gsplat evo e3nn hpsv2 qwen-vl-utils |
| 121 | +``` |
| 122 | + |
| 123 | +5. Optional acceleration packages, depending on your machine and base model setup: |
| 124 | + |
| 125 | +```bash |
| 126 | +pip install xformers bitsandbytes |
| 127 | +``` |
| 128 | + |
| 129 | +The launcher scripts default to `python3` and `torchrun` from `PATH`. If your environment uses different executables, set them explicitly: |
| 130 | + |
| 131 | +```bash |
| 132 | +export WAN_PYTHON=$(which python) |
| 133 | +export WAN_TORCHRUN=$(which torchrun) |
| 134 | +``` |
| 135 | + |
| 136 | +For WAN-based utilities, you can either pass `--model /path/to/checkpoint` or set: |
| 137 | + |
| 138 | +```bash |
| 139 | +export WORLD_R1_WAN_MODEL=/path/to/Wan-Diffusers-checkpoint |
| 140 | +``` |
| 141 | + |
| 142 | +A simple smoke test after setup: |
| 143 | + |
| 144 | +```bash |
| 145 | +python -c "import torch, diffusers, transformers, peft, flask, lpips; print('env ok')" |
| 146 | +``` |
| 147 | + |
| 148 | +## Training |
| 149 | + |
| 150 | +Single-node training with local reward servers: |
| 151 | + |
| 152 | +```bash |
| 153 | +MODEL_PATH=/path/to/Wan2.1-T2V-14B-Diffusers \ |
| 154 | +SERVER_VISIBLE_DEVICES=0,1 \ |
| 155 | +TRAIN_VISIBLE_DEVICES=2,3,4,5,6,7 \ |
| 156 | +NUM_PROCESSES=6 \ |
| 157 | +bash scripts/run_single_node.sh |
| 158 | +``` |
| 159 | + |
| 160 | +CogVideoX training: |
| 161 | + |
| 162 | +```bash |
| 163 | +MODEL_FAMILY=cogvideox \ |
| 164 | +MODEL_PATH=THUDM/CogVideoX1.5-5B \ |
| 165 | +TRAIN_CONFIG=config/world_r1.py:world_r1_cogvideox_5b \ |
| 166 | +SERVER_VISIBLE_DEVICES=0,1 \ |
| 167 | +TRAIN_VISIBLE_DEVICES=2,3,4,5,6,7 \ |
| 168 | +NUM_PROCESSES=6 \ |
| 169 | +bash scripts/run_single_node.sh |
| 170 | +``` |
| 171 | + |
| 172 | +If reward servers are already running, launch training directly: |
| 173 | + |
| 174 | +```bash |
| 175 | +REWARD_3D_SERVER_URL=http://127.0.0.1:18089 \ |
| 176 | +GENERAL_REWARD_SERVER_URL=http://127.0.0.1:18090 \ |
| 177 | +MODEL_PATH=/path/to/Wan2.1-T2V-14B-Diffusers \ |
| 178 | +NUM_PROCESSES=6 \ |
| 179 | +bash scripts/run_training.sh |
| 180 | +``` |
| 181 | + |
| 182 | +For multi-node training, start external reward servers first and then provide `REWARD_3D_SERVER_URL`, `GENERAL_REWARD_SERVER_URL`, `MASTER_ADDR`, `MASTER_PORT`, `NNODES`, and `NODE_RANK` before running `scripts/run_multi_node.sh`. |
| 183 | + |
| 184 | +## Reward Servers and Utilities |
| 185 | + |
| 186 | +Start the two release reward services independently: |
| 187 | + |
| 188 | +```bash |
| 189 | +bash scripts/run_reward_3d_server.sh |
| 190 | +bash scripts/run_general_reward_server.sh |
| 191 | +``` |
| 192 | + |
| 193 | +Useful utilities included in this release: |
| 194 | + |
| 195 | +- `scripts/train_world_r1.py`: main RL training entry point. |
| 196 | +- `scripts/infer_wan_lora.py`: batch inference for WAN checkpoints or LoRAs. |
| 197 | +- `scripts/noise_wrap_ablation.py`: visualize latent wrapping and generation effects. |
| 198 | +- `scripts/noise_wrap_strength_sweep.py`: sweep wrap strengths over a prompt set. |
| 199 | +- `scripts/serve_reward_3d.py` and `scripts/serve_general_reward.py`: reward service backends. |
| 200 | + |
| 201 | +## Dataset |
| 202 | + |
| 203 | +The repository ships prompt-only data used by the release: |
| 204 | + |
| 205 | +- `dataset/final/`: the base prompt split used for training, validation, and dynamic regularization. |
| 206 | +- `dataset/enhanced/`: expanded prompt variants used for richer post-training supervision. |
| 207 | + |
| 208 | +The prompt-processing helpers under `scripts/` are configured to read and write these repository-local directories by default. |
| 209 | + |
| 210 | +## License |
| 211 | + |
| 212 | +This project is licensed under the [MIT License](LICENSE). |
| 213 | + |
| 214 | +The top-level `licenses/` directory is reserved for bundled third-party source code that remains under its upstream license: |
| 215 | + |
| 216 | +- `flow_grpo/` includes bundled or adapted `Flow-GRPO` code. See [`licenses/FLOW_GRPO_LICENSE`](licenses/FLOW_GRPO_LICENSE). |
| 217 | +- `reward_server/depth_anything_3/` includes bundled or modified `Depth Anything 3` code. See [`licenses/DEPTH_ANYTHING_3_LICENSE`](licenses/DEPTH_ANYTHING_3_LICENSE). |
| 218 | + |
| 219 | +Unless a file states otherwise, the rest of this repository is covered by the root MIT license. Please review the third-party license files before redistributing derivative work based on the bundled upstream code. |
| 220 | + |
| 221 | +## Citation |
| 222 | + |
| 223 | +If you find this repository useful, please cite: |
| 224 | + |
| 225 | +```bibtex |
| 226 | +@misc{wang2026worldr1, |
| 227 | + title={World-R1: Reinforcing 3D Constraints for Text-to-Video Generation}, |
| 228 | + author={Weijie Wang and Xiaoxuan He and Youping Gu and Zeyu Zhang and Yefei He and Yanbo Ding and Xirui Hu and Donny Y. Chen and Zhiyuan He and Yuqing Yang and Yifan Yang and Bohan Zhuang}, |
| 229 | + year={2026}, |
| 230 | + note={ICML 2026} |
| 231 | +} |
| 232 | +``` |
| 233 | + |
| 234 | +## Acknowledgements |
| 235 | + |
| 236 | +World-R1 builds on top of several strong open-source projects and model ecosystems, including Wan, CogVideoX, Flow-GRPO, and Depth Anything 3. We thank the original authors and maintainers for making those foundations available. |
| 237 | + |
| 238 | +## Support |
| 239 | + |
| 240 | +See [SUPPORT.md](SUPPORT.md) for usage support and [SECURITY.md](SECURITY.md) for vulnerability reporting. |
12 | 241 |
|
13 | 242 | ## Contributing |
14 | 243 |
|
15 | | -This project welcomes contributions and suggestions. Most contributions require you to agree to a |
| 244 | +This project welcomes contributions and suggestions. Most contributions require you to agree to a |
16 | 245 | Contributor License Agreement (CLA) declaring that you have the right to, and actually do, grant us |
17 | 246 | the rights to use your contribution. For details, visit [Contributor License Agreements](https://cla.opensource.microsoft.com). |
18 | 247 |
|
19 | 248 | When you submit a pull request, a CLA bot will automatically determine whether you need to provide |
20 | | -a CLA and decorate the PR appropriately (e.g., status check, comment). Simply follow the instructions |
21 | | -provided by the bot. You will only need to do this once across all repos using our CLA. |
| 249 | +a CLA and decorate the PR appropriately. Simply follow the instructions provided by the bot. You will |
| 250 | +only need to do this once across all repos using our CLA. |
22 | 251 |
|
23 | 252 | This project has adopted the [Microsoft Open Source Code of Conduct](https://opensource.microsoft.com/codeofconduct/). |
24 | | -For more information see the [Code of Conduct FAQ](https://opensource.microsoft.com/codeofconduct/faq/) or |
25 | | -contact [opencode@microsoft.com](mailto:opencode@microsoft.com) with any additional questions or comments. |
| 253 | +For more information, see the [Code of Conduct FAQ](https://opensource.microsoft.com/codeofconduct/faq/) or |
| 254 | +contact [opencode@microsoft.com](mailto:opencode@microsoft.com) with additional questions or comments. |
26 | 255 |
|
27 | 256 | ## Trademarks |
28 | 257 |
|
|
0 commit comments