Reference:
- Wide Policy-Based optimization method support
- PPO
- DPO
- GRPO
- Value-Based Optimization methods support (Optional)
- Wide Environment Support
- Asynchronous execution of Trainer and Generator
- Zero-Bubble DDMA weight swapping (may require custom inference engine)
- KV-cache reuse after weight synchronization