Hi, thank you for your excellent work. I have two questions regarding the training process.
- During training, original CUT3R modules are frozen. Why are CUT3R losses (conf loss, rgb loss, pose loss, etc) still used?
- How to make sure the predicted SMPLX meshes are aligned with the point clouds in scale? The model is only trained in BEDLAM to learn the SMPLX translation. Is it enough to let the model know the spatial relationship between the SMPLX mesh and the scene? Do you have any additional constraints?
Hi, thank you for your excellent work. I have two questions regarding the training process.