The loss you propose seems to be not differentiable w.r.t to models' parameters <img width="836" height="137" alt="Image" src="https://github.com/user-attachments/assets/32f07a19-dedf-43a0-9f19-d8b61cd046ff" /> the reward are by construction not differentiable wrt to models' parameters. How could you train your model in such conditions?
The loss you propose seems to be not differentiable w.r.t to models' parameters
the reward are by construction not differentiable wrt to models' parameters.
How could you train your model in such conditions?