Thanks for your excellent work!
I am confused about the definition of open-vocabulary segmentation from two aspects:
- I note that the segmentation model (i.e., maskformer in the paper) is trained on full categories of PASCAL VOC and COCO while the data are synthetic from the Stable Diffusion.
- Can open-vocabulary segmentation protocol access the complete categories during training? In my opinion, the unseen(novel) class name should only be available at the test instead of training time. Otherwise, it is not really open-vocabulary.
Hope the authors could give me some help to make me better understand this paper!
Thanks!
Thanks for your excellent work!
I am confused about the definition of open-vocabulary segmentation from two aspects:
Hope the authors could give me some help to make me better understand this paper!
Thanks!