Hello, I would like to ask a question. I have the recording data of 5000 speakers, and each person has multiple recording segments. I use a neural network to extract features from the speech, the speaker identity is used as the label, and a clustering network is trained using the method provided by you. Then use this clustering network to extract features from the audio clips of other people and cluster them. Is it feasible in theory? Thank you very much for your answer
Hello, I would like to ask a question. I have the recording data of 5000 speakers, and each person has multiple recording segments. I use a neural network to extract features from the speech, the speaker identity is used as the label, and a clustering network is trained using the method provided by you. Then use this clustering network to extract features from the audio clips of other people and cluster them. Is it feasible in theory? Thank you very much for your answer