Leveraging speaker attribute information using multi task learning for speaker verification and diarization

10/27/2020
by   Chau Luu, et al.
0

Deep speaker embeddings have become the leading method for encoding speaker identity in speaker recognition tasks. The embedding space should ideally capture the variations between all possible speakers, encoding the multiple aspects that make up speaker identity. In this work, utilizing speaker age as an auxiliary variable in US Supreme Court recordings and speaker nationality with VoxCeleb, we show that by leveraging additional speaker attribute information in a multi task learning setting, deep speaker embedding performance can be increased for verification and diarization tasks, achieving a relative improvement of 17.8 compared to omitting the auxiliary task. Experimental code has been made publicly available.

READ FULL TEXT

Please sign up or login with your details

Forgot password? Click here to reset