Tom Ko scite author profile

This paper introduces a new method to extract speaker embeddings from a deep neural network (DNN) for text-independent speaker verification. Usually, speaker embeddings are extracted from a speaker-classification DNN that averages the hidden vectors over the frames of a speaker; the hidden vectors produced from all the frames are assumed to be equally important. We relax this assumption and compute the speaker embedding as a weighted average of a speaker's frame-level hidden vectors, and their weights are automatically determined by a self-attention mechanism. The effect of multiple attention heads are also investigated to capture different aspects of a speaker's input speech. Finally, a PLDA classifier is used to compare pairs of embeddings. The proposed self-attentive speaker embedding system is compared with a strong DNN embedding baseline on NIST SRE 2016. We find that the self-attentive embeddings achieve superior performance. Moreover, the improvement produced by the self-attentive speaker embeddings is consistent with both short and long testing utterances.

show abstract

JHU ASpIRE system: Robust LVCSR with TDNNS, iVector adaptation and RNN-LMS

Peddinti

Chen

Manohar

et al. 2015

View full text Add to dashboard Cite

An empirical exploration of CTC acoustic models

Miao

Gowayyed

et al. 2016

View full text Add to dashboard Cite

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

hi@scite.ai

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Tom Ko

Audio augmentation for speech recognition

A study on data augmentation of reverberant speech for robust speech recognition

Self-Attentive Speaker Embeddings for Text-Independent Speaker Verification

JHU ASpIRE system: Robust LVCSR with TDNNS, iVector adaptation and RNN-LMS

An empirical exploration of CTC acoustic models

Contact Info

Product

Resources

About