An Improved Deep Neural Network for Modeling Speaker Characteristics at Different Temporal Scales

Gu, Bin; Guo, Wenbin; Dai, Li-Rong; Du, Jun

doi:10.1109/icassp40776.2020.9054151

Cited by 7 publications

(2 citation statements)

References 26 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…• Single-head Baum-Welch statistics attention mechanism based statistics pooling [116]: To overcome the weakness of ( 12) which cannot fully mine the inner relationship between an utterance and its frames, [116] integrated the Baum-Welch statistics into the attention mechanism:…”

Section: Attention Pooling Methodsmentioning

confidence: 99%

See 1 more Smart Citation

Speaker Recognition Based on Deep Learning: An Overview

Bai

Zhang

2020

Preprint

View full text Add to dashboard Cite

Speaker recognition is a task of identifying persons from their voices. Recently, deep learning has dramatically revolutionized speaker recognition. However, there is lack of comprehensive reviews on the exciting progress. In this paper, we review several major subtasks of speaker recognition, including speaker verification, identification, diarization, and robust speaker recognition, with a focus on deep-learning-based methods. Because the major advantage of deep learning over conventional methods is its representation ability, which is able to produce highly abstract embedding features from utterances, we first pay close attention to deep-learning-based speaker feature extraction, including the inputs, network structures, temporal pooling strategies, and objective functions respectively, which are the fundamental components of many speaker recognition subtasks. Then, we make an overview of speaker diarization, with an emphasis of recent supervised, end-to-end, and online diarization. Finally, we survey robust speaker recognition from the perspectives of domain adaptation and speech enhancement, which are two major approaches of dealing with domain mismatch and noise problems. Popular and recently released corpora are listed at the end of the paper.

show abstract

Section: Attention Pooling Methodsmentioning

confidence: 99%

“…The key matrix K is calculated from the Baum-Welch statistics. Specifically, [116] first calculates the normalized first order statistics f c from the cth component of a GMM-UBM model Ω (see ( 5)), and then conducts the following nonlinear transform:…”

Section: Attention Pooling Methodsmentioning

confidence: 99%