An adaptive a priori SNR estimator for perceptual speech enhancement

Nahma, Lara; Yong, Pei Chee; Dam, Hai Huyen; Nordholm, Sven

doi:10.1186/s13636-019-0150-3

Cited by 8 publications

(3 citation statements)

References 39 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…To tackle speech signals corrupted by noise, in this field, some of previous studies [4,5] tended to recover original signals by removing noise. Some methods [6,7] focused on feature extraction from un-corrupted voices, and some methods [8,9] tried to estimated speech quality by computing signal-to-noise ratio (SNR). Although speech enhancement has been used for speaker recognition, in most of previous studies it was often processed individually.…”

Section: Introductionmentioning

confidence: 99%

Robust Speaker Recognition Using Speech Enhancement And Attention Model

Shi¹,

Huang²,

Hain³

2020

The Speaker and Language Recognition Workshop (Odyssey 2020)

View full text Add to dashboard Cite

In this paper, a novel architecture for speaker recognition is proposed by cascading speech enhancement and speaker processing. It aims to improve speaker recognition performance when speech signals are corrupted by noise. Instead of separately processing speech enhancement and speaker recognition, the two modules are integrated into one framework by a joint optimisation using deep neural networks. Furthermore, to increase the robustness against noise, a multi-stage attention mechanism is employed to highlight the speaker related features learned from context information in both time and frequency domains. To evaluate speaker identification and verification performance of the proposed approach, VoxCeleb1, one of mostly used benchmark datasets, is used. Moreover, the robustness evaluation is also conducted on VoxCeleb1 when its being corrupted by three types of interferences, general noise, music, and babble, at different signal-to-noise ratio (SNR) levels. The obtained results show that the proposed approach using speech enhancement and multi-stage attention models outperforms two strong baselines in different acoustic conditions in our experiments.

show abstract

Section: Introductionmentioning

confidence: 99%

Robust Speaker Recognition Using Speech Enhancement And Attention Model

Shi¹,

Huang²,

Hain³

2020

The Speaker and Language Recognition Workshop (Odyssey 2020)

View full text Add to dashboard Cite

show abstract

“…Some of previous studies [4,5] tended to recover original signals by removing noise. Some methods [6,7] focused on feature extraction from un-corrupted speech signals, and some methods [8,9] tried to estimated speech quality by computing signal-to-noise ratio (SNR).…”

Section: Introductionmentioning

confidence: 99%

Robust Speaker Recognition Using Speech Enhancement And Attention Model

Shi¹,

Huang²,

Hain³

2020

Preprint

View full text Add to dashboard Cite

In this paper, a novel architecture for speaker recognition is proposed by cascading speech enhancement and speaker processing. Its aim is to improve speaker recognition performance when speech signals are corrupted by noise. Instead of individually processing speech enhancement and speaker recognition, the two modules are integrated into one framework by a joint optimisation using deep neural networks. Furthermore, to increase robustness against noise, a multi-stage attention mechanism is employed to highlight the speaker related features learned from context information in time and frequency domain. To evaluate speaker identification and verification performance of the proposed approach, we test it on the dataset of VoxCeleb1, one of mostly used benchmark datasets. Moreover, the robustness of our proposed approach is also tested on VoxCeleb1 data when being corrupted by three types of interferences, general noise, music, and babble, at different signal-tonoise ratio (SNR) levels. The obtained results show that the proposed approach using speech enhancement and multi-stage attention models outperforms two strong baselines not using them in most acoustic conditions in our experiments.

show abstract

“…Previous studies [6,7,8] tended to recover original signals by removing noise. Other methods [9,10,11] focused on feature extraction from uncorrupted speech signals, and further methods [12,13] tried to estimated speech quality by computing signal-to-noise ratios (SNRs).…”

Section: Introductionmentioning

confidence: 99%

Speaker Re-identification with Speaker Dependent Speech Enhancement

Shi

Huang

Hain

2020

Preprint

View full text Add to dashboard Cite

While the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments. Here speech enhancement methods have traditionally allowed improved performance. The recent works have shown that adapting speech enhancement can lead to further gains. This paper introduces a novel approach that cascades speech enhancement and speaker recognition. In the first step, a speaker embedding vector is generated , which is used in the second step to enhance the speech quality and re-identify the speakers. Models are trained in an integrated framework with joint optimisation. The proposed approach is evaluated using the Voxceleb1 dataset, which aims to assess speaker recognition in real world situations. In addition three types of noise at different signal-noise-ratios were added for this work. The obtained results show that the proposed approach using speaker dependent speech enhancement can yield better speaker recognition and speech enhancement performances than two baselines in various noise conditions.

show abstract

An adaptive a priori SNR estimator for perceptual speech enhancement

Cited by 8 publications

References 39 publications

Robust Speaker Recognition Using Speech Enhancement And Attention Model

Robust Speaker Recognition Using Speech Enhancement And Attention Model

Robust Speaker Recognition Using Speech Enhancement And Attention Model

Speaker Re-identification with Speaker Dependent Speech Enhancement

Contact Info

Product

Resources

About