Streaming End-to-End Target-Speaker Automatic Speech Recognition and Activity Detection

Moriya, Takafumi; Satō, Hiroshi; Ochiai, Tsubasa; Delcroix, Marc; Shinozaki, Takahiro

doi:10.1109/access.2023.3243690

Cited by 1 publication

(1 citation statement)

References 42 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…Target speaker tracking is employed in many speech-related tasks to retrieve the information of a specific speaker, including target speaker automatic speech recognition (TS-ASR) [55], target speaker speech separation [56] and TS-VAD [19]. The similarity of these tasks is that they all use a speaker profile to focus on the speech of interest, which consequently refines results corresponding to that particular speaker.…”

Section: Target Speaker Voice Activity Detectionmentioning

confidence: 99%

Robust End-to-end Speaker Diarization with Generic Neural Clustering

Yang¹,

Wang²

2022

Interspeech 2022

View full text Add to dashboard Cite

This paper proposes an online target speaker voice activity detection system for speaker diarization tasks, which does not require a priori knowledge from the clustering-based diarization system to obtain the target speaker embeddings. By adapting the conventional target speaker voice activity detection for real-time operation, this framework can identify speaker activities using self-generated embeddings, resulting in consistent performance without permutation inconsistencies in the inference phase. During the inference process, we employ a front-end model to extract the frame-level speaker embeddings for each coming block of a signal. Next, we predict the detection state of each speaker based on these frame-level speaker embeddings and the previously estimated target speaker embedding. Then, the target speaker embeddings are updated by aggregating these framelevel speaker embeddings according to the predictions in the current block. Our model predicts the results for each block and updates the target speakers' embeddings until reaching the end of the signal. Experimental results show that the proposed method outperforms the offline clustering-based diarization system on the DIHARD III and AliMeeting datasets. The proposed method is further extended to multi-channel data, which achieves similar performance with the state-of-the-art offline diarization systems.

show abstract

Section: Target Speaker Voice Activity Detectionmentioning

confidence: 99%