Transcription Correction Using Group Delay Processing for Continuous Speech Recognition

Brunet, R. Golda; Murthy, A. Hema

doi:10.1007/s00034-017-0598-2

Search citation statements

Order By: Relevance

Paper Sections

Select...

Citation Types

Supporting

Mentioning

Contrasting

Year Published

2018

2023

Publication Types

Select...

Article3

Other1

Relationship

Self Cite0

Independent4

Authors

Journals

Cited by 4 publications

References 25 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

A comprehensive survey on automatic speech recognition using neural networks

Dhanjal,

Singh

2023

Multimed Tools Appl

View full text Add to dashboard Cite

A comprehensive survey on automatic speech recognition using neural networks

Dhanjal,

Singh

2023

Multimed Tools Appl

View full text Add to dashboard Cite

Importance of Signal Processing Cues in Transcription Correction for Low-Resource Indian Languages

Prakash

Rajan²,

Murthy

2019

ACM Trans. Asian Low-Resour. Lang. Inf. Process.

View full text Add to dashboard Cite

Accurate phonetic transcriptions are crucial for building robust acoustic models for speech recognition as well as speech synthesis applications. Phonetic transcriptions are not usually provided with speech corpora. A lexicon is used to generate phone-level transcriptions of speech corpora with sentence-level transcriptions. When lexical entries are not available, letter-to-sound (LTS) rules are used. Whether it is a lexicon or LTS, the rules for pronunciation are generic and may not match the spoken utterance. This can lead to transcription errors. The objective of this study is to address the issue of mismatch between the transcription and its acoustic realisation. In particular, the issue of vowel deletions is studied. Group-delay-based segmentation is used to determine insertion/deletion of vowels in the speech utterance. The transcriptions are corrected in the training data based on this. The corrected data are used in automatic speech recognition (ASR) and text to speech synthesis (TTS) systems. ASR and TTS systems built with the corrected transcriptions show improvements in the performance.

show abstract

Transcription Correction for Indian Languages Using Acoustic Signatures

JPrakash¹,

Rajan²,

Murthy

2018

Interspeech 2018

View full text Add to dashboard Cite

Accurate phonetic transcription of the speech corpus has a significant impact on the performance of speech processing applications especially for low resource languages. Mismatches between the transcriptions and their utterances occur often at phoneme level due to insertion/deletion/substitution errors. This is very common in Indian languages owing to schwa deletion in the context of vowels, and agglutination in the context of consonants. An attempt is made in this paper to use acoustic cues at the syllable level to remove vowels from the transcription when they are poorly articulated or absent. Hidden Markov model (HMM) based forced Viterbi alignment (FVA) and group delay (GD) based signal processing are employed in tandem to achieve this task. Disagreement between FVA (which produces vowel boundaries based on transcription) and GD boundaries (which uses signal processing cues for syllables) are used to correct the transcription. An increase in likelihood of 0.3% is observed across 3 Indian languages, namely, Gujarati, Telugu and Tamil.

show abstract

Transcription Correction Using Group Delay Processing for Continuous Speech Recognition

Cited by 4 publications

References 25 publications

A comprehensive survey on automatic speech recognition using neural networks

A comprehensive survey on automatic speech recognition using neural networks

Importance of Signal Processing Cues in Transcription Correction for Low-Resource Indian Languages

Transcription Correction for Indian Languages Using Acoustic Signatures

Contact Info

Product

Resources

About