Benedikt Hilmes scite author profile

Benedikt Hilmes

3Publications

4Citation Statements Received

91Citation Statements Given

How they've been cited

How they cite others

Affiliations

RWTH Aachen University, United States Military Academy

Publications

Order By: Most citations

Comparing the Benefit of Synthetic Training Data for Various Automatic Speech Recognition Architectures

Rossenbach¹,

Zeineldeen²,

Hilmes³

et al. 2021

Preprint

View full text Add to dashboard Cite

Recent publications on automatic-speech-recognition (ASR) have a strong focus on attention encoder-decoder (AED) architectures which work well for large datasets, but tend to overfit when applied in low resource scenarios. One solution to tackle this issue is to generate synthetic data with a trained text-tospeech system (TTS) if additional text is available. This was successfully applied in many publications with AED systems. We present a novel approach of silence correction in the data pre-processing for TTS systems which increases the robustness when training on corpora targeted for ASR applications. In this work we do not only show the successful application of synthetic data for AED systems, but also test the same method on a highly optimized state-of-the-art Hybrid ASR system and a competitive monophone based system using connectionisttemporal-classification (CTC). We show that for the later systems the addition of synthetic data only has a minor effect, but they still outperform the AED systems by a large margin on LibriSpeech-100h. We achieve a final word-error-rate of 3.3%/10.0% with a Hybrid system on the clean/noisy test-sets, surpassing any previous state-of-the-art systems that do not include unlabeled audio data.

show abstract

Comparing the Benefit of Synthetic Training Data for Various Automatic Speech Recognition Architectures

Rossenbach

Zeineldeen

Hilmes

et al. 2021

View full text Add to dashboard Cite

Synthetic data generated by text-to-speech (TTS) systems can be used to improve automatic speech recognition (ASR) systems in low-resource or domain mismatch tasks. It has been shown that TTS-generated outputs still do not have the same qualities as real data. In this work we focus on the temporal structure of synthetic data and its relation to ASR training. By using a novel oracle setup we show how much the degradation of synthetic data quality is influenced by duration modeling in non-autoregressive (NAR) TTS. To get reference phoneme durations we use two common alignment methods, a hidden Markov Gaussian-mixture model (HMM-GMM) aligner and a neural connectionist temporal classification (CTC) aligner. Using a simple algorithm based on random walks we shift phoneme duration distributions of the TTS system closer to real durations, resulting in an improvement of an ASR system using synthetic data in a semi-supervised setting.

show abstract

Using local area networks to improve learning efficiency

Hilmes

View full text Add to dashboard Cite

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

hi@scite.ai

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Benedikt Hilmes

Comparing the Benefit of Synthetic Training Data for Various Automatic Speech Recognition Architectures

Comparing the Benefit of Synthetic Training Data for Various Automatic Speech Recognition Architectures

Using local area networks to improve learning efficiency

Contact Info

Product

Resources

About