Tomas Koctur scite author profile

This paper describes a new Slovak speech recognition dedicated corpus built from TEDx talks and Jump Slovakia lectures. The proposed speech database consists of 220 talks and lectures in total duration of about 58 hours. Annotated speech database was generated automatically in an unsupervised manner by using acoustic speech segmentation based on principal component analysis and automatic speech transcription using two complementary speech recognition systems. The evaluation data consisting of 50 manually annotated talks and lectures in total duration of about 12 hours, has been created for evaluation of the quality of Slovak speech recognition. By unsupervised automatic annotation of TEDx talks and Jump Slovakia lectures we have obtained 21.26% of new speech segments with approximately 9.44% word error rate, suitable for retraining or adaptation of acoustic models trained beforehand.

show abstract

Speech corpus generation based on N-gram confidence measure classification

Koctur

Ondáš

Juhar

2017

View full text Add to dashboard Cite

Unsupervised acoustic corpora building based on variable confidence measure thresholding

Koctur

Staš

Juhar

2016

View full text Add to dashboard Cite

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

hi@scite.ai

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Tomas Koctur

Unsupervised speech transcription and alignment based on two complementary ASR systems

Automatic Transcription and Subtitling of Slovak Multi-genre Audiovisual Recordings

TEDxSK and JumpSK: A New Slovak Speech Recognition Dedicated Corpus

Speech corpus generation based on N-gram confidence measure classification

Unsupervised acoustic corpora building based on variable confidence measure thresholding

Contact Info

Product

Resources

About