Jerome R. Bellegarda scite author profile

A new approach is proposed for the clustering of words in a given vocabulary. The method is based on a paradigm first, formulated in the context, of information retrieval, called latent semuntac unulysis. This paradigm leads to a parsimonious vector representation of each word in a suitable vector space, where familiar clustering techniques can be applied. The distance measure selected in this space arises naturally from the problem formulation. Preliminary experiments indicate that the clusters produced are intuitively satisfactory. Because these clusters are semantic in nature, this approach may prove useful as a complement, to conventional class-based statistical language modeling techniques.

show abstract

Large vocabulary speech recognition with multispan statistical language models

Bellegarda

2000

IEEE Trans. Speech Audio Process.

View full text Add to dashboard Cite

Multispan language modeling refers to the integration of the various constraints, both local and global, present in the language. It was recently proposed to capture global constraints through the use of latent semantic analysis, while taking local constraints into account via the usual n-gram approach. This has led to several families of data-driven, multispan language models for large vocabulary speech recognition. Because of the inherent complementarity in the two types of constraints, the multispan performance, as measured by perplexity, has been shown to compare favorably with the corresponding n-gram performance. The objective of this work is to characterize the behavior of such multispan modeling in actual recognition. Major implementation issues are addressed, including search integration and context scope selection. Experiments are conducted on a subset of the Wall Street Journal (WSJ) speaker-independent, 20 000-word vocabulary, continuous speech task. Results show that, compared to standard n-gram, the multispan framework can lead to a reduction in average word error rate of over 20%. The paper concludes with a discussion of intrinsic multi-span tradeoffs, such as the influence of training data selection on the resulting performance.

show abstract

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

hi@scite.ai

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Jerome R. Bellegarda

Exploiting latent semantic information in statistical language modeling

Statistical language model adaptation: review and perspectives

Tied mixture continuous parameter modeling for speech recognition

A novel word clustering algorithm based on latent semantic analysis

Large vocabulary speech recognition with multispan statistical language models

Contact Info

Product

Resources

About