Multi-stream Asynchrony Modeling for Audio-Visual Speech Recognition

Lv, Guoyun; Jiang, Dongmei; Zhao, Ran; Hou, Yanze

doi:10.1109/ism.2007.4412354

Cited by 7 publications

(1 citation statement)

References 15 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…Usually the visual features are handled in a separate stream in a multi-stream HMM (MSHMM) [3]. However, since regular multi-stream HMM handles both channels synchronously and there may be asynchrony between audio and video channels, some extensions like coupled hidden Markov models (CHMM) [4], product hidden Markov models (PHMM) [3] and multi-stream asynchrony dynamic Bayesian networks [5] have been proposed to take this asynchrony into consideration. Also as This research is supported by The Scientific and Technological Research Council of Turkey (TUBITAK) under the scientific and technological research support program (code 1001), project number 107E015 entitled "Novel Approaches in Audio Visual Speech Recognition".…”

Section: Introductionmentioning

confidence: 99%

Using multiple visual tandem streams in audio-visual speech recognition

Topkaya

Erdoğan

2011

2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

View full text Add to dashboard Cite

The method which is called the "tandem approach" in speech recognition has been shown to increase performance by using classifier posterior probabilities as observations in a hidden Markov model. We study the effect of using visual tandem features in audio-visual speech recognition using a novel setup which uses multiple classifiers to obtain multiple visual tandem features. We adopt the approach of multi-stream hidden Markov models where visual tandem features from two different classifiers are considered as additional streams in the model. It is shown in our experiments that using multiple visual tandem features improve the recognition accuracy in various noise conditions. In addition, in order to handle asynchrony between audio and visual observations, we employ coupled hidden Markov models and obtain improved performance as compared to the synchronous model.

show abstract

Section: Introductionmentioning

confidence: 99%

Using multiple visual tandem streams in audio-visual speech recognition

Topkaya

Erdoğan

2011

2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

View full text Add to dashboard Cite

show abstract

Speech driven realistic mouth animation based on multi-modal unit selection

Jiang

Ravyse

Sahli

et al. 2008

J Multimodal User Interfaces

Self Cite

View full text Add to dashboard Cite

This paper presents a novel audio visual diviseme (viseme pair) instance selection and concatenation method for speech driven photo realistic mouth animation. Firstly, an audio visual diviseme database is built consisting of the audio feature sequences, intensity sequences and visual feature sequences of the instances. In the Viterbi based diviseme instance selection, we set the accumulative cost as the weighted sum of three items: 1) logarithm of concatenation smoothness of the synthesized mouth trajectory; 2) logarithm of the pronunciation distance; 3) logarithm of the audio intensity distance between the candidate diviseme instance and the target diviseme segment in the incoming speech. The selected diviseme instances are time warped and blended to construct the mouth animation. Objective and subjective evaluations on the synthesized mouth animations prove that the multimodal diviseme instance selection algorithm proposed in this paper outperforms the triphone unit selection algorithm in Video Rewrite. Clear, accurate, smooth mouth animations can be obtained matching well with the pronunciation and intensity changes in the incoming speech. Moreover, with the logarithm function in the accumulative cost, it is easy to set the weights to obtain optimal mouth animations.

show abstract