Correlation based speech-video synchronization

    Research output: Contribution to journalArticlepeer-review

    8 Citations (Scopus)


    This paper presents a novel Lip synchronization technique which investigates the correlation between the speech and lips movements. First, the speech signal is represented as a nonlinear time-varying model which involves a sum of AM–FM signals. Each of these signals is employed to model a single Formant frequency. The model is realized using Taylor series expansion in a way which provides the relationship between the lip shape (width and height) w.r.t. the speech amplitude and instantaneous frequency. Using lips width and height, a semi-speech signal is generated and correlated with the original speech signal over a span of delays then the delay between the speech and the video is estimated. Using real and noisy data from the VidTimit and in-house diastases, the proposed method was able to estimate small delays of 0.01–0.1 s in the case of noise-less and noisy signals respectively with a maximum absolute error of 0.0022 s.
    Original languageEnglish
    Pages (from-to)780 - 786
    JournalPattern Recognition Letters
    Issue number6
    Early online date9 Jan 2011
    Publication statusPublished - Apr 2011


    Dive into the research topics of 'Correlation based speech-video synchronization'. Together they form a unique fingerprint.

    Cite this