Joint audio-video (AV) generation is still a significant challenge in generative AI, primarily due to four critical requirements: quality of the generated samples, seamless multimodal synchronization, with audio tracks that match the visual data and vice versa, limitless video duration, and training and inference efficiency. In this paper, we present TECMO, a novel transformer-based architecture that addresses these key challenges of AV generation. After exploring three distinct cross-modality interaction modules, we propose a novel lightweight fusion module which emerged as the most effective and computationally efficient approach for aligning audio and visual modalities. We trained our model on AIST++ and Landscape datasets measuring metrics for both image quality (FVD and KVD) as well as audio quality (FAD) and AV syncronization (CAVP). Our experimental results demonstrate that TECMO outperforms existing state-of-the-art models in multimodal AV generation tasks.

TECMO: Efficient Temporal Cross-Interaction Module for Long Synchronized Audio-Video Generation / Ergasti, A., Tarollo, G., Botti, F., Fontanini, T., Ferrari, C., Bertozzi, M., Prati, A.. - In: IEEE TRANSACTIONS ON MULTIMEDIA. - ISSN 1520-9210. - (In corso di stampa), pp. 1-12. [10.1109/tmm.2026.3717470]

TECMO: Efficient Temporal Cross-Interaction Module for Long Synchronized Audio-Video Generation

Ergasti, Alex;Tarollo, Giuseppe;Botti, Filippo;Fontanini, Tomaso;Bertozzi, Massimo;Prati, Andrea
In corso di stampa

Abstract

Joint audio-video (AV) generation is still a significant challenge in generative AI, primarily due to four critical requirements: quality of the generated samples, seamless multimodal synchronization, with audio tracks that match the visual data and vice versa, limitless video duration, and training and inference efficiency. In this paper, we present TECMO, a novel transformer-based architecture that addresses these key challenges of AV generation. After exploring three distinct cross-modality interaction modules, we propose a novel lightweight fusion module which emerged as the most effective and computationally efficient approach for aligning audio and visual modalities. We trained our model on AIST++ and Landscape datasets measuring metrics for both image quality (FVD and KVD) as well as audio quality (FAD) and AV syncronization (CAVP). Our experimental results demonstrate that TECMO outperforms existing state-of-the-art models in multimodal AV generation tasks.
In corso di stampa
TECMO: Efficient Temporal Cross-Interaction Module for Long Synchronized Audio-Video Generation / Ergasti, A., Tarollo, G., Botti, F., Fontanini, T., Ferrari, C., Bertozzi, M., Prati, A.. - In: IEEE TRANSACTIONS ON MULTIMEDIA. - ISSN 1520-9210. - (In corso di stampa), pp. 1-12. [10.1109/tmm.2026.3717470]
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11381/3068274
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact