Joint audio-video (AV) generation is still a significant challenge in generative AI, primarily due to four critical requirements: quality of the generated samples, seamless multimodal synchronization, with audio tracks that match the visual data and vice versa, limitless video duration, and training and inference efficiency. In this paper, we present TECMO, a novel transformer-based architecture that addresses these key challenges of AV generation. After exploring three distinct cross-modality interaction modules, we propose a novel lightweight fusion module which emerged as the most effective and computationally efficient approach for aligning audio and visual modalities. We trained our model on AIST++ and Landscape datasets measuring metrics for both image quality (FVD and KVD) as well as audio quality (FAD) and AV syncronization (CAVP). Our experimental results demonstrate that TECMO outperforms existing state-of-the-art models in multimodal AV generation tasks.
TECMO: Efficient Temporal Cross-Interaction Module for Long Synchronized Audio-Video Generation / Ergasti, A., Tarollo, G., Botti, F., Fontanini, T., Ferrari, C., Bertozzi, M., Prati, A.. - In: IEEE TRANSACTIONS ON MULTIMEDIA. - ISSN 1520-9210. - (In corso di stampa), pp. 1-12. [10.1109/tmm.2026.3717470]
TECMO: Efficient Temporal Cross-Interaction Module for Long Synchronized Audio-Video Generation
Ergasti, Alex;Tarollo, Giuseppe;Botti, Filippo;Fontanini, Tomaso;Bertozzi, Massimo;Prati, Andrea
In corso di stampa
Abstract
Joint audio-video (AV) generation is still a significant challenge in generative AI, primarily due to four critical requirements: quality of the generated samples, seamless multimodal synchronization, with audio tracks that match the visual data and vice versa, limitless video duration, and training and inference efficiency. In this paper, we present TECMO, a novel transformer-based architecture that addresses these key challenges of AV generation. After exploring three distinct cross-modality interaction modules, we propose a novel lightweight fusion module which emerged as the most effective and computationally efficient approach for aligning audio and visual modalities. We trained our model on AIST++ and Landscape datasets measuring metrics for both image quality (FVD and KVD) as well as audio quality (FAD) and AV syncronization (CAVP). Our experimental results demonstrate that TECMO outperforms existing state-of-the-art models in multimodal AV generation tasks.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


