Text-guided style transfer aims to stylize a content image according to a textual description while preserving the image structure. Unfortunately, many recent approaches rely on either attention-based fusion or diffusion backbones. On one hand, transformers require high memory consumption due to their quadratic complexity, while diffusion-based methods require iterative denoising, making inference slow and computationally expensive. In this work, we propose a lightweight architecture based on Mamba state-space models. Our key idea is to perform cross-modal fusion inside the Mamba dynamics by conditioning the state-space parameters on the text embedding while processing content tokens, avoiding attention and enabling linear-complexity fusion. Additionally, we represent the content image as a sequence of multi-scale frozen VGG features instead of learned patch embeddings, reducing sequence length and improving content encoding. Experiments at 512×512 resolution show that our method achieves strong prompt alignment while maintaining high structural similarity to the input and it provides a favorable quality-efficiency trade-off compared to open baselines. Code is available at https://github.com/FilippoBotti/TGSTM
Revisiting Text-Guided Style Transfer Through Mamba: A Fast and Memory-Efficient Architecture / Botti, F., Calzetti, T., Ergasti, A., Fontanini, T., Prati, A.. - In: IEEE ACCESS. - ISSN 2169-3536. - 14:(2026), pp. 74541-74552. [10.1109/ACCESS.2026.3692952]
Revisiting Text-Guided Style Transfer Through Mamba: A Fast and Memory-Efficient Architecture
Botti F.
;Calzetti T.;Ergasti A.;Fontanini T.;Prati A.
2026-01-01
Abstract
Text-guided style transfer aims to stylize a content image according to a textual description while preserving the image structure. Unfortunately, many recent approaches rely on either attention-based fusion or diffusion backbones. On one hand, transformers require high memory consumption due to their quadratic complexity, while diffusion-based methods require iterative denoising, making inference slow and computationally expensive. In this work, we propose a lightweight architecture based on Mamba state-space models. Our key idea is to perform cross-modal fusion inside the Mamba dynamics by conditioning the state-space parameters on the text embedding while processing content tokens, avoiding attention and enabling linear-complexity fusion. Additionally, we represent the content image as a sequence of multi-scale frozen VGG features instead of learned patch embeddings, reducing sequence length and improving content encoding. Experiments at 512×512 resolution show that our method achieves strong prompt alignment while maintaining high structural similarity to the input and it provides a favorable quality-efficiency trade-off compared to open baselines. Code is available at https://github.com/FilippoBotti/TGSTMI documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


