Text-based person anomaly search (TPAS) refers to the task of retrieving people exhibiting normal or anomalous behaviors from natural language descriptions. Existing TPAS models often learn a single joint embedding where appearance, action, and background information are entangled, causing over-reliance on identity cues, poor alignment for action-centric queries, and limited semantic connection between actions and places where they occur. To address these issues, we propose Semantic Decoupled Alignment (SeDA), a disentangled vision–language retrieval framework that explicitly factorizes both visual and textual representations into appearance, action, and background components. SeDA introduces Semantic Token Projection, which derives three semantic queries from the global [CLS] token, softly aggregates modality tokens relevant to each factor, and recomposes the resulting factor tokens into a compact retrieval embedding. To enforce factor-specific semantics, we decompose each caption into appearance/action/background sub-captions and supervise the corresponding tokens with a Feature Decoupling Loss, combined with contrastive and image–text matching objectives. On the Person Anomaly Benchmark (1M pairs), SeDA achieves 86.45\% R@1 (+1.52 over SOTA), improves average multi-weather R@1 by +2.43, and gains +3.74 R@1 under out-of-distribution evaluation.

Divide and Align: Disentangled Vision-Language Learning for Text-Based Person Anomaly Search / Ergasti, A., Fontanini, T., Ferrari, C., Bertozzi, M., Prati, A.. - (2026), pp. 638-653. (19th European Conference on Computer Vision -- ECCV 2026 Malmo ) [10.1007/978-3-032-37595-7_35].

Divide and Align: Disentangled Vision-Language Learning for Text-Based Person Anomaly Search

Ergasti, Alex;Fontanini, Tomaso;Ferrari, Claudio;Bertozzi, Massimo;Prati, Andrea
2026-01-01

Abstract

Text-based person anomaly search (TPAS) refers to the task of retrieving people exhibiting normal or anomalous behaviors from natural language descriptions. Existing TPAS models often learn a single joint embedding where appearance, action, and background information are entangled, causing over-reliance on identity cues, poor alignment for action-centric queries, and limited semantic connection between actions and places where they occur. To address these issues, we propose Semantic Decoupled Alignment (SeDA), a disentangled vision–language retrieval framework that explicitly factorizes both visual and textual representations into appearance, action, and background components. SeDA introduces Semantic Token Projection, which derives three semantic queries from the global [CLS] token, softly aggregates modality tokens relevant to each factor, and recomposes the resulting factor tokens into a compact retrieval embedding. To enforce factor-specific semantics, we decompose each caption into appearance/action/background sub-captions and supervise the corresponding tokens with a Feature Decoupling Loss, combined with contrastive and image–text matching objectives. On the Person Anomaly Benchmark (1M pairs), SeDA achieves 86.45\% R@1 (+1.52 over SOTA), improves average multi-weather R@1 by +2.43, and gains +3.74 R@1 under out-of-distribution evaluation.
2026
9783032375940
9783032375957
Divide and Align: Disentangled Vision-Language Learning for Text-Based Person Anomaly Search / Ergasti, A., Fontanini, T., Ferrari, C., Bertozzi, M., Prati, A.. - (2026), pp. 638-653. (19th European Conference on Computer Vision -- ECCV 2026 Malmo ) [10.1007/978-3-032-37595-7_35].
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11381/3073694
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact