Text-based person anomaly search (TPAS) refers to the task of retrieving people exhibiting normal or anomalous behaviors from natural language descriptions. Existing TPAS models often learn a single joint embedding where appearance, action, and background information are entangled, causing over-reliance on identity cues, poor alignment for action-centric queries, and limited semantic connection between actions and places where they occur. To address these issues, we propose Semantic Decoupled Alignment (SeDA), a disentangled vision–language retrieval framework that explicitly factorizes both visual and textual representations into appearance, action, and background components. SeDA introduces Semantic Token Projection, which derives three semantic queries from the global [CLS] token, softly aggregates modality tokens relevant to each factor, and recomposes the resulting factor tokens into a compact retrieval embedding. To enforce factor-specific semantics, we decompose each caption into appearance/action/background sub-captions and supervise the corresponding tokens with a Feature Decoupling Loss, combined with contrastive and image–text matching objectives. On the Person Anomaly Benchmark (1M pairs), SeDA achieves 86.45\% R@1 (+1.52 over SOTA), improves average multi-weather R@1 by +2.43, and gains +3.74 R@1 under out-of-distribution evaluation.
Divide and Align: Disentangled Vision-Language Learning for Text-Based Person Anomaly Search / Ergasti, A., Fontanini, T., Ferrari, C., Bertozzi, M., Prati, A.. - (2026), pp. 638-653. (19th European Conference on Computer Vision -- ECCV 2026 Malmo ) [10.1007/978-3-032-37595-7_35].
Divide and Align: Disentangled Vision-Language Learning for Text-Based Person Anomaly Search
Ergasti, Alex;Fontanini, Tomaso;Ferrari, Claudio;Bertozzi, Massimo;Prati, Andrea
2026-01-01
Abstract
Text-based person anomaly search (TPAS) refers to the task of retrieving people exhibiting normal or anomalous behaviors from natural language descriptions. Existing TPAS models often learn a single joint embedding where appearance, action, and background information are entangled, causing over-reliance on identity cues, poor alignment for action-centric queries, and limited semantic connection between actions and places where they occur. To address these issues, we propose Semantic Decoupled Alignment (SeDA), a disentangled vision–language retrieval framework that explicitly factorizes both visual and textual representations into appearance, action, and background components. SeDA introduces Semantic Token Projection, which derives three semantic queries from the global [CLS] token, softly aggregates modality tokens relevant to each factor, and recomposes the resulting factor tokens into a compact retrieval embedding. To enforce factor-specific semantics, we decompose each caption into appearance/action/background sub-captions and supervise the corresponding tokens with a Feature Decoupling Loss, combined with contrastive and image–text matching objectives. On the Person Anomaly Benchmark (1M pairs), SeDA achieves 86.45\% R@1 (+1.52 over SOTA), improves average multi-weather R@1 by +2.43, and gains +3.74 R@1 under out-of-distribution evaluation.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


