Text-based person anomaly search (TPAS) refers to the task of retrieving people exhibiting normal or anomalous behaviors from natural language descriptions. Existing TPAS models often learn a single joint embedding where appearance, action, and background information are entangled, causing over-reliance on identity cues, poor alignment for action-centric queries, and limited semantic connection between actions and places where they occur. To address these issues, we propose Semantic Decoupled Alignment (SeDA), a disentangled vision–language retrieval framework that explicitly factorizes both visual and textual representations into appearance, action, and background components. SeDA introduces Semantic Token Projection, which derives three semantic queries from the global [CLS] token, softly aggregates modality tokens relevant to each factor, and recomposes the resulting factor tokens into a compact retrieval embedding. To enforce factor-specific semantics, we decompose each caption into appearance/action/background sub-captions and supervise the corresponding tokens with a Feature Decoupling Loss, combined with contrastive and image–text matching objectives. On the Person Anomaly Benchmark (1M pairs), SeDA achieves 86.45% R@1 (+1.52 over SOTA), improves average multi-weather R@1 by +2.43, and gains +3.74 R@1 under out-of-distribution evaluation. Code is available in the supplementary materials.
Divide and Align: Disentangled Vision-Language Learning for Text-Based Person Anomaly Search / Ergasti, A., Fontanini, T., Ferrari, C., Bertozzi, M., Prati, A.. - (In corso di stampa). (19th European Conference on Computer Vision ).
Divide and Align: Disentangled Vision-Language Learning for Text-Based Person Anomaly Search
Alex Ergasti;Tomaso Fontanini;Claudio Ferrari;Massimo Bertozzi;Andrea Prati
In corso di stampa
Abstract
Text-based person anomaly search (TPAS) refers to the task of retrieving people exhibiting normal or anomalous behaviors from natural language descriptions. Existing TPAS models often learn a single joint embedding where appearance, action, and background information are entangled, causing over-reliance on identity cues, poor alignment for action-centric queries, and limited semantic connection between actions and places where they occur. To address these issues, we propose Semantic Decoupled Alignment (SeDA), a disentangled vision–language retrieval framework that explicitly factorizes both visual and textual representations into appearance, action, and background components. SeDA introduces Semantic Token Projection, which derives three semantic queries from the global [CLS] token, softly aggregates modality tokens relevant to each factor, and recomposes the resulting factor tokens into a compact retrieval embedding. To enforce factor-specific semantics, we decompose each caption into appearance/action/background sub-captions and supervise the corresponding tokens with a Feature Decoupling Loss, combined with contrastive and image–text matching objectives. On the Person Anomaly Benchmark (1M pairs), SeDA achieves 86.45% R@1 (+1.52 over SOTA), improves average multi-weather R@1 by +2.43, and gains +3.74 R@1 under out-of-distribution evaluation. Code is available in the supplementary materials.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


