TY - GEN
T1 - No More Sibling Rivalry
T2 - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
AU - Yang, Bin
AU - Zhang, Yulin
AU - Zhou, Hong Yu
AU - Yang, Sibei
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Detection transformers have been applied to humanobject interaction (HOI) detection, enhancing the localization and recognition of human-action-object triplets in images. Despite remarkable progress, this study identifies a critical issue - 'Toxic Siblings' bias - which hinders the interaction decoder's learning, as numerous similar yet distinct HOI triplets interfere with and even compete against each other both input side and output side to the interaction decoder. This bias arises from high confusion among sibling triplets/categories, where increased similarity paradoxically reduces precision, as one's gain comes at the expense of its toxic sibling's decline. To address this, we propose two novel debiasing learning objectives - 'contrastive-thencalibration' and 'merge-then-split' - targeting the input and output perspectives, respectively. The former samples sibling-like incorrect HOI triplets and reconstructs them into correct ones, guided by strong positional priors. The latter first learns shared features among sibling categories to distinguish them from other groups, then explicitly refines intra-group differentiation to preserve uniqueness. Experiments show that we significantly outperform both the baseline (+9.18% mAP on HICO-Det) and the state-of-the-art (+3.59% mAP) across various settings.
AB - Detection transformers have been applied to humanobject interaction (HOI) detection, enhancing the localization and recognition of human-action-object triplets in images. Despite remarkable progress, this study identifies a critical issue - 'Toxic Siblings' bias - which hinders the interaction decoder's learning, as numerous similar yet distinct HOI triplets interfere with and even compete against each other both input side and output side to the interaction decoder. This bias arises from high confusion among sibling triplets/categories, where increased similarity paradoxically reduces precision, as one's gain comes at the expense of its toxic sibling's decline. To address this, we propose two novel debiasing learning objectives - 'contrastive-thencalibration' and 'merge-then-split' - targeting the input and output perspectives, respectively. The former samples sibling-like incorrect HOI triplets and reconstructs them into correct ones, guided by strong positional priors. The latter first learns shared features among sibling categories to distinguish them from other groups, then explicitly refines intra-group differentiation to preserve uniqueness. Experiments show that we significantly outperform both the baseline (+9.18% mAP on HICO-Det) and the state-of-the-art (+3.59% mAP) across various settings.
KW - class-imbalance learning
KW - hico-det
KW - hoi
KW - human-object interaction
KW - v-coco
UR - https://www.scopus.com/pages/publications/105044173931
U2 - 10.1109/ICCV51701.2025.02108
DO - 10.1109/ICCV51701.2025.02108
M3 - 会议稿件
AN - SCOPUS:105044173931
T3 - Proceedings of the IEEE International Conference on Computer Vision
SP - 22707
EP - 22717
BT - Proceedings - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 19 October 2025 through 23 October 2025
ER -