Logo image
Unpaired multi-modal multi-label learning for detecting endometriosis signs
Journal article   Peer reviewed

Unpaired multi-modal multi-label learning for detecting endometriosis signs

Yuan Zhang, Hu Wang, Yutong Xie, Minh-Son To, Steven Knox, Mathew Leonardi, George Condous, Jodie C. Avery, M. Louise Hull and Gustavo Carneiro
Artificial intelligence in medicine, Vol.181, p.103503
11/2026
PMID: 42594747

Abstract

Bowel nodule Endometriosis Knowledge distillation Multi-label classification Pouch of Douglas obliteration Unpaired multi-modal learning
Endometriosis is a widespread gynecological disorder causing severe pain and infertility, with diagnosis currently relying on slow, costly, and risky laparoscopy. This highlights the critical need for non-invasive imaging diagnostics using transvaginal ultrasound (TVUS) and magnetic resonance imaging (MRI). A key challenge is that patients typically receive only one scan modality in practice, despite TVUS and MRI offering differing diagnostic strengths for endometriosis signs like Pouch of Douglas (POD) obliteration and bowel nodules (BN). Previous work partially addressed this challenge by leveraging unpaired multi-modal data for detecting a single marker: Pouch of Douglas (POD) obliteration. However, this is restrictive because endometriosis signs, such as POD obliteration and bowel nodules (BN), often provide correlated diagnostic cues. Capturing these correlations is essential for accurate detection of endometriosis imaging signs, particularly when combined with multi-modal learning, as each modality offers complementary strengths for different signs. To overcome these limitations, we propose EndoFusion, a novel unpaired multi-modal, multi-label learning framework that enables the detection of POD and BN from TVUS and MRI. Our approach introduces three key innovations: (1) label-based pairing, mixup, and cross-modal feature exchange for robust single-modality inference; (2) Dynamic Mutual Knowledge Distillation (DMKD), which adaptively selects teachers using a worst-student-oriented strategy for effective cross-modal transfer; and (3) label correlations modeling with multi-head attention and a specialized loss to handle imbalance and boost accuracy. This design ensures that knowledge from the superior modality and from co-occurring signs is effectively transferred, mitigating modality-specific weaknesses and improving robustness in imaging sign detection. Experiments on our endometriosis dataset show that our method significantly outperforms all comparison methods, achieving an average AUC of 0.827 (95% CI: 0.790–0.861) when evaluated using single-modality inference. These results represent an initial proof-of-concept toward multi-modal, non-invasive assessment of selected endometriosis imaging signs from MRI and TVUS. •A multi-modal, multi-label framework detects multiple endometriosis imaging signs.•Multi-modal, multi-label learning exploits complementary MRI and TVUS information.•Dynamic mutual knowledge distillation enables robust single-modality inference.

Metrics

1 Record Views

Details

Logo image

Usage Policy