Information collapse
About 10K visual tokens are often supervised by only a short diagnostic phrase.
Two failure modes motivate a structured, evidence-aware solution.
Research gaps
About 10K visual tokens are often supervised by only a short diagnostic phrase.
More than 300 conditions may share similar wording and overlapping anatomical views.
FetalMind response
Assigns each image to its anatomical view before multi-image reasoning.
Encodes view and disease priors as structured tokens.
Links diagnoses to salient views and optimizes evidence-aware preferences.
Spatial alignment and token injection. FetalMind first assigns each image to a gestation-specific anatomical view, then maps key view and disease terms to explicit tokens. This organizes a variable number of ultrasound images and reduces confusion between clinically distinct abnormalities with similar descriptions.
Salient Epistemic Disentanglement. An expert-curated disease-view graph identifies the views that carry diagnostic evidence. FetalMind swaps only matched salient views to construct hard preference pairs, while SVPO trains the model to favor evidence-consistent diagnoses and reports.
Multi-center supervision across early, mid, and late pregnancy.
FetalMind leads across report generation, diagnosis, and gestational stages.
Ablation: removing SED causes the largest overall decline, from 45.2 to 40.5 average score.
A guideline-informed template standardizes findings, measurements, impression, and diagnosis.
Please cite our paper if you find the dataset or model useful.
@misc{he2025fetalmind,
title={Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation},
author={Xiao He and Huangxuan Zhao and Guojia Wan and Jiancheng Pan and Yanxing Liu and Xin Zou and Yong Luo and Yongchao Xu and Juhua Liu and Wei Zhou and Dacheng Tao and Bo Du},
year={2025},
eprint={2510.12953},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2510.12953}
}