Multimodal Semantic-Guided Scribble-Supervised Medical Image Segmentation

LI Tu-chao, QU Mei-jie, CHEN Zhao-yang

Computer & Telecommunication ›› 2025 ›› Issue (12) : 26-34.

Computer & Telecommunication ›› 2025 ›› Issue (12) : 26-34.

Multimodal Semantic-Guided Scribble-Supervised Medical Image Segmentation

  • LI Tu-chao, QU Mei-jie, CHEN Zhao-yang
Author information +
History +

Abstract

Recently, scribble-based weakly-supervised learning methods have shown great potential in medical image segmentation, achieving promising results at a low annotation cost. However, the limited supervisory signals often lead models to overemphasize the main body of the target while overlooking its precise boundaries, consequently resulting in significant segmentation errors along the edges. To tackle this issue, we propose a Weak-Spatially-Agile-Mamba-UNet (WSAM-UNet) model. This architecture integrates a convolutional U-Net sub-model for local feature extraction, a dynamic adaptive attention U-Net sub-model for context-aware pixel perception, and a Visual Mamba U-Net sub-model for global contextual modeling. Extensive experiments on the ACDC dataset demonstrate that our WSAM-UNet achieves an overlap rate of 0.893 19 with the ground truth, while reducing the Hausdorff Distance and Average Surface Distance to 4.186 5 and 1.035 2, respectively. Compared to several strong baselines, WSAM-UNet achieves an optimal balance between overall segmentation accuracy and boundary error control.

Key words

medical image segmentation / weakly supervised learning / sparse annotation / multimodal information fusion / scribble supervision

Cite this article

Download Citations
LI Tu-chao, QU Mei-jie, CHEN Zhao-yang. Multimodal Semantic-Guided Scribble-Supervised Medical Image Segmentation[J]. Computer & Telecommunication. 2025(12): 26-34

References

[1] Ronneberger O,Fischer P,Brox T.U-Net:Convolutional Networks for Biomedical Image Segmentation[C]//Medical Image Computing and Computer-Assisted Intervention - MICCAI 2015.Cham:Springer,2015:234-241.
[2] Isensee F,Jaeger P F,Kohl S A A,et al.NnU-Net:A Self-Configuring Method for Deep Learning-Based Biomedical Image Segmentation[J].Nature Methods,2021,18(2):203-211.
[3] Chen J,Lu Y,Yu Q,et al.TransUNet:Transformers Make Strong Encoders for Medical Image Segmentation[EB/OL].2021:arXiv:2102.04306. https://arxiv.org/abs/2102.04306.
[4] 胡玉泽. 基于一致性正则化的半监督医学图像分割方法研究[D].济南:齐鲁工业大学,2025.
[5] Dorent R,Joutard S,Shapey J,et al.Inter Extreme Points Geodesics ForEnd-to-End Weakly Supervised ImageSegmentation[C]//Medical Image Computing and Computer Assisted Intervention - MICCAI 2021.Cham:Springer,2021:615-624.
[6] 周明珠,吕笑妍.弱监督学习在医学图像分割中的应用[J].中国科技信息,2025(18):131-133.
[7] Dorent R,Joutard S,Shapey J,et al.Inter Extreme Points Geodesics ForEnd-to-End Weakly Supervised ImageSegmentation[C]//Medical Image Computing and Computer Assisted Intervention - MICCAI 2021.Cham:Springer,2021:615-624.
[8] Lin D,Dai J,Jia J,et al.ScribbleSup:Scribble-Supervised Convolutional Networks for Semantic Segmentation[C]//2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR).IEEE,2016:3159-3167.
[9] Isensee F,Jaeger P F,Kohl S A A,et al.NnU-Net:A Self-Configuring Method for Deep Learning-Based Biomedical Image Segmentation[J].Nature Methods,2021,18(2):203-211.
[10] Ronneberger O,Fischer P,Brox T.U-Net:Convolutional Networks for Biomedical Image Segmentation[C]//Medical Image Computing and Computer-Assisted Intervention - MICCAI 2015.Cham:Springer,2015:234-241.
[11] Wang H,Cao P,Wang J,et al.UCTransNet:Rethinking the Skip Connections in U-Net from a Channel-Wise Perspective with Transformer[J]. Proceedings of the AAAI Conference on Artificial Intelligence,2022,36(3):2441-2449.
[12] Oktay O,Schlemper J,Le Folgoc L,et al.Attention U-Net:Learning where to Look for the Pancreas[EB/OL].2018:arXiv:1804.03999. https://arxiv.org/abs/1804.03999.
[13] Dosovitskiy A,Beyer L,Kolesnikov A,et al.An Image Is Worth 16x16 Words:Transformers for Image Recognition at Scale[EB/OL].2020:arXiv:2010.11929.https://arxiv.org/abs/2010.11929.
[14] Liu Z,Lin Y,Cao Y,et al.Swin Transformer:Hierarchical Vision Transformer Using Shifted Windows[C]//2021 IEEE/CVF International Conference on Computer Vision (ICCV).IEEE,2022:9992-10002.
[15] Zhou H Y,Guo J,Zhang Y,et al.NnFormer:Volumetric Medical Image Segmentation via a 3D Transformer[J].IEEE Transactions on Image Processing,2023,32:4036-4045.
[16] Li Z,Zheng Y,Shan D,et al.ScribFormer:Transformer Makes CNN Work Better for Scribble-Based Medical Image Segmentation[J].IEEE Transactions on Medical Imaging,2024,43(6):2254-2265.
[17] Hatamizadeh A,Tang Y,Nath V,et al.UNETR:Transformers for 3D Medical Image Segmentation[C]//2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV).IEEE,2022:1748-1758.
[18] Gu A.Modeling Sequences with Structured State Spaces[M].Stanford University,2023.
[19] Zhu L,Liao B,Zhang Q,et al.Vision Mamba:Efficient Visual Representation Learning with Bidirectional State Space Model[EB/OL].2024:arXiv:2401.09417.https://arxiv.org/abs/2401.09417.
[20] Wang Z,Zheng J Q,Zhang Y,et al.Mamba-unet:Unet-like Pure Visual Mamba for Medical Image Segmentation[J].arXiv preprint arXiv:2402.05079,2024.
[21] Ma C,Wang Z.Semi-Mamba-UNet:Pixel-Level Contrastive and Cross-Supervised Visual Mamba-Based UNet for Semi-Supervised Medical Image Segmentation[J].Knowledge-Based Systems,2024,300:112203.
[22] Tang M,Perazzi F,Djelouah A,et al.On Regularized Losses for Weakly-Supervised CNN Segmentation[C]//Computer Vision-ECCV 2018.Cham:Springer,2018:524-540.
[23] Qiu P,Yang J,Kumar S,et al.AgileFormer:Spatially Agile Transformer UNet for Medical Image Segmentation[EB/OL].2024:arXiv:2404.00122.https://arxiv.org/abs/2404.00122.
[24] Dai J,Qi H,Xiong Y,et al.Deformable Convolutional Networks[C]//2017 IEEE International Conference on Computer Vision (ICCV).IEEE,2017:764-773.
[25] Hassani A,Walton S,Li J,et al.Neighborhood Attention Transformer[C]//2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).IEEE,2023:6185-6194.
[26] Zhu X,Su W,Lu L,et al.Deformable DETR:Deformable Transformers for End-to-End Object Detection[EB/OL].2020:arXiv:2010.04159. https://arxiv.org/abs/2010.04159.
[27] Wang Z,Ma C.Weak-Mamba-UNet:Visual Mamba Makes CNN and ViT Work Better for Scribble-Based Medical Image Segmentation[EB/OL].2024:arXiv:2402.10887.https://arxiv.org/abs/2402.10887.
[28] Lee H,Jeong W K.Scribble2Label:Scribble-Supervised Cell Segmentation via Self-Generating Pseudo-Labels with Consistency[M]//Medical Image Computing and Computer Assisted Intervention - MICCAI 2020.Cham:Springer International Publishing,2020:14-23.
[29] Luo X,Hu M,Liao W,et al.Scribble-Supervised Medical Image Segmentation ViaDual-Branch Network andDynamically Mixed Pseudo Labels Supervision[C]//Medical Image Computing and Computer Assisted Intervention - MICCAI 2022.Cham:Springer,2022:528-538.
[30] Verma V,Kawaguchi K,Lamb A,et al.Interpolation Consistency Training for Semi-Supervised Learning[J].Neural Networks,2022,145:90-106.
[31] Ouali Y,Hudelot C,Tami M.Semi-Supervised Semantic Segmentation with Cross-Consistency Training[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).IEEE,2020:12671-12681.
[32] Bernard O,Lalande A,Zotti C,et al.Deep Learning Techniques for Automatic MRI Cardiac Multi-Structures Segmentation and Diagnosis:Is the Problem Solved?[J].IEEE Transactions on Medical Imaging,2018,37(11):2514-2525.
[33] Grandvalet Y,Bengio Y.Semi-supervised Learning by Entropy Minimization[J].Advances in Neural Information Processing Systems,2004,17.
[34] Obukhov A,Georgoulis S,Dai D,et al.Gated CRF Loss for Weakly Supervised Semantic Image Segmentation[EB/OL].2019:arXiv:1906.04651.https://arxiv.org/abs/1906.04651.
[35] Kim B,Ye J C.Mumford-Shah Loss Functional for Image Segmentation with Deep Learning[J].2020,29:1856-1866.
[36] 张家波,李倩,代玥.融合注意力调制的弱监督语义分割方法[J/OL].计算机工程与应用, 1-12[2025-10-23].https://doi.org/10.3778/j.issn.1002-8331.2505-0382.

Accesses

Citation

Detail

Sections
Recommended

/

〈 〉