近年来,基于稀疏标注的涂鸦式弱监督学习方法,在医学影像分割中展现出了以较低标注成本而达到良好分割效果的潜力。但弱监督信号有限,模型极易出现侧重目标主体而忽视目标边缘的问题,进一步导致了较大的边界分割误差。为解决该问题,本研究提出Weak-Spatially-Agile-Mamba-UNet(WSAM-UNet)模型,利用卷积神经U型子模型进行局部特征提取、使用动态自适应注意力U型子模型进行相关像素动态感知,以及通过视觉曼巴U型子模型进行全局上下文建模。最终,经过在ACDC数据集上的实验验证,WSAM-UNet生成的分割掩模与真实标签在重叠率上达到0.893 19,在豪斯多夫距离误差和平均表面距离误差指标上分别降低至4.186 5和1.035 2,与其他多个基线模型相比,该模型在整体分割精度和边界误差控制方面达到了最优平衡。
Abstract
Recently, scribble-based weakly-supervised learning methods have shown great potential in medical image segmentation, achieving promising results at a low annotation cost. However, the limited supervisory signals often lead models to overemphasize the main body of the target while overlooking its precise boundaries, consequently resulting in significant segmentation errors along the edges. To tackle this issue, we propose a Weak-Spatially-Agile-Mamba-UNet (WSAM-UNet) model. This architecture integrates a convolutional U-Net sub-model for local feature extraction, a dynamic adaptive attention U-Net sub-model for context-aware pixel perception, and a Visual Mamba U-Net sub-model for global contextual modeling. Extensive experiments on the ACDC dataset demonstrate that our WSAM-UNet achieves an overlap rate of 0.893 19 with the ground truth, while reducing the Hausdorff Distance and Average Surface Distance to 4.186 5 and 1.035 2, respectively. Compared to several strong baselines, WSAM-UNet achieves an optimal balance between overall segmentation accuracy and boundary error control.
关键词
医学图像分割 /
弱监督学习 /
稀疏注解 /
多模态信息融合 /
涂鸦监督
Key words
medical image segmentation /
weakly supervised learning /
sparse annotation /
multimodal information fusion /
scribble supervision
{{custom_sec.title}}
{{custom_sec.title}}
{{custom_sec.content}}
参考文献
[1] Ronneberger O,Fischer P,Brox T.U-Net:Convolutional Networks for Biomedical Image Segmentation[C]//Medical Image Computing and Computer-Assisted Intervention - MICCAI 2015.Cham:Springer,2015:234-241.
[2] Isensee F,Jaeger P F,Kohl S A A,et al.NnU-Net:A Self-Configuring Method for Deep Learning-Based Biomedical Image Segmentation[J].Nature Methods,2021,18(2):203-211.
[3] Chen J,Lu Y,Yu Q,et al.TransUNet:Transformers Make Strong Encoders for Medical Image Segmentation[EB/OL].2021:arXiv:2102.04306. https://arxiv.org/abs/2102.04306.
[4] 胡玉泽. 基于一致性正则化的半监督医学图像分割方法研究[D].济南:齐鲁工业大学,2025.
[5] Dorent R,Joutard S,Shapey J,et al.Inter Extreme Points Geodesics ForEnd-to-End Weakly Supervised ImageSegmentation[C]//Medical Image Computing and Computer Assisted Intervention - MICCAI 2021.Cham:Springer,2021:615-624.
[6] 周明珠,吕笑妍.弱监督学习在医学图像分割中的应用[J].中国科技信息,2025(18):131-133.
[7] Dorent R,Joutard S,Shapey J,et al.Inter Extreme Points Geodesics ForEnd-to-End Weakly Supervised ImageSegmentation[C]//Medical Image Computing and Computer Assisted Intervention - MICCAI 2021.Cham:Springer,2021:615-624.
[8] Lin D,Dai J,Jia J,et al.ScribbleSup:Scribble-Supervised Convolutional Networks for Semantic Segmentation[C]//2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR).IEEE,2016:3159-3167.
[9] Isensee F,Jaeger P F,Kohl S A A,et al.NnU-Net:A Self-Configuring Method for Deep Learning-Based Biomedical Image Segmentation[J].Nature Methods,2021,18(2):203-211.
[10] Ronneberger O,Fischer P,Brox T.U-Net:Convolutional Networks for Biomedical Image Segmentation[C]//Medical Image Computing and Computer-Assisted Intervention - MICCAI 2015.Cham:Springer,2015:234-241.
[11] Wang H,Cao P,Wang J,et al.UCTransNet:Rethinking the Skip Connections in U-Net from a Channel-Wise Perspective with Transformer[J]. Proceedings of the AAAI Conference on Artificial Intelligence,2022,36(3):2441-2449.
[12] Oktay O,Schlemper J,Le Folgoc L,et al.Attention U-Net:Learning where to Look for the Pancreas[EB/OL].2018:arXiv:1804.03999. https://arxiv.org/abs/1804.03999.
[13] Dosovitskiy A,Beyer L,Kolesnikov A,et al.An Image Is Worth 16x16 Words:Transformers for Image Recognition at Scale[EB/OL].2020:arXiv:2010.11929.https://arxiv.org/abs/2010.11929.
[14] Liu Z,Lin Y,Cao Y,et al.Swin Transformer:Hierarchical Vision Transformer Using Shifted Windows[C]//2021 IEEE/CVF International Conference on Computer Vision (ICCV).IEEE,2022:9992-10002.
[15] Zhou H Y,Guo J,Zhang Y,et al.NnFormer:Volumetric Medical Image Segmentation via a 3D Transformer[J].IEEE Transactions on Image Processing,2023,32:4036-4045.
[16] Li Z,Zheng Y,Shan D,et al.ScribFormer:Transformer Makes CNN Work Better for Scribble-Based Medical Image Segmentation[J].IEEE Transactions on Medical Imaging,2024,43(6):2254-2265.
[17] Hatamizadeh A,Tang Y,Nath V,et al.UNETR:Transformers for 3D Medical Image Segmentation[C]//2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV).IEEE,2022:1748-1758.
[18] Gu A.Modeling Sequences with Structured State Spaces[M].Stanford University,2023.
[19] Zhu L,Liao B,Zhang Q,et al.Vision Mamba:Efficient Visual Representation Learning with Bidirectional State Space Model[EB/OL].2024:arXiv:2401.09417.https://arxiv.org/abs/2401.09417.
[20] Wang Z,Zheng J Q,Zhang Y,et al.Mamba-unet:Unet-like Pure Visual Mamba for Medical Image Segmentation[J].arXiv preprint arXiv:2402.05079,2024.
[21] Ma C,Wang Z.Semi-Mamba-UNet:Pixel-Level Contrastive and Cross-Supervised Visual Mamba-Based UNet for Semi-Supervised Medical Image Segmentation[J].Knowledge-Based Systems,2024,300:112203.
[22] Tang M,Perazzi F,Djelouah A,et al.On Regularized Losses for Weakly-Supervised CNN Segmentation[C]//Computer Vision-ECCV 2018.Cham:Springer,2018:524-540.
[23] Qiu P,Yang J,Kumar S,et al.AgileFormer:Spatially Agile Transformer UNet for Medical Image Segmentation[EB/OL].2024:arXiv:2404.00122.https://arxiv.org/abs/2404.00122.
[24] Dai J,Qi H,Xiong Y,et al.Deformable Convolutional Networks[C]//2017 IEEE International Conference on Computer Vision (ICCV).IEEE,2017:764-773.
[25] Hassani A,Walton S,Li J,et al.Neighborhood Attention Transformer[C]//2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).IEEE,2023:6185-6194.
[26] Zhu X,Su W,Lu L,et al.Deformable DETR:Deformable Transformers for End-to-End Object Detection[EB/OL].2020:arXiv:2010.04159. https://arxiv.org/abs/2010.04159.
[27] Wang Z,Ma C.Weak-Mamba-UNet:Visual Mamba Makes CNN and ViT Work Better for Scribble-Based Medical Image Segmentation[EB/OL].2024:arXiv:2402.10887.https://arxiv.org/abs/2402.10887.
[28] Lee H,Jeong W K.Scribble2Label:Scribble-Supervised Cell Segmentation via Self-Generating Pseudo-Labels with Consistency[M]//Medical Image Computing and Computer Assisted Intervention - MICCAI 2020.Cham:Springer International Publishing,2020:14-23.
[29] Luo X,Hu M,Liao W,et al.Scribble-Supervised Medical Image Segmentation ViaDual-Branch Network andDynamically Mixed Pseudo Labels Supervision[C]//Medical Image Computing and Computer Assisted Intervention - MICCAI 2022.Cham:Springer,2022:528-538.
[30] Verma V,Kawaguchi K,Lamb A,et al.Interpolation Consistency Training for Semi-Supervised Learning[J].Neural Networks,2022,145:90-106.
[31] Ouali Y,Hudelot C,Tami M.Semi-Supervised Semantic Segmentation with Cross-Consistency Training[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).IEEE,2020:12671-12681.
[32] Bernard O,Lalande A,Zotti C,et al.Deep Learning Techniques for Automatic MRI Cardiac Multi-Structures Segmentation and Diagnosis:Is the Problem Solved?[J].IEEE Transactions on Medical Imaging,2018,37(11):2514-2525.
[33] Grandvalet Y,Bengio Y.Semi-supervised Learning by Entropy Minimization[J].Advances in Neural Information Processing Systems,2004,17.
[34] Obukhov A,Georgoulis S,Dai D,et al.Gated CRF Loss for Weakly Supervised Semantic Image Segmentation[EB/OL].2019:arXiv:1906.04651.https://arxiv.org/abs/1906.04651.
[35] Kim B,Ye J C.Mumford-Shah Loss Functional for Image Segmentation with Deep Learning[J].2020,29:1856-1866.
[36] 张家波,李倩,代玥.融合注意力调制的弱监督语义分割方法[J/OL].计算机工程与应用, 1-12[2025-10-23].https://doi.org/10.3778/j.issn.1002-8331.2505-0382.