Precise recognition of medical images plays a critical role in the diagnosis and treatment of major diseases. In data-scarce scenarios, few-shot learning technology offers an effective solution for the diagnosis of rare diseases of medical imaging. Current mainstream methods predominantly adopt a metric-learning paradigm, which first extracts image features via a backbone network to construct prototype representations, followed by classifying test samples based on spatial metrics. Existing research typically employs Convolutional Neural Networks (CNNs) or Transformer architectures as feature extraction backbones. However, convolution operation is limited by local receptive field, resulting in insufficient global feature capture of the image, while Transformers has the problem of excessively high secondary computational complexity. To address the aforementioned issues, a novel backbone network architecture based on a Frequency-Domain Collaborative State Space Model (DCT-Mamba) is proposed. This model leverages the Discrete Cosine Transform (DCT) to enhance frequency-domain features and combines the State Space Model (SSM) to achieve long-range modeling with linear complexity. Systematic experimental validation on medical imaging datasets demonstrates that the proposed model outperforms both CNN and Vision Transformer structures in meta-task scenarios under few-shot learning settings. Experiment results indicate that, compared with traditional architectures, DCT-Mamba significantly improves recognition accuracy while maintaining linear computational complexity.
Mamba的SS2D架构基于选择性状态空间模型(Selective State Space Model, SS2D),通过动态调整跨空间维度的信息交互路径,实现图像长程依赖的建模[11]。该架构的核心创新体现在两个关键机制:选择性扫描机制和空间自适应状态转移模块。选择性扫描机制针对医学图像中病灶与正常组织的空间分布差异,采用门控网络动态激活不同方向的扫描路径,使模型能够捕捉病灶边缘区域的纹理细节特征[12];空间自适应状态转移模块通过可学习的状态转移矩阵,根据图像内容自动调节特征传播强度,有效克服了基于注意力方法中注意力权重导致的图像细节模糊问题。相较于传统CNN的局部感受野限制和Transformer的全局计算复杂度过高的问题,SS2D架构在医学图像处理中具有更高的效率:通过动态特征选择机制,模型能在保证长程依赖建模效率的同时,实现对病灶区域的自适应特征提取[13],这对于提升下游任务的精度具有重要价值[14]。
RONNEBERGERO, FISCHERP, BROXT. U-Net:Convolutional networks for biomedical image segmentation[C]//International Conference on Medical Image Computing and Computer-Assisted Intervention. Cham:Springer,2015:234-241.
[2]
FRID-ADARM, KLANGE, AMITAIM,et al. GAN-based synthetic medical image augmentation for improved CNN performance in liver lesion classification[J]. IEEE Transactions on Medical Imaging,2018,37(12):2280-2290.
[3]
RAGHUM, ZHANGC, KLEINBERGJ,et al. Transfusion:Understanding transfer learning for medical imaging[C]//Advances in Neural Information Processing Systems. New York:Curran Associates,2019:3347-3357.
[4]
ZHOUJ, CUIG, HUS,et al.Graph neural networks:A review of methods and applications[J]. AI Open,2021,2:57-81.
[5]
FINNC, ABBEELP, LEVINES. Model-agnostic meta-learning for fast adaptation of deep networks[C]//International Conference on Machine Learning. Sydney:PMLR,2017:1126-1135.
[6]
SNELLJ, SWERSKYK, ZEMELR. Prototypical networks for few-shot learning[C]//Advances in Neural Information Processing Systems. Long Beach:Curran Associates,2017:4077-4087.
[7]
AHMEDN, NATARAJANT, RAOK R. Discrete cosine transform[J]. IEEE Transactions on Acoustics,1974,23(1):90-93.
[8]
ZHANGW, SALMIA, YANGC,et al. Innovative Noise Extraction and Denoising in Low-Dose CT Using a Supervised Deep Learning Framework[J]. Electronics,2020,65:101768.
[9]
LIY, ZHAOZ, YUANJ,et al. MFENet:Multi-Scale and Local Frequency Enhancement Network for Skin Lesion Classification[C]//Proceedings of the Computer Graphics International Conference (CGI 2024). Cham:Springer Nature Switzerland,2024:192-203.
[10]
CHENX, LIM, ZHOUT,et al.Adaptive DCT in pulmonary nodule detection[C]//Medical Image Computing and Computer Assisted Intervention. Cham:Springer,2023:100-110.
[11]
GUA,DAO T. Mamba:Linear-Time Sequence Modeling with Selective State Spaces[C]//Proceedings of the First Conference on Language Modeling. New York:Curran Associates,2024:1-10.
[12]
GONGH, KANGL, WANGY, alel. NNMamba:3D Biomedical Image Segmentation,Classification and Landmark Detection with State Space Model[C]//Proceedings of the 2025 IEEE 22nd International Symposium on Biomedical Imaging (ISBI 2025). Los Alamitos:IEEE Computer Society,2025:1-5.
[13]
GALEA G, WORTHINGTONB S. The Utility of Scanning Strategies in Radiology[M]//Eye Movements and Psychological Functions. Abingdon:Routledge,2021:169-191.
[14]
ZHANGY, LIUQ, WUT,et al. Efficient long-range modeling for medical images[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle:IEEE Computer Society,2024:2345-2355.
BOUDIAFM, MASNICHIL, TORRMANCHEA,et al.Transductive information maximization for few-shot learning[C]//Advances in Neural Information Processing Systems 33 (NeurIPS 2020). New York:Curran Associates,2020:1234-1245.
[17]
DHILLONG, CHAUDHARIP, RAVICHANDRANA,et al. Transductive fine-tuning for few-shot learning[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle:IEEE,2020:1126-1135.
[18]
LIUY, ZHANGH, WANGL,et al. Bayesian deep learning for few-shot medical image classification[J]. IEEE Transactions on Medical Imaging,2021,40(3):123-135.
[19]
YEH, HUG, LIUJ,et al. FEAT:Few-shot embedding adaptation with transformer[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle:IEEE,2020:2345-2355.
[20]
POGORELOVK, RANHEIM RANDELK, GRIWODZC,et al. Kvasir:A multi-class image dataset for computer aided gastrointestinal disease detection[C]//Proceedings of the 8th ACM Multimedia Systems Conference (MMSys 2017). Taipei:ACM,2017:164-169.
[21]
SHASTRIS, KANSALI, KUMARS,et al. CheXImageNet:a novel architecture for accurate classification of Covid-19 with chest x-ray digital images using deep convolutional neural networks[J]. Health and Technology,2022,12(2):193-204.