School of Computer Science and Technology,Shanxi Provincial Key Laboratory of Biomedical Imaging and Imaging Big Data,North University of China,Taiyuan 030051,China
Existing deep learning-based multimodal medical image fusion methods suffer from insufficient high-level feature extraction and easy loss of low-level features. To tackle these problems, this paper proposed a multimodal medical image fusion method based on dilated convolution and graph attention aggregation. The method was comprised of three components: a dual-branch encoder, a fusion module, and a decoder. The dual-branch encoder consisted of a convolution-based low-level encoder and a graph-convolution-based high-level encoder. The convolution-based low-level encoder employed dilated convolution to mitigate the loss of low-level features like texture details and provided initialized node features for the high-level encoder. The graph-convolution-based high-level encoder mainly adopted the graph attention aggregation module to effectively capture high-level features such as deep semantics. The graph attention aggregation module constructed a node adjacency matrix by integrating multi-head attention with edge encoding and then performed deep aggregation of nodes through graph convolution based on this adjacency matrix. The fusion module fused the extracted features, and the decoder reconstructed the fused image. The method was compared with six state-of-the-art image fusion methods on subjective vision and objective evaluation metrics. The results show that this method improves 2.4% on EN compared to the IGNet method, 3.53% and 5.06% on AG and MI compared to the DATFuse method, and 1.18%, 6.24%, and 3% on SD, SF, and SCD compared to the SwinFusion method, respectively, while the fused image obtained by this method retains more texture detail information. The comprehensive experimental results demonstrate that this method achieves effective fusion of multimodal medical images, offering more reliable image support for clinical diagnosis.
DIWAKARM, SINGHP, RAVIV, et al. A non-conventional review on multi-modality-based medical image fusion[J]. Diagnostics, 2023, 13(5): 820.
[2]
HUANGB, YANGF, YINM, et al. A review of multimodal medical image fusion techniques [J]. Computational and Mathematical Methods in Medicine, 2020, 2020: 8279342.
[3]
DUJ, LIW. Two - scale image decomposition based image fusion using structure tensor [J]. International Journal of Imaging Systems and Technology, 2020, 30(2): 271-284.
[4]
XIAJ M, CHENY M, CHENA Y, et al. Medical image fusion based on sparse representation and PCNN in NSCT domain [J]. Computational and Mathematical Methods in Medicine, 2018(1): 2806047.
[5]
WEIX, QIUY, XUX, et al. ECINFusion: A novel explicit channel-wise interaction network for unified multi-modal medical image fusion[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2025, 35(5): 4011-4025.
[6]
CHENJ, DINGJ, YUY, et al. THFuse: An infrared and visible image fusion network using transformer and hybrid feature extractor[J]. Neurocomputing, 2023, 527: 71-82.
[7]
XUH, MAJ, JIANGJ, et al. U2Fusion: A unified unsupervised image fusion network [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 44(1): 502-518.
[8]
LIH, XUT, WUX J, et al. LRRNet: A novel representation learning guided fusion network for infrared and visible images [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(9): 11040-11052.
[9]
SHAOD, YANGH, MAL, et al. AFPNet: An adaptive frequency-domain optimized progressive medical image fusion network[J]. Biomedical Signal Processing and Control, 2025, 103: 107357.
[10]
LIUZ, LINY, CAOY, et al. Swin transformer: Hierarchical vision transformer using shifted windows [DB/OL]. (2021-08-17) [2025-01-06].
[11]
CHENJ, DINGJ, MAJ. HitFusion: Infrared and visible image fusion for high-level vision tasks using transformer[J]. IEEE Transactions on Multimedia, 2024, 26: 10145-10159.
[12]
LIX, HEH, SHIJ. HDCCT: Hybrid densely connected CNN and transformer for infrared and visible image fusion[J]. Electronics, 2024, 13(17): 3470.
[13]
LINZ, SUNW, TANGB, et al. Semantic segmentation network with multi-path structure, attention reweighting and multi-scale encoding[J]. The Visual Computer, 2023, 39(2): 597-608.
[14]
XUJ, BIANQ, LIX, et al. Contrastive graph pooling for explainable classification of brain networks[J]. IEEE Transactions on Medical Imaging, 2024, 43(9): 3292-3305.
[15]
LIJ, CHENJ, LIUJ, et al. Learning a graph neural network with cross modality interaction for image fusion [DB/OL]. (2023-08-07) [2025-01-06].
[16]
LIJ, BAIL, YANGB, et al. Graph representation learning for infrared and visible image fusion[DB/OL]. (2023-11-01) [2025-01-06].
[17]
MAJ, LIX, ZHANGY, et al. U-Convnext network for infrared small target detection[C]//IEEE International Conference on Image Processing(ICIP), 2024: 1371-1376.
[18]
ZHOUM, XUX, ZHANGY. An attention-based multi-scale feature learning network for multimodal medical image fusion [DB/OL]. (2022-12-09) [2025-01-06].
MAJiquan, ZHAOShumin, KONGFanhui. Semantic image segmentation by using multi-scale strip pooling and channel attention[J]. Journal of Image and Graphics, 2022, 27(12): 3530-3541. (in Chinese)
[23]
LIANGF, QIANC, YUW, et al. Survey of graph neural networks and applications[J]. Wireless Communications and Mobile Computing, 2022(1): 9261537.
XUZhihong, ZHANGTianrun, WANGLiqin, et al. Temporal knowledge graph reasoning with graph reconstruction [J]. Computer Engineering and Applications, 2024, 60(9): 181-187. (in Chinese)
[26]
PANCINON, GALLEGATIC, ROMAGNOLIF, et al. Protein–protein interfaces: A graph neural network approach [J]. International Journal of Molecular Sciences, 2024, 25(11): 5870.
[27]
ACHANTAR, SHAJIA, SMITHK, et al. SLIC superpixels compared to state-of-the-art superpixel methods[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2012, 34(11): 2274-2282.
[28]
ZHOUH, LUOF, ZHUANGH, et al. Attention multihop graph and multiscale convolutional fusion network for hyperspectral image classification[J]. IEEE Transactions on Geoscience and Remote Sensing, 2023, 61: 5508614.
[29]
SHIC, WUH, WANGL. CEGAT: A CNN and enhanced-GAT based on key sample selection strategy for hyperspectral image classification[J]. Neural Networks, 2023, 168: 105-122.
[30]
PENGF, LUW, TANW, et al. Multi-output network combining GNN and CNN for remote sensing scene classification[J]. Remote Sensing, 2022, 14(6): 1478.
[31]
YINGC, CAIT, LUOS, et al. Do transformers really perform bad for graph representation? [DB/OL]. (2021-06-09) [2025-01-06].
[32]
HANK, WANGY, GUOJ, et al. Vision GNN: An image is worth graph of nodes [DB/OL]. (2022-11-04) [2025-01-06].
[33]
CHENJ, LIUW, HUANGZ, et al. Universal deep GNNs: Rethinking residual connection in GNNs from a path decomposition perspective for preventing the over-smoothing[DB/OL]. (2022-05-30) [2025-01-06].
[34]
ZHAOY, ZHENGQ, ZHUP, et al. TUFusion: A transformer-based universal fusion algorithm for multimodal images[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34(3): 1712-1725.
[35]
MAJ, TANGL, FANF, et al. SwinFusion: Cross-domain long-range learning for general image fusion via swin transformer[J]. IEEE/CAA Journal of Automatica Sinica, 2022, 9(7): 1200-1217.
[36]
LIJ, CHENJ, LIUJ, et al. Learning a graph neural network with cross modality interaction for image fusion[DB/OL].(2023-08-07) [2025-01-06].
[37]
TANGW, HEF, LIUY. ITFuse: An interactive transformer for infrared and visible image fusion[J]. Pattern Recognition, 2024, 156: 110822.
[38]
TANGW, HEF, LIUY, et al. DATFuse: Infrared and visible image fusion via dual attention transformer[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2023, 33(7): 3159-3172.
[39]
AZAMM A, KHANK B, SALAHUDDINS, et al. A review on multimodal medical image fusion: Compendious analysis of medical modalities, multimodal databases, fusion techniques and quality metrics[J]. Computers in Biology and Medicine, 2022, 144: 105253.