South-Central Minzu University,a. College of Computer Science(College of Artificial Intelligence); b. College of Electronic and Information Engineering,Wuhan 430074,Hubei China
With the continuous advancement of computer vision technology, image fusion, as an effective way to obtain and optimize information, has received extensive attention. In recent years, diffusion models have become a key research direction in the field of image fusion due to their outstanding advantages in generation and reconstruction. This paper provides a systematic review of the latest research progress in using diffusion models for image fusion. Firstly, the basic principles of diffusion models are elaborated in detail, including the denoising diffusion probabilistic model, the noise-conditioned score network, and the diffusion model based on stochastic differential equations. Key elements such as the forward diffusion process, the reverse process, and the training objectives are introduced, and the application potential of this model in image fusion is analyzed. Secondly, based on the model architecture, training strategies, and practical application scenarios, the existing image fusion methods based on diffusion models are analyzed in depth, with a focus on comparing the characteristics and innovative aspects of various methods. At the same time, relevant fusion datasets and evaluation indicators are discussed, and qualitative and quantitative evaluation comparison experiments of multiple algorithms are carried out, clearly demonstrating the advantages and disadvantages of different algorithms. Finally, the future development directions are prospected. It is proposed that efforts should be made to improve the fusion performance, optimize the model training and application strategies, and strengthen cross-disciplinary cooperation to promote the further application and development of diffusion models in image fusion technology.
MitianoudisN, StathakiT. Pixel-based and region-based image fusion schemes using ICA bases[J]. Information Fusion, 2007, 8(2): 131-142.
[3]
AslantasV, ToprakA N. A pixel based multi-focus image fusion method[J]. Optics Communications, 2014, 332: 350-358.
[4]
SreejaP, HariharanS. An improved feature based image fusion technique for enhancement of liver lesions[J]. Biocybernetics and Biomedical Engineering, 2018, 38(3): 611-623.
[5]
HossnyM, NahavandiS, CrieghtonD. Feature-based image fusion quality metrics[C]// Intelligent Robotics and Applications. Berlin: Springer Berlin Heidelberg, 2008: 469-478.
[6]
KausarN, MajidA. Random forest-based scheme using feature and decision levels information for multi-focus image fusion[J]. Pattern Analysis and Applications, 2016, 19(1): 221-236.
[7]
ZhangY, LiuY, SunP, et al. IFCNN: A general image fusion framework based on convolutional neural network[J]. Information Fusion, 2020, 54: 99-118.
[8]
LiuY, ChenX, PengH, et al. Multi-focus image fusion with a deep convolutional neural network[J]. Information Fusion, 2017, 36: 191-207.
[9]
RrenX, MengF, HuT, et al. Infrared-visible image fusion based on convolutional neural networks (CNN)[C]// Intelligence Science and Big Data Engineering. Cham: Springer International Publishing, 2018: 301-307.
[10]
WangL J, HanJ, ZhangY, et al. Image fusion via feature residual and statistical matching[J]. IET Computer Vision, 2016, 10(6): 551-558.
[11]
WuY, HuangM, LiY, et al. A distributed fusion framework of multispectral and panchromatic images based on residual network[J]. Remote Sensing, 2021, 13(13): 2556.
[12]
XuH, MaJ, ZhangX P. MEF-GAN: Multi-exposure image fusion via generative adversarial networks[J]. IEEE Transactions on Image Processing, 2020, 29: 7203-7216.
[13]
ZhangH, LeZ, ShaoZ, et al. MFF-GAN: An unsupervised generative adversarial network with adaptive and gradient joint constraints for multi-focus image fusion[J]. Information Fusion, 2021, 66: 40-53.
[14]
ZhangJ, JiaoL, MaW, et al. Transformer based conditional GAN for multimodal image fusion[J]. IEEE Transactions on Multimedia, 2023, 25: 8988-9001.
[15]
DhariwalP, NicholA. Diffusion models beat gans on image synthesis[J]. Advances in Neural Information Processing Systems, 2021, 34: 8780-8794.
AustinJ, JohnsonD D, HO J, et al. Structured denoising diffusion models in discrete state-spaces[J]. Advances in Neural Information Processing Systems, 2021, 34: 17981-17993.
[19]
HoJ, JainA, AbbeelP. Denoising diffusion probabilistic models[J]. Advances in Neural Information Processing Systems, 2020, 33: 6840-6851.
[20]
CroitoruF A, HondruV, IonescuR T, et al. Diffusion models in vision: A survey[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(9): 10850-10869.
[21]
SongY, Sohl-DicksteinJ, KingmaD P, et al. Score-based generative modeling through stochastic differential equations[J]. arXiv:2020: 2011.13456.
[22]
LiM, PeiR, ZhengT, et al. FusionDiff: Multi-focus image fusion using denoising diffusion probabilistic models[J]. Expert Systems with Applications, 2024, 238: 121664.
[23]
LiX, ZongS, DuanZ, et al. A new generative method for multi-focus image fusion of underwater micro bubbles[J]. Scientific Reports, 2024, 14: 30280.
[24]
SinghP, DiwakarM. Wavelet-based multi-focus image fusion using average method noise diffusion[J]. Recent Advances in Computer Science and Communications, 2021, 14(8): 2436-2448.
[25]
DorjsembeZ, OdonchimedS, XiaoF. Three-dimensional medical image synthesis with denoising diffusion probabilistic models[C]//Medical Imaging with Deep Learning. Zurich: MIDL, 2022:1-3.
[26]
ZhaoZ, BaiH, ZhuY, et al. DDFM: denoising diffusion model for multi-modality image fusion[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Paris: IEEE/CVF, 2023: 8082-8093.
[27]
YangB, JiangZ, PanD, et al. Multimodal image fusion based on diffusion model[C]//Proceedings of the 2024 International Conference on Advanced Robotics, Automation Engineering and Machine Learning. Hangzhou:ACM, 2024: 75-81.
[28]
LuoZ, ChenD, ZhangY, et al. Notice of removal: VideoFusion: Decomposed diffusion models for high-quality video generation[C]//2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver:IEEE, 2023: 10209-10218.
[29]
LeD T, ShiH, CaiJ, et al. Diffusion model for robust multi-sensor fusion in 3D object detection and BEV segmentation[C]//Computer Vision-ECCV 2024. Cham: Springer,2024: 232-249.
[30]
SunJ, CaoN, BiH, et al. DiffRecon: Diffusion-based CT reconstruction with cross-modal deformable fusion for DR-guided non-coplanar radiotherapy[J]. Computers in Biology and Medicine, 2024, 179: 108868.
ZhangC, HanJ, ZhuJ, et al. Variational diffusion method for remote sensing image fusion[J]. IEEE Geoscience and Remote Sensing Letters, 2024, 21: 4008405.
[33]
PopS, LavialleO, TerebesR, et al. A PDE-based approach for image fusion[C]//Advanced Concepts for Intelligent Vision Systems. Berlin: Springer Berlin Heidelberg, 2007: 121-131.
[34]
XuY, LiX, JieY, et al. Simultaneous tri-modal medical image fusion andSuper-resolution using conditional diffusion model[C]//Medical Image Computing and Computer Assisted Intervention-MICCAI 2024. Cham: Springer, 2024: 635-645.
[35]
Muller-FranzesG, NiehuesJ M, KhaderF, et al. A multimodal comparison of latent denoising diffusion probabilistic models and generative adversarial networks for medical image synthesis[J]. Scientific Reports, 2023, 13(1): 12098.
[36]
MengQ, ShiW, LiS, et al. PanDiff: A novel pansharpening method based on denoising diffusion probabilistic model[J]. IEEE Transactions on Geoscience and Remote Sensing, 2023, 61: 5611317.
[37]
YiX, TangL, ZhangH, et al. Diff-IF: Multi-modality image fusion via diffusion model with fusion knowledge prior[J]. Information Fusion, 2024, 110: 102450.
[38]
ZhangH, CaoL, MaJ. Text-DiFuse: An interactive multi-modal image fusion framework based on text-modulated diffusion model[C]//The Thirty-eighth Annual Conference on Neural Information Processing Systems. Vancouver: NeurIPS Foundation, 2024: 1-24.
[39]
CaoB, XuX, ZhuP, et al. Conditional controllable image fusion[C]//The Thirty-eighth Annual Conference on Neural Information Processing Systems. Vancouver: NeurIPS Foundation, 2024: 1-19.
LiS, LiS, ZhangL. Hyperspectral and panchromatic images fusion based on the dual conditional diffusion models[J]. IEEE Transactions on Geoscience and Remote Sensing, 2023, 61: 5526315.
[42]
HuangH, HeW, ZhangH, et al. STFDiff: Remote sensing image spatiotemporal fusion with diffusion models[J]. Information Fusion, 2024, 111: 102505.
[43]
RanW, YuanW, ShibasakiR. Few-shot depth completion using denoising diffusion probabilistic model[C]//2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). Vancouver:IEEE, 2023: 6559-6567.
[44]
Bar-TalO, YarivL, LipmanY, et al. MultiDiffusion: Fusing diffusion paths for controlled image generation[J]. Proceedings of Machine Learning Research, 2023, 202: 1737-1752.
[45]
TangL, DengY, YiX, et al. DRMF: Degradation-robust multi-modal image fusion via composable diffusion prior[C]//Proceedings of the 32nd ACM International Conference on Multimedia. Melbourne:ACM, 2024: 8546-8555.
[46]
LinJ, WangY, TaoZ, et al. Adaptive multi-modal fusion of spatially rariant kernel refinement with diffusion model for blind image super-resolution[C]//Computer Vision-ECCV. Cham: Springer, 2024: 363-380.
[47]
PanX, ZhaiH, YangY, et al. Improving multi-focus image fusion through noisy image and feature difference network[J]. Image and Vision Computing, 2024, 142: 104891.
[48]
JiJ, ZhangY, LinZ, et al. Infrared and visible image fusion based on iterative control of anisotropic diffusion and regional gradient structure[J]. J Sensors, 2022, 2022: 1-10.
[49]
DucayR, MessingerD W. Image fusion of hyperspectral and multispectral imagery using nearest-neighbor diffusion[J]. Journal of Applied Remote Sensing, 2023, 17(2): 1-14.
[50]
DucayR, MessingerD W. Hyperspectral-multispectral image fusion using nearest-neighbor diffusion-based sharpening algorithm[C]//Algorithms, Technologies, and Applications for Multispectral and Hyperspectral Imaging XXVIII. Orlando:SPIE, 2022: 22.
[51]
YangX, ZhouD, FengJ, et al. Diffusion probabilistic model made slim[C]//2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver:IEEE, 2023: 22552-22562.
[52]
PanS, WangT, QiuR L J, et al. 2D medical image synthesis using transformer-based denoising diffusion probabilistic model[J]. Physics in Medicine and Biology, 2023, 68(10): 105004.
[53]
YangB, JiangZ, PanD, et al. LFDT-Fusion: A latent feature-guided diffusion Transformer model for general image fusion[J]. Information Fusion, 2025, 113: 102639.
[54]
YangZ, QinH, ZengS, et al. DGFusion: A novel infrared and visible image fusion method based on diffusion and generative adversarial networks[J]. IEEE Access, 2024, 12: 147051-147064.
[55]
ZhaoL, SongD, ChenW, et al. Coloring and fusing architectural sketches by combining a Y-shaped generative adversarial network and a denoising diffusion implicit model[J]. Computer-Aided Civil and Infrastructure Engineering, 2024, 39(7): 1003-1018.
[56]
LuX, JiangY, HongH, et al. DCAFuse: Dual-branch diffusion-CNN complementary feature aggregation network for multi-modality image fusion[C]//Proceedings of the 32nd ACM International Conference on Multimedia. Melbourne:ACM, 2024: 1524-1533.
[57]
HanG, ZhangX, HuangY. An end-to-end based on semantic region guidance for infrared and visible image fusion[J]. Signal, Image and Video Processing, 2024, 18(1): 295-303.
[58]
GuoH, ChenM, LiK, et al. GLAD: A global-attention-based diffusion model for infrared and visible image fusion[C]//Advanced Intelligent Computing Technology and Applications. Singapore:Springer, 2024: 345-356.
[59]
FizaS, SafinazS. Multi-focus image fusion using edge discriminative diffusion filter for satellite images[J]. Multimedia Tools and Applications, 2024, 83(25): 66087-66106.
[60]
KhaderF, Muller-FranzesG, Tayebi ArastehS, et al. Denoising diffusion probabilistic models for 3D medical image generation[J]. Scientific Reports, 2023, 13(1): 7303.
[61]
DingW, GengS, WwangH, et al. FDiff-Fusion: Denoising diffusion fusion network based on fuzzy learning for 3D medical image segmentation[J]. Information Fusion, 2024, 112: 102540.
[62]
ZhuR, LiX, ZhangX, et al. MRI and CT medical image fusion based on synchronized-anisotropic diffusion model[J]. IEEE Access, 2020, 8: 91336-91350.
[63]
AakaaramV, BachuS. MRI and CT image fusion using synchronized anisotropic diffusion equation with DT-CWT decomposition[C]//2022 Smart Technologies, Communication and Robotics (STCR). Sathyamangalam:IEEE, 2022: 1-5.
[64]
GoyalB, DograA, KhoondR, et al. Medical image fusion based on anisotropic diffusion and non-subsampled contourlet transform[J]. Computers, Materials & Continua, 2023, 76(1): 311-327.
[65]
LepchaD C, DograA, GoyalB, et al. Multimodal medical image fusion based on pixel significance using anisotropic diffusion and cross bilateral filter[J]. Human-Centered Computing and Information Sciences, 2022, 12:1-20.
[66]
ZhangC, WangZ, HanJ. Diffusion posterior sampling for remote sensing image fusion[C]//7th International Conference on Vision, Image and Signal Processing (ICVISP 2023). London: IET, 2023: 28-33.
[67]
ZhangC, ChangY, WuY, et al. Semantic information guided diffusion posterior sampling for remote sensing image fusion[J]. Scientific Reports, 2024, 14(1): 27259.
[68]
ZhangC, ChangY, WuY, et al. Binary diffusion method for image fusion with despeckling[C]//2024 2nd International Conference on Algorithm, Image Processing and Machine Vision (AIPMV). Zhenjiang: IEEE, 2024: 243-248.
[69]
WeiJ, GanL, TangW, et al. Diffusion models for spatio-temporal-spectral fusion of homogeneous Gaofen-1 satellite platforms[J]. International Journal of Applied Earth Observation and Geoinformation, 2024, 128: 103752.
[70]
MaY, WangQ, WeiJ. Spatiotemporal fusion via conditional diffusion model[J]. IEEE Geoscience and Remote Sensing Letters, 2024, 21: 5002405.
[71]
CaoZ, CaoS, DengL J, et al. Diffusion model with disentangled modulations for sharpening multispectral and hyperspectral images[J]. Information Fusion, 2024, 104: 102158.
[72]
YueJ, FangL, XiaS, et al. Dif-fusion: Toward high color fidelity in infrared and visible image fusion with diffusion models[J]. IEEE Transactions on Image Processing, 2023, 32: 5705-5720.