Key Laboratory of Aerospace Information Security and Trusted Computing,Ministry of Education,School of Cyber Science and Engineering,Wuhan University,Wuhan 430072,Hubei,China
Most mainstream deepfake detection methods are based on visual single-modality, which detects deepfake videos by identifying fake artifacts in video frames. However, different deepfake methods may introduce different artifacts; therefore such methods have limited performance and poor generalization ability. Meanwhile, the current work utilizing audio information does not fully leverage audio information to uncover tampering in the visual modality. Natural videos have intrinsic synergy between speeches and faces, and deepfake methods this disrupt the intrinsic synergy between speeches and faces; this paper proposes a speech-face synergy-driven deepfake detection algorithm, called SFSD (Speech-Face Synergy Detection). The core of the algorithm is the audio-face synergy contrastive learning strategy, which constructs samples on a general video dataset to simulate the disruption of audio-face synergy using forgery methods and pre-trains the detection model on these samples. This strategy utilizes a large number of unlabeled real videos and enhances model performance and generalization capability. A multimodal model named SFformer (Speech-Face Transformer) is constructed, which utilizes an attention bottleneck to guide the condensation and fusion of essential information in the audio-face modality, reduces interference from redundant information, improves the model's feature extraction capability, and enhances detection performance. Numerous experiments on the public dataset FakeAVCeleb demonstrate that the accuracy of SFSD can reach 72.12% after pre-training, surpassing some benchmark methods, and the accuracy reaches 89.51% after transfer learning, which is higher than that of previous studies, and improves the generalization ability.
KORSHUNOVAI, SHIW Z, DAMBREJ, et al. Fast face-swap using convolutional neural networks[DB/OL]. [2023-02-02]. DOI: 10.1109/iccv.2017.397 .
[4]
NIRKINY, KELLERY, HASSNERT. FSGAN: Subject agnostic face swapping and reenactmen[DB/OL]. [2023-02-02]. DOI: 10.1109/iccv.2019.00728 .
[5]
KOUJANM R, DOUKASM C, ROUSSOSA, et al. Head2Head: Video-based neural head synthesis[C]//2020 15th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2020). New York: IEEE Press, 2020: 16-23. DOI: 10.1109/FG47880.2020.00048 .
[6]
PRAJWALK R, MUKHOPADHYAYR, NAMBOODIRIV P, et al. A lip sync expert is all you need for speech to lip generation in the wild[DB/OL]. [2023-02-02]. DOI: 10.1145/3394171.3413532 .
[7]
PEROVI, GAOD H, CHERVONIYN, et al. DeepFaceLab: Integrated, flexible and extensible face-swapping framework[EB/OL]. 2020: arXiv: 2005.05535.
[8]
NGUYENH H, YAMAGISHIJ, ECHIZENI. Capsule-forensics: Using capsule networks to detect forged images and videos[C]//2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). New York: IEEE Press, 2019: 2307-2311. DOI: 10.1109/ICASSP.2019.8682602 .
[9]
LIY Z, LYUS W. Exposing DeepFake videos by detecting face warping artifacts[EB/OL]. 2018: arXiv: 1811.00656. DOI: 10.1109/wifs.2018.8630787 .
[10]
RÖSSLERA, COZZOLINOD, VERDOLIVAL, et al. FaceForensics ++: Learning to detect manipulated facial images[DB/OL]. [2023-01-12]. DOI: 10.1109/iccv.2019.00009 .
[11]
AFCHARD, NOZICKV, YAMAGISHIJ, et al. MesoNet: A compact facial video forgery detection network[C]//2018 IEEE International Workshop on Information Forensics and Security (WIFS). New York: IEEE Press, 2018: 1-7. DOI: 10.1109/WIFS.2018.8630761 .
[12]
CHUGHK, GUPTAP, DHALLA, et al. Not made for each other- audio-visual dissonance-based deepfake detection and localization[DB/OL]. [2023-02-03]. DOI: 10.1145/3394171.3413700 .
MITTALT, BHATTACHARYAU, CHANDRAR, et al. Emotions don’t lie: An audio-visual deepfake detection method using affective cues[DB/OL]. [2023-02-01]. DOI: 10.1145/3394171.3413570 .
[15]
VASWANIA, SHAZEERN, PARMARN, et al. Attention is all you need[DB/OL]. [2024-10-11].
[16]
XUP, ZHUX T, CLIFTOND A. Multimodal learning with transformers: A survey[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(10): 12113-12132. DOI: 10.1109/TPAMI.2023.3275156 .
[17]
NAGRANIA, YANGS, ARNABA, et al. Attention bottlenecks for multimodal fusion[EB/OL]. 2021: arXiv: 2107.00135.
[18]
NGUYENT T, NGUYENQ V H, NGUYEND T, et al. Deep learning for deepfakes creation and detection: A survey[J]. Computer Vision and Image Understanding, 2022, 223: 103525. DOI: 10.1016/j.cviu.2022.103525 .
[19]
LIL Z, BAOJ M, YANGH, et al. Advancing high fidelity identity swapping for forgery detection[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2020: 5074-5083. DOI: 10.1109/CVPR42600.2020.00512 .
[20]
JAMALUDINA, CHUNGJ S, ZISSERMANA. You said that: Synthesising talking faces from audio[J]. International Journal of Computer Vision, 2019, 127(11): 1767-1779. DOI: 10.1007/s11263-019-01150-y .
[21]
ZHOUH, SUNY S, WUW, et al. Pose-controllable talking face generation by implicitly modularized audio-visual representation[C]//2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2021: 4176-4186. DOI: 10.1109/CVPR46437.2021.00416 .
SUNP, LANGY B, GONGJ C, et al. Authentication method for splicing manipulation with inconsistencies in color shift[J]. Journal of Computer-Aided Design & Computer Graphics, 2017, 29(8): 1408-1415 (Ch).
ZHANGY X, LIG, CAOY, et al. A method for detecting human-face-tampered videos based on interframe difference[J]. Journal of Cyber Security, 2020, 5(2): 49-72. DOI: 10.19363/J.cnki.cn10-1380/tn.2020.02.05 (Ch ).
[26]
MASII, KILLEKARA, MASCARENHASR M, et al. Two-branch recurrent network for isolating deepfakes in videos[M]// Computer Vision - ECCV 2020. Cham: Springer International Publishing, 2020: 667-684. DOI: 10.1007/978-3-030-58571-6_39 .
[27]
WUX, XIEZ, GAOY T, et al. SSTNet: Detecting manipulated faces through spatial, steganalysis and temporal features[C]//2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). New York: IEEE Press, 2020: 2952-2956. DOI: 10.1109/ICASSP40776.2020.9053969 .
[28]
QIANY Y, YING J, SHENGL, et al. Thinking in frequency: Face forgery detection by mining frequency-aware clues[C]//European Conference on Computer Vision. Cham: Springer, 2020: 86-103.10.1007/978-3-030-58610-2_6. DOI: 10.1007/978-3-030-58610-2_6 .
[29]
WANGS Y, WANGO, ZHANGR, et al. CNN-generated images are surprisingly easy to spot for now[DB/OL]. [2023-01-05]. DOI: 10.1109/cvpr42600.2020.00872 .
[30]
ZHOUY P, LIMS N. Joint audio-visual deepfake detection[C]//2021 IEEE/CVF International Conference on Computer Vision (ICCV). New York: IEEE Press, 2021: 14780-14789. DOI: 10.1109/ICCV48922.2021.01453 .
[31]
ILYASH, JAVEDA, MALIKK M. AVFakeNet: A unified end-to-end Dense Swin Transformer deep learning model for audio-visualdeepfakes detection[J]. Applied Soft Computing, 2023, 136: 110124. DOI: 10.1016/j.asoc.2023.110124 .
[32]
LIUZ, LINY T, CAOY, et al. Swin Transformer: Hierarchical vision transformer using shifted windows[C]//2021 IEEE/CVF International Conference on Computer Vision (ICCV). New York: IEEE Press, 2021: 9992-10002. DOI: 10.1109/ICCV48922.2021.00986 .
[33]
CAIZ X, STEFANOVK, DHALLA, et al. Do you really mean that? content driven audio-visual deepfake dataset and multimodal method for temporal forgery localization[C]//2022 International Conference on Digital Image Computing: Techniques and Applications (DICTA). New York: IEEE Press, 2022: 1-10. DOI: 10.1109/DICTA56598.2022.10034605 .
[34]
YANGW Y, ZHOUX Y, CHENZ K, et al. AVoiD-DF: Audio-visual joint learning for detecting deepfake[J]. IEEE Transactions on Information Forensics and Security, 2023, 18: 2015-2029. DOI: 10.1109/TIFS.2023.3262148 .
[35]
KAMACHIM, HILLH, LANDERK, et al. Putting the face to the voice: Matching identity across modality [J]. Current Biology, 2003, 13(19): 1709-1714. DOI: 10.1016/j.cub.2003.09.005 .
[36]
NAGRANIA, ALBANIES, ZISSERMANA. Seeing voices and hearing faces: Cross-modal biometric matching[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2018: 8427-8436. DOI: 10.1109/CVPR.2018.00879 .
[37]
WANGR, LIUX, CHEUNGY M, et al. Learning discriminative joint embeddings for efficient face and voice association[DB/OL]. [2023-04-02]. DOI: 10.1145/3397271.3401302 .
[38]
ZHUB Q, XUK L, WANGC J, et al. Unsupervised voice-face representation learning by cross-modal prototype contrast[DB/OL]. [2023-03-12]. DOI: 10.24963/ijcai.2022/526 .
[39]
MALIKM, MALIKM K, MEHMOODK, et al. Automatic speech recognition: A survey[J]. Multimedia Tools and Applications, 2021, 80(6): 9411-9457. DOI: 10.1007/s11042-020-10073-7 .
[40]
CHUNGJ S, SENIORA, VINYALSO, et al. Lip reading sentences in the wild[DB/OL]. [2023-03-14]. DOI: 10.1109/cvpr.2017.367 .
[41]
CHUNGJ S, ZISSERMANA. Out of time: Automated lip sync in the wild[DB/OL]. [2023-01-23]. DOI: 10.1007/978-3-319-54427-4_19 .
[42]
ZHOUH, LIUY, LIUZ W, et al. Talking face generation by adversarially disentangled audio-visual representation[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2019, 33(1): 9299-9306. DOI: 10.1609/aaai.v33i01.33019299 .
[43]
DOSOVITSKIYA, BEYERL, KOLESNIKOVA, et al. An image is worth 16x16 words: Transformers for image recognition at scale[EB/OL]. 2020: arXiv: 2010.11929.
[44]
HEK M, FANH Q, WUY X, et al. Momentum contrast for unsupervised visual representation learning[DB/OL]. [2023-02-06]. DOI: 10.1109/cvpr42600.2020.00975 .
AFOURAST, CHUNGJ S, SENIORA, et al. Deep audio-visual speech recognition[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 44(12): 8717-8727. DOI: 10.1109/TPAMI.2018.2889052 .
[47]
KHALIDH, TARIQS, KIMM, et al. FakeAVCeleb: A novel audio-video multimodal deepfake dataset[EB/OL]. 2021: arXiv: 2108.05080. DOI: 10.1145/3476099.3484315 .
[48]
KORSHUNOVP, MARCELS, KORSHUNOVP, et al. DeepFakes: A new threat to face recognition? assessment and detection[EB/OL]. 2018: arXiv: 1812.08685.
[49]
SANDERSONC, LOVELLB C. Multi-region probabilistic histograms for robust and scalable identity inference[C]//International Conference on Biometrics. Heidelberg: Springer, 2009: 199-208. DOI: 10.1007/978-3-642-01793-3_21 .
[50]
KINGMAD P, BAJ, HAMMADM M. Adam: A method for stochastic optimization[EB/OL]. 2014: arXiv: 1412.6980.
[51]
KHALIDH, KIMM, TARIQS, et al. Evaluation of an audio-video multimodal deepfake dataset using unimodal and multimodal detectors[DB/OL]. [2023-02-10]. DOI: 10.1145/3476099.3484315 .