Loading…
10th WorldS4 2026 has ended
Thursday July 30, 2026 2:00pm - 3:30pm BST

Authors - Meghali Kalyankar, Prashant Lahane
Abstract - As deepfake generation technologies have rapidly advanced, establishing authenticity for multimedia content on digital platforms has become a major challenge[1]. Current deepfake detection approaches primarily rely on unimodal analysis and face challenges in encoding joint AV inconsistencies, temporal consistency, and adversarial attacks. To overcome these problems, a novel Multimodal Attention and Adversarial Deepfake Network (MMAD-Net) framework for robust audiovisual deepfake detection is proposed. The proposed solution adopts a framework that combines fine-grained visual feature extraction provided by VideoMAE v2 [2], audio representation learning provided by HuBERT [3] and temporal dependency modeling provided by TimeSformer [4] to model the fine-grained spatial and temporal inconsistency in manipulated media. In addition, a Multimodal CoAttention Transformer (MCAT) is used for better cross-modal interaction between audio and visual streams, and a hybrid HOA-COA optimization scheme optimizes discriminative feature representations to ensure better feature separation and remove redundancy. Adversarial Consistency Training (ACT) is embedded in the learning process to enhance adversarial robustness against adversarial perturbation and unseen adversarial manipulation. FakeAVCeleb and Celeb-DF are used for testing the proposed model with several performance metrics. Experimental results show that MMAD-Net can provide stable and general detection performance while maintaining a high level of robustness for current multimodal deepfake detection methods.
Paper Presenters
Thursday July 30, 2026 2:00pm - 3:30pm BST
Virtual Room E London, UK

Sign up or log in to save this to your schedule, view media, leave feedback and see who's attending!

Share Modal

Share this link via

Or copy link