Authors - Caitlin Loh, Susmitha Vekkot, Pancham Shukla Abstract - Deepfakes generated using modern machine learning techniques pose growing risks to digital trust by enabling realistic manipulation of both audio and video content. Many existing detection approaches rely on a single modality, limiting robustness when confronted with increasingly sophisticated forgeries. This paper evaluates a multimodal deepfake detection framework that integrates audio and visual information using deep learning. The proposed system combines stateof-the-art audio and visual encoders within a modular architecture and conducts a systematic comparison of three fusion strategies: early fusion, late fusion, and cross-attention. Experiments are conducted on the PolyGlotFake dataset, a multilingual benchmark containing synthetic and authentic audio–visual media. Results show that multimodal approaches substantially outperform unimodal baselines, with late fusion achieving an AUROC of 0.955 and cross-attention models reaching accuracies of up to 0.996. These findings provide a controlled comparison of fusion strategies and demonstrate that multimodal fusion significantly improves detection performance and highlights its potential for building more robust deepfake detection systems.