A MULTIMODAL DEEP LEARNING FRAMEWORK FOR DETECTING AI-GENERATED CONTENT ON SOCIAL MEDIA
Keywords:
Multimodal deep learning, AI-generated content, deepfake detection, misinformation, social media, explainable AI, cross-modal fusionAbstract
Generative artificial intelligence has reached a point where machine-written text, synthetic photographs, cloned voices, and manipulated video clips are often impossible to tell apart from authentic material by eye alone. Because social media platforms carry all of these content types side by side, and because a single misleading post frequently combines several of them at once, detection systems that look at only one channel text alone, or an image alone tend to miss a large share of manipulated content and struggle badly when confronted with a generator they were never trained on. This paper reviews the current state of AI-generated content detection across the text, image, audio, and behavioural-context channels, and proposes a multimodal deep learning framework that fuses these channels through a cross-modal attention layer rather than treating each as a separate problem. The framework combines a transformer-based text encoder, a convolutional/vision-transformer image branch strengthened with frequency-domain artifact analysis, an optional audio branch for voice and video-audio content, and a lightweight contextual branch that reads engagement and propagation metadata, before merging their representations into a single credibility judgment accompanied by a human-readable explanation. Drawing on thirty peer-reviewed and indexed studies published mainly between 2020 and 2025, the paper synthesises reported detection accuracies across modalities and finds that fused multimodal systems consistently outperform single-channel detectors by a margin of roughly five to ten percentage points, while also being more resistant to the kind of distribution shift that occurs when a new generative model appears. The paper closes with a discussion of explainability, watermarking, and content provenance as complementary rather than competing strategies, and outlines the empirical work still needed to validate the architecture on real, in-the-wild social media traffic.


