Authors - Ndaula Kelvin, Wu Jun Abstract - Multimodal sentiment analysis often fails due to the modality gap between semantic text images and GIFs. To address these challenges this paper introduces Fusion Core a novel hardware agnostic heterogeneous pipeline that bridges the gap between high level AI with low level systems engineering to facilitate the deciphering of combined sentiment of these three modalities. Through the integration of GPGPU accelerated OpenCL kernels for 3D temporal extraction with an ONNX/DirectML inference engine which ensures cross platform portability. To address the issue of inconsistent real world data distributions the preprocessing system was introduced with an adaptive multi-head attention mechanism for late feature fusion. The ablation studies performed also revealed the integration of spatiotemporal GIF layers resolves contextual ambiguities missed by static analysis (Text, Images). The extensive testing on a balanced dataset of 13,964 samples the model achieved a 100% success rate showing the robustness of the proposed model for industrial scale deployment.