WVisionBERT-VL: a multimodal model architecture for toxicity classification on social media platforms using large language models

Witta Listiya Ningrum, Achmad Benny Mutiara, Diana Ikasari

Abstract


The increasing prevalence of toxic content on social media, conveyed through text, images, and videos, poses significant challenges for automated content moderation systems. Although prior studies have reported promising results in unimodal and bimodal settings, they often fail to capture implicit and contextual toxicity emerging from interactions across multiple modalities, particularly in non-English environments. This paper proposes WVisionBERT-VL, an end-to-end multimodal framework for toxicity detection that integrates text, image, and video modalities within a unified architecture. The proposed model incorporates modality-specific encoders, bidirectional multi-head cross attention (BMHCA) for cross-modal synchronization, and an adaptive fusion gate to dynamically balance modality contributions. A balanced multimodal dataset is constructed from social media platforms, including X, Instagram, and TikTok, and refined using a model-based labeling strategy with limited human-in-the-loop validation. Experimental results on a custom Indonesian dataset demonstrate strong in-domain performance, achieving an accuracy of 94.12%, a macro-F1 of 0.9407, and a receiver operating characteristic - area under the curve (ROC-AUC) of 0.9721, with robustness further validated through five-fold cross-validation. Cross-dataset evaluation highlights challenges related to domain shift, underscoring the need for future research on robust and domain-adaptive multimodal toxicity detection.

Keywords


Content moderation; Cross-modal attention; Multimodal toxicity detection; Social media analysis; Transformer-based fusio

Full Text:

PDF


DOI: http://doi.org/10.11591/ijai.v15.i4.pp3365-3375

Refbacks

  • There are currently no refbacks.


Copyright (c) 2026 Witta Listiya Ningrum, Achmad Benny Mutiara, Diana Ikasari

Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

IAES International Journal of Artificial Intelligence (IJ-AI)
ISSN/e-ISSN 2089-4872/2252-8938 
This journal is published by the Institute of Advanced Engineering and Science (IAES).

View IJAI Stats