Request Correction
Use this form to request corrections to the paper metadata. Select the fields that need correction and provide the correct information.
Correction Guidelines
- Click the edit button next to a field to report a correction.
- Fill in the suggested correction value for each field you want to correct.
- Provide your name and email so we can contact you if needed.
Correction requests are currently disabled.
Paper Information
Human-Centered Multimodal Fusion for Sexism Detection in Memes with Eye-Tracking, Heart Rate, and EEG Signals
Paper Fields
Click the edit button next to a field to report a correction.
Human-Centered Multimodal Fusion for Sexism Detection in Memes with Eye-Tracking, Heart Rate, and EEG Signals
The automated detection of sexism in memes is a notoriously challenging task due to multimodal ambiguity, cultural nuance, and the use of humor to provide plausible deniability. As a result, content-only models often fail to capture the complexity of human perception. To address this fundamental limitation, we introduce and validate a human-centered paradigm that augments standard content features with rich physiological data. We created a novel resource by recording Eye-Tracking (ET), Heart Rate (HR), and Electroencephalography (EEG) from 16 subjects (8 per experiment) while they viewed 3,984 memes from the EXIST 2025 dataset. Our statistical analysis reveals significant physiological differences in how subjects process sexist versus non-sexist content. Sexist memes were associated with higher cognitive load (evidenced by increased fixation counts and longer reaction times), and with differences in EEG spectral power across the Alpha, Beta, and Gamma frequency bands. This pattern, commonly linked in previous research to increased attentional engagement and cognitive effort during visual processing, suggests that sexist memes may elicit more demanding neural activity compared to non-sexist ones. Building on these findings, we propose a novel multimodal fusion model that integrates these physiological signals with enriched textual-visual features derived from a Vision-Language Model (VLM). Our final model achieves an AUC of 0.794 in binary sexism detection, a statistically significant 3.4% improvement over a powerful VLM-based baseline. The fusion of physiological data proves particularly effective for nuanced and ambiguous cases, boosting the F1-score for the most challenging fine-grained category, *Misogyny and Non-Sexual Violence*, by an unprecedented 26.3%. Our work demonstrates that human physiological responses provide a robust, objective signal of perception that can significantly enhance the accuracy and human-awareness of automated systems for countering online sexism.
Authors
Expand an author to correct their information. Use the remove button to request author removal, or add a new author.
PDF Attachment
You may attach a PDF as a corrected version of the paper. Max file size: 10MB. Only PDF files are accepted.
Your Information
Author Declaration *
Select at least one field to correct using the edit buttons above.