An Image is Worth a Thousand Toxic Words: A Metamorphic Testing Framework for Content Moderation Software

08/18/2023
by   Wenxuan Wang, et al.
0

The exponential growth of social media platforms has brought about a revolution in communication and content dissemination in human society. Nevertheless, these platforms are being increasingly misused to spread toxic content, including hate speech, malicious advertising, and pornography, leading to severe negative consequences such as harm to teenagers' mental health. Despite tremendous efforts in developing and deploying textual and image content moderation methods, malicious users can evade moderation by embedding texts into images, such as screenshots of the text, usually with some interference. We find that modern content moderation software's performance against such malicious inputs remains underexplored. In this work, we propose OASIS, a metamorphic testing framework for content moderation software. OASIS employs 21 transform rules summarized from our pilot study on 5,000 real-world toxic contents collected from 4 popular social media applications, including Twitter, Instagram, Sina Weibo, and Baidu Tieba. Given toxic textual contents, OASIS can generate image test cases, which preserve the toxicity yet are likely to bypass moderation. In the evaluation, we employ OASIS to test five commercial textual content moderation software from famous companies (i.e., Google Cloud, Microsoft Azure, Baidu Cloud, Alibaba Cloud and Tencent Cloud), as well as a state-of-the-art moderation research model. The results show that OASIS achieves up to 100 models with the test cases generated by OASIS, the robustness of the moderation model can be improved without performance degradation.

READ FULL TEXT
research
02/11/2023

MTTM: Metamorphic Testing for Textual Content Moderation Software

The exponential growth of social media platforms such as Twitter and Fac...
research
05/23/2023

Validating Multimedia Content Moderation Software via Semantic Fusion

The exponential growth of social media platforms, such as Facebook and T...
research
01/11/2022

Captcha Attack: Turning Captchas Against Humanity

Nowadays, people generate and share massive content on online platforms ...
research
09/12/2023

Catch You Everything Everywhere: Guarding Textual Inversion via Concept Watermarking

AIGC (AI-Generated Content) has achieved tremendous success in many appl...
research
02/19/2019

Fusing Visual, Textual and Connectivity Clues for Studying Mental Health

With ubiquity of social media platforms, millions of people are sharing ...
research
10/03/2019

From Senseless Swarms to Smart Mobs: Tuning Networks for Prosocial Behaviour

Social media have been seen to accelerate the spread of negative content...
research
05/03/2022

Themes of Revenge: Automatic Identification of Vengeful Content in Textual Data

Revenge is a powerful motivating force reported to underlie the behavior...

Please sign up or login with your details

Forgot password? Click here to reset