The Challenges of Machine Learning for Trust and Safety: A Case Study on Misinformation Detection

08/23/2023
by   Madelyne Xiao, et al.
0

We examine the disconnect between scholarship and practice in applying machine learning to trust and safety problems, using misinformation detection as a case study. We systematize literature on automated detection of misinformation across a corpus of 270 well-cited papers in the field. We then examine subsets of papers for data and code availability, design missteps, reproducibility, and generalizability. We find significant shortcomings in the literature that call into question claimed performance and practicality. Detection tasks are often meaningfully distinct from the challenges that online services actually face. Datasets and model evaluation are often non-representative of real-world contexts, and evaluation frequently is not independent of model training. Data and code availability is poor. Models do not generalize well to out-of-domain data. Based on these results, we offer recommendations for evaluating machine learning applications to trust and safety problems. Our aim is for future work to avoid the pitfalls that we identify.

READ FULL TEXT

page 10

page 12

research
04/24/2023

Can we Trust Chatbots for now? Accuracy, reproducibility, traceability; a Case Study on Leonardo da Vinci's Contribution to Astronomy

Large Language Models (LLM) are studied. Applications to chatbots and ed...
research
07/03/2023

Adversarial Learning in Real-World Fraud Detection: Challenges and Perspectives

Data economy relies on data-driven systems and complex machine learning ...
research
09/16/2019

Searching for Better Test Case Prioritization Schemes: a Case Study of AI-assisted Systematic Literature Review

Given the large numbers of publications in the SE field, it is difficult...
research
11/27/2019

To Trust, or Not to Trust? A Study of Human Bias in Automated Video Interview Assessments

Supervised systems require human labels for training. But, are humans th...
research
01/11/2021

A Framework for Assurance of Medication Safety using Machine Learning

Medication errors continue to be the leading cause of avoidable patient ...
research
12/03/2020

Ethical Testing in the Real World: Evaluating Physical Testing of Adversarial Machine Learning

This paper critically assesses the adequacy and representativeness of ph...
research
10/05/2022

A Pilot Study of Sidewalk Equity in Seattle Using Crowdsourced Sidewalk Assessment Data

We examine the potential of using large-scale open crowdsourced sidewalk...

Please sign up or login with your details

Forgot password? Click here to reset