A Cross-Verification Approach for Protecting World Leaders from Fake and Tampered Audio

by   Mengyi Shan, et al.

This paper tackles the problem of verifying the authenticity of speech recordings from world leaders. Whereas previous work on detecting deep fake or tampered audio focus on scrutinizing an audio recording in isolation, we instead reframe the problem and focus on cross-verifying a questionable recording against trusted references. We present a method for cross-verifying a speech recording against a reference that consists of two steps: aligning the two recordings and then classifying each query frame as matching or non-matching. We propose a subsequence alignment method based on the Needleman-Wunsch algorithm and show that it significantly outperforms dynamic time warping in handling common tampering operations. We also explore several binary classification models based on LSTM and Transformer architectures to verify content at the frame level. Through extensive experiments on tampered speech recordings of Donald Trump, we show that our system can reliably detect audio tampering operations of different types and durations. Our best model achieves 99.7 and a 0.43 non-matching.



There are no comments yet.


page 1

page 2

page 3

page 4


Speech watermarking: an approach for the forensic analysis of digital telephonic recordings

In this article, the authors discuss the problem of forensic authenticat...

Low Resource Audio-to-Lyrics Alignment From Polyphonic Music Recordings

Lyrics alignment in long music recordings can be memory exhaustive when ...

Half-Truth: A Partially Fake Audio Detection Dataset

Diverse promising datasets have been designed to hold back the developme...

Parkinson's disease diagnostics using AI and natural language knowledge transfer

In this work, the issue of Parkinson's disease (PD) diagnostics using no...

Improving Audio Anomalies Recognition Using Temporal Convolutional Attention Network

Anomalous audio in speech recordings is often caused by speaker voice di...

Diarization of Legal Proceedings. Identifying and Transcribing Judicial Speech from Recorded Court Audio

United States Courts make audio recordings of oral arguments available a...

Reliability of Power System Frequency on Times-Stamping Digital Recordings

Power system frequency could be captured by digital recordings and extra...
This week in AI

Get the week's most popular data science and artificial intelligence research sent straight to your inbox every Saturday.