Can You Fool AI by Doing a 180? x2013 A Case Study on Authorship Analysis of Texts by Arata Osada

by   Jagna Nieuwazny, et al.

This paper is our attempt at answering a twofold question covering the areas of ethics and authorship analysis. Firstly, since the methods used for performing authorship analysis imply that an author can be recognized by the content he or she creates, we were interested in finding out whether it would be possible for an author identification system to correctly attribute works to authors if in the course of years they have undergone a major psychological transition. Secondly, and from the point of view of the evolution of an author's ethical values, we checked what it would mean if the authorship attribution system encounters difficulties in detecting single authorship. We set out to answer those questions through performing a binary authorship analysis task using a text classifier based on a pre-trained transformer model and a baseline method relying on conventional similarity metrics. For the test set, we chose works of Arata Osada, a Japanese educator and specialist in the history of education, with half of them being books written before the World War II and another half in the 1950s, in between which he underwent a transformation in terms of political opinions. As a result, we were able to confirm that in the case of texts authored by Arata Osada in a time span of more than 10 years, while the classification accuracy drops by a large margin and is substantially lower than for texts by other non-fiction writers, confidence scores of the predictions remain at a similar level as in the case of a shorter time span, indicating that the classifier was in many instances tricked into deciding that texts written over a time span of multiple years were actually written by two different people, which in turn leads us to believe that such a change can affect authorship analysis, and that historical events have great impact on a person's ethical outlook as expressed in their writings.


page 16

page 18


Analyzing Stylistic Variation across Different Political Regimes

In this article we propose a stylistic analysis of texts written across ...

A comparison of several AI techniques for authorship attribution on Romanian texts

Determining the author of a text is a difficult task. Here we compare mu...

BERT-based Authorship Attribution on the Romanian Dataset called ROST

Being around for decades, the problem of Authorship Attribution is still...

BERT in Plutarch's Shadows

The extensive surviving corpus of the ancient scholar Plutarch of Chaero...

Why Molière most likely did write his plays

As for Shakespeare, a hard-fought debate has emerged about Molière, a su...

The Trumpiest Trump? Identifying a Subject's Most Characteristic Tweets

The sequence of documents produced by any given author varies in style a...

QuALITY: Question Answering with Long Input Texts, Yes!

To enable building and testing models on long-document comprehension, we...

Please sign up or login with your details

Forgot password? Click here to reset