Czech Dataset for Cross-lingual Subjectivity Classification

04/29/2022
by   Pavel Přibáň, et al.
0

In this paper, we introduce a new Czech subjectivity dataset of 10k manually annotated subjective and objective sentences from movie reviews and descriptions. Our prime motivation is to provide a reliable dataset that can be used with the existing English dataset as a benchmark to test the ability of pre-trained multilingual models to transfer knowledge between Czech and English and vice versa. Two annotators annotated the dataset reaching 0.83 of the Cohen's ąp̨p̨ą inter-annotator agreement. To the best of our knowledge, this is the first subjectivity dataset for the Czech language. We also created an additional dataset that consists of 200k automatically labeled sentences. Both datasets are freely available for research purposes. Furthermore, we fine-tune five pre-trained BERT-like models to set a monolingual baseline for the new dataset and we achieve 93.56 English dataset for which we obtained results that are on par with the current state-of-the-art results. Finally, we perform zero-shot cross-lingual subjectivity classification between Czech and English to verify the usability of our dataset as the cross-lingual benchmark. We compare and discuss the cross-lingual and monolingual results and the ability of multilingual models to transfer knowledge between languages.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
02/15/2021

Beyond the English Web: Zero-Shot Cross-Lingual and Lightweight Monolingual Classification of Registers

We explore cross-lingual transfer of register classification for web doc...
research
09/24/2021

Monolingual and Cross-Lingual Acceptability Judgments with the Italian CoLA corpus

The development of automated approaches to linguistic acceptability has ...
research
10/19/2022

Leveraging a New Spanish Corpus for Multilingual and Crosslingual Metaphor Detection

The lack of wide coverage datasets annotated with everyday metaphorical ...
research
09/28/2021

Multilingual Counter Narrative Type Classification

The growing interest in employing counter narratives for hatred interven...
research
04/03/2023

LAHM : Large Annotated Dataset for Multi-Domain and Multilingual Hate Speech Identification

Current research on hate speech analysis is typically oriented towards m...
research
01/25/2023

Cross-lingual Argument Mining in the Medical Domain

Nowadays the medical domain is receiving more and more attention in appl...
research
06/13/2023

Monolingual and Cross-Lingual Knowledge Transfer for Topic Classification

This article investigates the knowledge transfer from the RuQTopics data...

Please sign up or login with your details

Forgot password? Click here to reset