Assessing Linguistic Generalisation in Language Models: A Dataset for Brazilian Portuguese

05/23/2023
by   Rodrigo Wilkens, et al.
0

Much recent effort has been devoted to creating large-scale language models. Nowadays, the most prominent approaches are based on deep neural networks, such as BERT. However, they lack transparency and interpretability, and are often seen as black boxes. This affects not only their applicability in downstream tasks but also the comparability of different architectures or even of the same model trained using different corpora or hyperparameters. In this paper, we propose a set of intrinsic evaluation tasks that inspect the linguistic information encoded in models developed for Brazilian Portuguese. These tasks are designed to evaluate how different language models generalise information related to grammatical structures and multiword expressions (MWEs), thus allowing for an assessment of whether the model has learned different linguistic phenomena. The dataset that was developed for these tasks is composed of a series of sentences with a single masked word and a cue phrase that helps in narrowing down the context. This dataset is divided into MWEs and grammatical structures, and the latter is subdivided into 6 tasks: impersonal verbs, subject agreement, verb agreement, nominal agreement, passive and connectors. The subset for MWEs was used to test BERTimbau Large, BERTimbau Base and mBERT. For the grammatical structures, we used only BERTimbau Large, because it yielded the best results in the MWE task.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
08/24/2018

Under the Hood: Using Diagnostic Classifiers to Investigate and Improve how Language Models Track Agreement Information

How do neural language models keep track of number agreement between sub...
research
10/31/2022

Emergent Linguistic Structures in Neural Networks are Fragile

Large language models (LLMs) have been reported to have strong performan...
research
05/24/2021

Neural Language Models for Nineteenth-Century English

We present four types of neural language models trained on a large histo...
research
05/22/2022

Blackbird's language matrices (BLMs): a new benchmark to investigate disentangled generalisation in neural networks

Current successes of machine learning architectures are based on computa...
research
10/05/2020

Linguistic Profiling of a Neural Language Model

In this paper we investigate the linguistic knowledge learned by a Neura...
research
09/05/2020

Visually Analyzing Contextualized Embeddings

In this paper we introduce a method for visually analyzing contextualize...

Please sign up or login with your details

Forgot password? Click here to reset