Few-Shot and Zero-Shot Learning for Historical Text Normalization

03/12/2019
by   Marcel Bollmann, et al.
0

Historical text normalization often relies on small training datasets. Recent work has shown that multi-task learning can sometimes lead to significant improvements by exploiting synergies with related datasets, but there has been no systematic study of multi-task learning strategies across different datasets from different languages. This paper evaluates 63 multi-task learning strategies for sequence-to-sequence-based historical text normalization across ten datasets from eight languages, using autoencoding, grapheme-to-phoneme mapping, and lemmatization as auxiliary tasks. We observe consistent, significant improvements across languages when training data for the target task is limited, but minimal or no improvements when training data is abundant. Finally, we show that zero-shot learning outperforms the simple, but relatively strong, identity baseline.

READ FULL TEXT
research
12/14/2022

Multi-task Learning for Cross-Lingual Sentiment Analysis

This paper presents a cross-lingual sentiment analysis of news articles ...
research
12/23/2014

A Unified Perspective on Multi-Domain and Multi-Task Learning

In this paper, we provide a new neural-network based perspective on mult...
research
10/25/2016

Improving historical spelling normalization with bi-directional LSTMs and multi-task learning

Natural-language processing of historical documents is complicated by th...
research
12/17/2022

Improving Cross-task Generalization of Unified Table-to-text Models with Compositional Task Configurations

There has been great progress in unifying various table-to-text tasks us...
research
04/03/2019

A Large-Scale Comparison of Historical Text Normalization Systems

There is no consensus on the state-of-the-art approach to historical tex...
research
04/29/2022

Task Embedding Temporal Convolution Networks for Transfer Learning Problems in Renewable Power Time-Series Forecast

Task embeddings in multi-layer perceptrons for multi-task learning and i...
research
10/21/2022

Generalizing over Long Tail Concepts for Medical Term Normalization

Medical term normalization consists in mapping a piece of text to a larg...

Please sign up or login with your details

Forgot password? Click here to reset