Topical Segmentation of Spoken Narratives: A Test Case on Holocaust Survivor Testimonies

10/25/2022
by   Eitan Wagner, et al.
0

The task of topical segmentation is well studied, but previous work has mostly addressed it in the context of structured, well-defined segments, such as segmentation into paragraphs, chapters, or segmenting text that originated from multiple sources. We tackle the task of segmenting running (spoken) narratives, which poses hitherto unaddressed challenges. As a test case, we address Holocaust survivor testimonies, given in English. Other than the importance of studying these testimonies for Holocaust research, we argue that they provide an interesting test case for topical segmentation, due to their unstructured surface level, relative abundance (tens of thousands of such testimonies were collected), and the relatively confined domain that they cover. We hypothesize that boundary points between segments correspond to low mutual information between the sentences proceeding and following the boundary. Based on this hypothesis, we explore a range of algorithmic approaches to the task, building on previous work on segmentation that uses generative Bayesian modeling and state-of-the-art neural machinery. Compared to manually annotated references, we find that the developed approaches show considerable improvements over previous work.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
09/29/2015

Automatically Segmenting Oral History Transcripts

Dividing oral histories into topically coherent segments can make them m...
research
03/25/2018

Text Segmentation as a Supervised Learning Task

Text segmentation, the task of dividing a document into contiguous segme...
research
04/15/2021

Neural Sequence Segmentation as Determining the Leftmost Segments

Prior methods to text segmentation are mostly at token level. Despite th...
research
08/28/2019

Onto Word Segmentation of the Complete Tang Poems

We aim at segmenting words in the Complete Tang Poems (CTP). Although it...
research
10/24/2022

Don't Discard Fixed-Window Audio Segmentation in Speech-to-Text Translation

For real-life applications, it is crucial that end-to-end spoken languag...
research
06/16/2022

Simultaneous Bone and Shadow Segmentation Network using Task Correspondence Consistency

Segmenting both bone surface and the corresponding acoustic shadow are f...
research
05/19/2023

Unsupervised Scientific Abstract Segmentation with Normalized Mutual Information

The abstracts of scientific papers consist of premises and conclusions. ...

Please sign up or login with your details

Forgot password? Click here to reset