Capturing Logical Structure of Visually Structured Documents with Multimodal Transition Parser

05/01/2021
by   Yuta Koreeda, et al.
0

While many NLP papers, tasks and pipelines assume raw, clean texts, many texts we encounter in the wild are not so clean, with many of them being visually structured documents (VSDs) such as PDFs. Conventional preprocessing tools for VSDs mainly focused on word segmentation and coarse layout analysis, while fine-grained logical structure analysis (such as identifying paragraph boundaries and their hierarchies) of VSDs is underexplored. To that end, we proposed to formulate the task as prediction of transition labels between text fragments that maps the fragments to a tree, and developed a feature-based machine learning system that fuses visual, textual and semantic cues. Our system significantly outperformed baselines in identifying different structures in VSDs. For example, our system obtained a paragraph boundary detection F1 score of 0.951 which is significantly better than a popular PDF-to-text tool with a F1 score of 0.739.

READ FULL TEXT
POST COMMENT

Comments

There are no comments yet.

Authors

page 1

page 2

page 3

page 4

11/30/2020

Floods Detection in Twitter Text and Images

In this paper, we present our methods for the MediaEval 2020 Flood Relat...
05/22/2020

Robust Layout-aware IE for Visually Rich Documents with Pre-trained Language Models

Many business documents processed in modern NLP and IR pipelines are vis...
11/09/2020

Chapter Captor: Text Segmentation in Novels

Books are typically segmented into chapters and sections, representing c...
10/08/2018

An AMR Aligner Tuned by Transition-based Parser

In this paper, we propose a new rich resource enhanced AMR aligner which...
09/06/2020

UPB at SemEval-2020 Task 8: Joint Textual and Visual Modeling in a Multi-Task Learning Architecture for Memotion Analysis

Users from the online environment can create different ways of expressin...
03/22/2021

Identifying Machine-Paraphrased Plagiarism

Employing paraphrasing tools to conceal plagiarized text is a severe thr...
07/12/2021

Lumen: A Machine Learning Framework to Expose Influence Cues in Text

Phishing and disinformation are popular social engineering attacks with ...
This week in AI

Get the week's most popular data science and artificial intelligence research sent straight to your inbox every Saturday.