ICDAR 2021 Competition on Scientific Literature Parsing

06/08/2021
by   Antonio Jimeno Yepes, et al.
0

Scientific literature contain important information related to cutting-edge innovations in diverse domains. Advances in natural language processing have been driving the fast development in automated information extraction from scientific literature. However, scientific literature is often available in unstructured PDF format. While PDF is great for preserving basic visual elements, such as characters, lines, shapes, etc., on a canvas for presentation to humans, automatic processing of the PDF format by machines presents many challenges. With over 2.5 trillion PDF documents in existence, these issues are prevalent in many other important application domains as well. Our ICDAR 2021 Scientific Literature Parsing Competition (ICDAR2021-SLP) aims to drive the advances specifically in document understanding. ICDAR2021-SLP leverages the PubLayNet and PubTabNet datasets, which provide hundreds of thousands of training and evaluation examples. In Task A, Document Layout Recognition, submissions with the highest performance combine object detection and specialised solutions for the different categories. In Task B, Table Recognition, top submissions rely on methods to identify table components and post-processing methods to generate the table structure and content. Results from both tasks show an impressive performance and opens the possibility for high performance practical applications.

READ FULL TEXT
research
05/05/2021

PingAn-VCGroup's Solution for ICDAR 2021 Competition on Scientific Literature Parsing Task B: Table Recognition to HTML

This paper presents our solution for ICDAR 2021 competition on scientifi...
research
05/30/2021

ICDAR 2021 Competition on Scientific Table Image Recognition to LaTeX

Tables present important information concisely in many scientific docume...
research
10/31/2022

Tables to LaTeX: structure and content extraction from scientific tables

Scientific documents contain tables that list important information in a...
research
11/16/2022

ChartParser: Automatic Chart Parsing for Print-Impaired

Infographics are often an integral component of scientific documents for...
research
11/15/2022

Deep learning for table detection and structure recognition: A survey

Tables are everywhere, from scientific journals, papers, websites, and n...
research
05/05/2021

PingAn-VCGroup's Solution for ICDAR 2021 Competition on Scientific Table Image Recognition to Latex

This paper presents our solution for the ICDAR 2021 Competition on Scien...
research
03/16/2023

Grab What You Need: Rethinking Complex Table Structure Recognition with Flexible Components Deliberation

Recently, Table Structure Recognition (TSR) task, aiming at identifying ...

Please sign up or login with your details

Forgot password? Click here to reset