A Supervised Learning Approach For Heading Detection

08/31/2018
by   Sahib Singh Budhiraja, et al.
0

As the Portable Document Format (PDF) file format increases in popularity, research in analysing its structure for text extraction and analysis is necessary. Detecting headings can be a crucial component of classifying and extracting meaningful data. This research involves training a supervised learning model to detect headings with features carefully selected through recursive feature elimination. The best performing classifier had an accuracy of 96.95 heading detection contributes to the field of PDF based text extraction and can be applied to the automation of large scale PDF text analysis in a variety of professional and policy based contexts.

READ FULL TEXT

Please sign up or login with your details

Forgot password? Click here to reset