A Joint Probabilistic Classification Model of Relevant and Irrelevant Sentences in Mathematical Word Problems

11/21/2014
by   Suleyman Cetintas, et al.
0

Estimating the difficulty level of math word problems is an important task for many educational applications. Identification of relevant and irrelevant sentences in math word problems is an important step for calculating the difficulty levels of such problems. This paper addresses a novel application of text categorization to identify two types of sentences in mathematical word problems, namely relevant and irrelevant sentences. A novel joint probabilistic classification model is proposed to estimate the joint probability of classification decisions for all sentences of a math word problem by utilizing the correlation among all sentences along with the correlation between the question sentence and other sentences, and sentence text. The proposed model is compared with i) a SVM classifier which makes independent classification decisions for individual sentences by only using the sentence text and ii) a novel SVM classifier that considers the correlation between the question sentence and other sentences along with the sentence text. An extensive set of experiments demonstrates the effectiveness of the joint probabilistic classification model for identifying relevant and irrelevant sentences as well as the novel SVM classifier that utilizes the correlation between the question sentence and other sentences. Furthermore, empirical results and analysis show that i) it is highly beneficial not to remove stopwords and ii) utilizing part of speech tagging does not make a significant improvement although it has been shown to be effective for the related task of math word problem type classification.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
03/14/2021

Learning a Word-Level Language Model with Sentence-Level Noise Contrastive Estimation for Contextual Sentence Probability Estimation

Inferring the probability distribution of sentences or word sequences is...
research
05/16/2020

Learning Probabilistic Sentence Representations from Paraphrases

Probabilistic word embeddings have shown effectiveness in capturing noti...
research
08/09/2018

Arithmetic Word Problem Solver using Frame Identification

Automatic Word problem solving has always posed a great challenge for th...
research
06/13/2023

CipherSniffer: Classifying Cipher Types

Ciphers are a powerful tool for encrypting communication. There are many...
research
11/03/2018

Unsupervised Identification of Study Descriptors in Toxicology Research: An Experimental Study

Identifying and extracting data elements such as study descriptors in pu...
research
09/22/2021

A Simple Approach to Jointly Rank Passages and Select Relevant Sentences in the OBQA Context

In the open question answering (OBQA) task, how to select the relevant i...
research
10/31/2018

SURFACE: Semantically Rich Fact Validation with Explanations

Judging the veracity of a sentence making one or more claims is an impor...

Please sign up or login with your details

Forgot password? Click here to reset