An improved Bayesian TRIE based model for SMS text normalization

08/04/2020
by   Abhinava Sikdar, et al.
0

Normalization of SMS text, commonly known as texting language, is being pursued for more than a decade. A probabilistic approach based on the Trie data structure was proposed in literature which was found to be better performing than HMM based approaches proposed earlier in predicting the correct alternative for an out-of-lexicon word. However, success of the Trie based approach depends largely on how correctly the underlying probabilities of word occurrences are estimated. In this work we propose a structural modification to the existing Trie-based model along with a novel training algorithm and probability generation scheme. We prove two theorems on statistical properties of the proposed Trie and use them to claim that is an unbiased and consistent estimator of the occurrence probabilities of the words. We further fuse our model into the paradigm of noisy channel based error correction and provide a heuristic to go beyond a Damerau Levenshtein distance of one. We also run simulations to support our claims and show superiority of the proposed scheme over previous works.

READ FULL TEXT
research
09/17/2023

A novel approach to measuring patent claim scope based on probabilities obtained from (large) language models

This work proposes to measure the scope of a patent claim as the recipro...
research
06/02/2021

Evidence-based Factual Error Correction

This paper introduces the task of factual error correction: performing e...
research
10/04/2021

Error Correction for FrodoKEM Using the Gosset Lattice

We consider FrodoKEM, a lattice-based cryptosystem based on LWE, and pro...
research
04/30/2018

Inherent Biases in Reference-based Evaluation for Grammatical Error Correction and Text Simplification

The prevalent use of too few references for evaluating text-to-text gene...
research
02/17/2023

Deep Joint Source-Channel Coding with Iterative Source Error Correction

In this paper, we propose an iterative source error correction (ISEC) de...

Please sign up or login with your details

Forgot password? Click here to reset