XSTEM: An exemplar-based stemming algorithm

05/09/2022
by   Kirk Baker, et al.
0

Stemming is the process of reducing related words to a standard form by removing affixes from them. Existing algorithms vary with respect to their complexity, configurability, handling of unknown words, and ability to avoid under- and over-stemming. This paper presents a fast, simple, configurable, high-precision, high-recall stemming algorithm that combines the simplicity and performance of word-based lookup tables with the strong generalizability of rule-based methods to avert problems with out-of-vocabulary words.

READ FULL TEXT

Please sign up or login with your details

Forgot password? Click here to reset