Improving Named Entity Recognition for Chinese Social Media with Word Segmentation Representation Learning

03/02/2016
by   Nanyun Peng, et al.
0

Named entity recognition, and other information extraction tasks, frequently use linguistic features such as part of speech tags or chunkings. For languages where word boundaries are not readily identified in text, word segmentation is a key first step to generating features for an NER system. While using word boundary tags as features are helpful, the signals that aid in identifying these boundaries may provide richer information for an NER system. New state-of-the-art word segmentation systems use neural models to learn representations for predicting word boundaries. We show that these same representations, jointly trained with an NER system, yield significant improvements in NER for Chinese social media. In our experiments, jointly training NER and word segmentation with an LSTM-CRF model yields nearly 5 absolute improvement over previously published results.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
02/27/2020

Integrating Boundary Assembling into a DNN Framework for Named Entity Recognition in Chinese Social Media Text

Named entity recognition is a challenging task in Natural Language Proce...
research
04/14/2020

Incorporating Uncertain Segmentation Information into Chinese NER for Social Media Text

Chinese word segmentation is necessary to provide word-level information...
research
11/14/2016

F-Score Driven Max Margin Neural Network for Named Entity Recognition in Chinese Social Media

We focus on named entity recognition (NER) for Chinese social media. Wit...
research
10/20/2018

Named Entity Recognition on Twitter for Turkish using Semi-supervised Learning with Word Embeddings

Recently, due to the increasing popularity of social media, the necessit...
research
09/22/2019

Using Chinese Glyphs for Named Entity Recognition

Most Named Entity Recognition (NER) systems use additional features like...
research
10/24/2015

Combine CRF and MMSEG to Boost Chinese Word Segmentation in Social Media

In this paper, we propose a joint algorithm for the word segmentation on...
research
09/29/2019

Gated Task Interaction Framework for Multi-task Sequence Tagging

Recent studies have shown that neural models can achieve high performanc...

Please sign up or login with your details

Forgot password? Click here to reset