Sentence Object Notation: Multilingual sentence notation based on Wordnet

01/03/2018
by   Abdelkrime Aries, et al.
0

The representation of sentences is a very important task. It can be used as a way to exchange data inter-applications. One main characteristic, that a notation must have, is a minimal size and a representative form. This can reduce the transfer time, and hopefully the processing time as well. Usually, sentence representation is associated to the processed language. The grammar of this language affects how we represent the sentence. To avoid language-dependent notations, we have to come up with a new representation which don't use words, but their meanings. This can be done using a lexicon like wordnet, instead of words we use their synsets. As for syntactic relations, they have to be universal as much as possible. Our new notation is called STON "SenTences Object Notation", which somehow has similarities to JSON. It is meant to be minimal, representative and language-independent syntactic representation. Also, we want it to be readable and easy to be created. This simplifies developing simple automatic generators and creating test banks manually. Its benefit is to be used as a medium between different parts of applications like: text summarization, language translation, etc. The notation is based on 4 languages: Arabic, English, Franch and Japanese; and there are some cases where these languages don't agree on one representation. Also, given the diversity of grammatical structure of different world languages, this annotation may fail for some languages which allows more future improvements.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
01/09/2013

Syntactic Analysis Based on Morphological Characteristic Features of the Romanian Language

This paper refers to the syntactic analysis of phrases in Romanian, as a...
research
04/13/2017

Learning Joint Multilingual Sentence Representations with Neural Machine Translation

In this paper, we use the framework of neural machine translation to lea...
research
01/25/2023

Distilling Text into Circuits

This paper concerns the structure of meanings within natural language. E...
research
06/27/2022

Center-Embedding and Constituency in the Brain and a New Characterization of Context-Free Languages

A computational system implemented exclusively through the spiking of ne...
research
05/23/2023

Towards Massively Multi-domain Multilingual Readability Assessment

We present ReadMe++, a massively multi-domain multilingual dataset for a...
research
01/18/2018

Natural Language Multitasking: Analyzing and Improving Syntactic Saliency of Hidden Representations

We train multi-task autoencoders on linguistic tasks and analyze the lea...
research
02/20/2020

Contextual Lensing of Universal Sentence Representations

What makes a universal sentence encoder universal? The notion of a gener...

Please sign up or login with your details

Forgot password? Click here to reset