The optimality of word lengths. Theoretical foundations and an empirical study

08/22/2022
by   Sonia Petrini, et al.
0

One of the most robust patterns found in human languages is Zipf's law of abbreviation, that is, the tendency of more frequent words to be shorter. Since Zipf's pioneering research, this law has been viewed as a manifestation of compression, i.e. the minimization of the length of forms - a universal principle of natural communication. Although the claim that languages are optimized has become trendy, attempts to measure the degree of optimization of languages have been rather scarce. Here we demonstrate that compression manifests itself in a wide sample of languages without exceptions, and independently of the unit of measurement. It is detectable for both word lengths in characters of written language as well as durations in time in spoken language. Moreover, to measure the degree of optimization, we derive a simple formula for a random baseline and present two scores that are dualy normalized, namely, they are normalized with respect to both the minimum and the random baseline. We analyze the theoretical and statistical pros and cons of these and other scores. Harnessing the best score, we quantify for the first time the degree of optimality of word lengths in languages. This indicates that languages are optimized to 62 or 67 percent on average (depending on the source) when word lengths are measured in characters, and to 65 percent on average when word lengths are measured in time. In general, spoken word durations are more optimized than written word lengths in characters. Beyond the analyses reported here, our work paves the way to measure the degree of optimality of the vocalizations or gestures of other species, and to compare them against written, spoken, or signed human languages.

READ FULL TEXT
research
03/17/2023

Direct and indirect evidence of compression of word lengths. Zipf's law of abbreviation revisited

Zipf's law of abbreviation, the tendency of more frequent words to be sh...
research
03/06/2017

Word forms - not just their lengths- are optimized for efficient communication

The inverse relationship between the length of a word and the frequency ...
research
07/29/2019

A Mathematical Model for Linguistic Universals

Inspired by chemical kinetics and neurobiology, we propose a mathematica...
research
09/12/2018

Multimodal neural pronunciation modeling for spoken languages with logographic origin

Graphemes of most languages encode pronunciation, though some are more e...
research
09/18/2021

Dependency distance minimization predicts compression

Dependency distance minimization (DDm) is a well-established principle o...
research
07/30/2020

The optimality of syntactic dependency distances

It is often stated that human languages, as other biological systems, ar...
research
05/02/2018

Robustness of sentence length measures in written texts

Hidden structural patterns in written texts have been subject of conside...

Please sign up or login with your details

Forgot password? Click here to reset