The empirical structure of word frequency distributions

01/09/2020
by   Michael Ramscar, et al.
0

The frequencies at which individual words occur across languages follow power law distributions, a pattern of findings known as Zipf's law. A vast literature argues over whether this serves to optimize the efficiency of human communication, however this claim is necessarily post hoc, and it has been suggested that Zipf's law may in fact describe mixtures of other distributions. From this perspective, recent findings that Sinosphere first (family) names are geometrically distributed are notable, because this is actually consistent with information theoretic predictions regarding optimal coding. First names form natural communicative distributions in most languages, and I show that when analyzed in relation to the communities in which they are used, first name distributions across a diverse set of languages are both geometric and, historically, remarkably similar, with power law distributions only emerging when empirical distributions are aggregated. I then show this pattern of findings replicates in communicative distributions of English nouns and verbs. These results indicate that if lexical distributions support efficient communication, they do so because their functional structures directly satisfy the constraints described by information theory, and not because of Zipf's law. Understanding the function of these information structures is likely to be key to explaining humankind's remarkable communicative capacities.

READ FULL TEXT
research
08/23/2022

Universality and diversity in word patterns

Words are fundamental linguistic units that connect thoughts and things ...
research
07/05/2018

Zipf's law in 50 languages: its structural pattern, linguistic interpretation, and cognitive motivation

Zipf's law has been found in many human-related fields, including langua...
research
12/08/2014

Optimization models of natural communication

A family of information theoretic models of communication was introduced...
research
05/14/2021

From Multisets over Distributions to Distributions over Multisets

A well-known challenge in the semantics of programming languages is how ...
research
06/09/2020

Re-evaluating phoneme frequencies

Causal processes can give rise to distinctive distributions in the lingu...
research
05/11/2019

Semantic categories of artifacts and animals reflect efficient coding

It has been argued that semantic categories across languages reflect pre...
research
11/04/2019

Global Regularity and Individual Variability in Dynamic Behaviors of Human Communication

A new model, called "Human Dynamics", has been recently proposed that in...

Please sign up or login with your details

Forgot password? Click here to reset