Analyzing Bantu History with a Coalescent Theory Model: Horizontal Language Contact Erased Traces of Early Diversification

Background
Interdisciplinary collaboration among archaeology, genetics, and linguistics is essential when tracing human migration and the spread of civilizations. The expansion of the Bantu language family, encompassing the region south of the Sahara in Africa, is considered one of the largest language dispersal events in human history. Numerous scholars have endeavored to elucidate the early diversification process of this language family. Traditional phylogenetic approaches have primarily relied on tree-like analyses, borrowed from models of biological diversification. This methodology operates under the assumption that languages, during their divergence, experienced little to no interaction.
However, unlike biological species, languages frequently undergo horizontal transmission, where vocabulary mixes due to contact with neighboring groups. The Bantu language family likely experienced extensive interaction and exchange of vocabulary among different groups throughout its history. Simple phylogenetic models that exclude this language contact can distort the true evolutionary trajectory of languages. Previous analyses treated language contact as mere noise, failing to provide a clear conclusion on how the Bantu language family actually diversified and spread thousands of years ago. It is crucial to fully incorporate the actual interaction, i.e., the mixing of vocabulary, into the computational model.
Key Findings
A multinational research team led by Dr. Patrícia Santos applied coalescent theory, originally developed in population genetics, to linguistics. The team proposes a novel mathematical computational model that treats language contact as a natural evolutionary process. This model incorporates the concepts of gene flow and mutation from genetics, representing them as lexical borrowing between languages and the rate of independent vocabulary change, respectively. To validate this model, the researchers employed an Approximate Bayesian Computation (ABC) approach, comparing it with actual Bantu vocabulary data.
The analysis yielded surprising results. It revealed that the currently available vocabulary data of the Bantu language family has largely lost its early historical information. The rate of vocabulary change and the intensity of language contact within the Bantu language family were significantly higher than expected. This rapid change and frequent contact completely obscured the original signals from the early stages of language diversification. Consequently, it is now impossible to reconstruct the early diversification history of the Bantu language family using only the currently available vocabulary dataset.
Significance and Prospects
This research serves as a serious warning for interdisciplinary studies aimed at reconstructing human migration routes. Previous archaeological and genetic studies have used linguistic phylogenetic analysis as a triangulation tool to validate their hypotheses. However, if the inherent lexical mixing in language data makes early phylogenetic reconstruction fundamentally impossible, then the hypotheses based on it are also questionable. The research team explicitly states that the linguistic diversification structure of the Bantu language family should not be hastily cited as definitive evidence for reconstructing prehistoric human history.
Future research should focus on refining the model by incorporating linguistic markers that are less susceptible to contact, such as grammatical structures or phonological features, in addition to vocabulary. Just as phylogenomic analysis in bioinformatics overcame the limitations of single-gene analysis, a multi-marker model should be developed in linguistics. The validated mathematical framework can be usefully applied not only in Africa but also in analyzing Native American languages and large language families in Eurasia. By acknowledging the limitations of linguistics, we can pave the way for more precise reconstructions of prehistory.
Proceedings of the National Academy of Sciences, Volume 123, Issue 31, August 2026. SignificanceWe build a coalescent-theoretic model for studying the early prehistory of the Bantu language family. Prior computational phylogenetic work on the Bantu uses models that assume that language contact is absent or negligible. In contrast, we ...
This modeling technique can be applied not only to the study of human history but also to improve the performance of natural language processing (NLP) artificial intelligence (AI) that handles multilingual translation. Modern large language models (LLMs) often make errors when evaluating the semantic similarity of words in multilingual environments with frequent historical contact. A promising alternative is to incorporate a horizontal contact calculation algorithm based on coalescent theory into the training process of LLMs. For example, in complex multilingual environments such as Swahili, which is a fusion of several Bantu languages and Arabic, the model can trace the origin and time of introduction of each vocabulary item, thereby minimizing translation errors. Furthermore, it can assist in predicting the vocabulary loss patterns of endangered minority languages and designing effective language restoration scenarios.