Education, Science, Technology, Innovation and Life
Open Access
Sign In

Characterization and Mechanisms of Lexical Complexity in AI-Generated Texts: A Comparative Corpus-Based Study

Download as PDF

DOI: 10.23977/langl.2026.090120 | Downloads: 2 | Views: 89

Author(s)

Lulu Chen 1

Affiliation(s)

1 School of International Education, Tianjin Foreign Studies University, Tianjin, 300270, China

Corresponding Author

Lulu Chen

ABSTRACT

Based on a corpus-based methodology, this study analyzes the intrinsic reasons for the high level of lexical complexity observed in Artificial Intelligence Generated Content (AIGC). The research compares 24 English argumentative essays written by AI with 24 second-language (L2) learner essays reaching the IELTS Writing Task 2 Band 7 level. Under controlled conditions of identical genre and topic, the study performs quantitative statistics across three dimensions: lexical sophistication, semantic abstraction, and information density. Statistical results indicate that the frequency of advanced vocabulary in AI texts is significantly higher, approximately 2.3 times that of human texts. The proportion of abstract nouns reached 9.14%, far exceeding the 3.32% found in human texts, suggesting that AI expressions tend toward nominalization and conceptualization. Regarding overall information organization, the lexical density of AI texts was 69.9%, also surpassing the 60.3% of human texts, reflecting a stronger tendency for information condensation and phrasal structures. The analysis points out that the complexity of AI text primarily stems from its mechanism of selecting vocabulary based on probability distributions. This mechanism favors longer words, abstract nouns, and words with high semantic content, thereby forming a highly compact linguistic surface. Such complexity is essentially a formal feature at the statistical level and is not entirely equivalent to the proficiency levels corresponding to human L2 acquisition. These findings provide empirical references for AI text identification, the refinement of writing evaluation standards, and L2 writing pedagogy.

KEYWORDS

AI-generated text, Corpus, L2 writing, Lexical complexity

CITE THIS PAPER

Lulu Chen. Characterization and Mechanisms of Lexical Complexity in AI-Generated Texts: A Comparative Corpus-Based Study. Lecture Notes on Language and Literature (2026). Vol. 9, No.1, 139-148. DOI: http://dx.doi.org/10.23977/langl.2026.090120.

REFERENCES

[1] X. Zhai (2023). ChatGPT for next generation science learning. XRDS: Crossroads, The ACM Magazine for Students, vol. 29, no. 3, p. 42-46.
[2] F. Jiang and X.L. Li (2026). A study on lexical richness and syntactic complexity of AI-generated texts: A comparison between ChatGPT and native college students’ writing. Foreign Languages, vol.49, no. 2, p. 43-51.
[3] J.H. Zhu, M.Y. Wang, E.H. Yang, J.R. Nie, L.E. Yang and Y.J. Wang (2024). A comparative study of linguistic features between large language model-generated responses and human responses. Journal of Chinese Information Processing, vol. 38, no. 4, p. 17-27.
[4] E. Kasneci, K. Sessler, S. Küchemann, M. Bannert, D. Dementieva, F. Fischer et al. (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences, vol. 103, p. 102274. 
[5] Z.H. Wang (2024). Research on identification of ChatGPT-generated texts based on semantic features. Master's thesis, Yanbian University.
[6] D. R. E. Cotton, P. A. Cotton and J. R. Shipway (2024). Chatting and cheating: Ensuring academic integrity in the era of ChatGPT. Innovations in Education and Teaching International, vol. 61, no. 2, p. 228-239. 
[7] C.A. Gao, F.M. Howard, N.S. Markov, E.C. Dyer, S. Wheeler, A. Leiter and A.T. Pearson (2023). Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence detector. bioRxiv. 
[8] A. Shah, P. Ranka, U. Dedhia, S. Prasad, S. Muni and K. Bhowmick (2023). Detecting and unmasking AI-generated texts through explainable artificial intelligence using stylistic features. International Journal of Advanced Computer Science and Applications, vol. 14, no. 10.
[9] S. Herbold, A. Hautli-Janisz, U. Heuer, Z. Kikteva and A. Trautsch (2023). A large-scale comparison of human-written versus ChatGPT-generated essays. Scientific Reports, vol. 13, no. 1, p. 18699. 
[10] A. Housen, F. Kuiken and I. Vedder (2012). Dimensions of L2 performance and proficiency: Complexity, accuracy and fluency in SLA. John Benjamins.
[11] X. Lu (2010). The relationship of lexical richness to the quality of ESL learners' oral narratives. The Modern Language Journal, vol. 94, no. 3, p. 474-496.
[12] P.M. McCarthy and S. Jarvis (2010). MTLD, vocd-D, and HD-D: A validation study of sophisticated approaches to lexical diversity assessment. Behavior Research Methods, vol. 42, no. 2, p. 381-392.
[13] T. McEnery and A. Hardie (2011). Corpus linguistics: Method, theory and practice. Cambridge University Press.
[14] D. Biber and B. Gray (2016). Grammatical complexity in academic English: Linguistic change in writing. Cambridge University Press.
[15] M.A.K. Halliday (1985). An introduction to functional grammar. Edward Arnold.
[16] C. Zhang, X.Y. Ma and Q.Y. Yan (2026). A comparative study on linguistic complexity between GenAI-generated texts and students' English argumentative essays. Shandong Foreign Language Teaching, vol. 47, no. 1, p. 77-86.
[17] X. Zhao (2025). A comparative analysis of AI-generated texts and PEP English textbook texts from the perspective of language and culture: A case study of senior high school reading materials. Master’s thesis, Beijing Foreign Studies University.
[18] T. Nkhobo and C. Chaka (2023). Student-written versus ChatGPT-generated discursive essays: A comparative Coh-Metrix analysis of lexical diversity, syntactic complexity, and referential cohesion. International Journal of Education and Development using Information and Communication Technology, vol. 19, no. 3, p. 69-84.

All published work is licensed under a Creative Commons Attribution 4.0 International License.

Copyright © 2016 - 2031 Clausius Scientific Press Inc. All Rights Reserved.