The shrinking landscape of linguistic diversity in the age of large language models
Pangram verdict · v3.3
We believe that this entire text is human-written.
AI likelihood · overall
HumanArticle text · 1,367 words · 1 segments analyzed
ReferencesOrwell, G. Nineteen Eighty-Four 69–70 (Penguin Books, 1949).Park, G. et al. Automatic personality assessment through social media language. J. Pers. Soc. Psychol. 108, 934–952 (2015).Article PubMed Google Scholar Oberlander, J. & Gill, A. J. Language with character: a stratified corpus comparison of individual differences in e-mail communication. Discourse Process. 42, 239–270 (2006).Article Google Scholar Moreno, J. D., Martinez-Huertas, J. A., Olmos, R., Jorge-Botana, G. & Botella, J. Can personality traits be measured analyzing written language? A meta-analytic study on computational methods. Pers. Individ. Dif. 177, 110818 (2021).Article Google Scholar Mairesse, F., Walker, M. A., Mehl, M. R. & Moore, R. K. Using linguistic cues for the automatic recognition of personality in conversation and text. J. Artif. Intell. Res. 30, 457–500 (2007).Article Google Scholar Schwartz, H. A. et al. Personality, gender, and age in the language of social media: the open-vocabulary approach. PLoS ONE 8, 73791 (2013).Article Google Scholar Kramsch, C. Language and culture. AILA Rev. 27, 30–55 (2014).Article Google Scholar Gumperz, J. The speech community. Int. Encycl. Soc. Sci. 9, 381–386 (1968). Google Scholar Nguyen, D. & Rosé, C. P. Language use as a reflection of socialization in online communities. In Proc. Workshop on Language in Social Media (eds Nagarajan, M. & Gamon, M.) 76–85 (Association for Computational Linguistics, 2011).Bamman, D., Eisenstein, J. & Schnoebelen, T. Gender identity and lexical variation in social media. J. Socioling. 18, 135–160 (2014).Article Google Scholar Pennebaker, J. W. The secret life of pronouns. New Sci. 211, 42–45 (2011).Article Google Scholar Robinson, M. D., Boyd, R. L., Fetterman, A. K. & Persich, M. R. The mind versus the body in political (and nonpolitical) discourse: linguistic evidence for an ideological signature in US politics. J. Lang. Soc. Psychol. 36, 438–461 (2017).Article Google Scholar Huang, Y., Guo, D., Kasakoff, A. & Grieve, J. Understanding US regional linguistic variation with Twitter data analysis. Comput. Environ. Urban Syst. 59, 244–255 (2016).Article Google Scholar Eisenstein, J., O’Connor, B., Smith, N. A. & Xing, E. A latent variable model for geographic lexical variation. In Proc. 2010 Conference on Empirical Methods in Natural Language Processing (eds Li, H. & Màrquez, L.) 1277–1287 (Association for Computational Linguistics, 2010).Peterson, K., Hohensee, M. & Xia, F. Email formality in the workplace: a case study on the enron corpus. In Proc. Workshop on Language in Social Media (LSM 2011) (eds Nagarajan, M. & Gamon, M.) 86–95 (Association for Computational Linguistics, 2011).Stamatatos, E. A survey of modern authorship attribution methods. J. Assoc. Inf. Sci. Technol. 60, 538–556 (2009).Article Google Scholar Grieve, J. Quantitative authorship attribution: an evaluation of techniques. Lit. Ling. Comput. 22, 251–270 (2007).Article Google Scholar Cassell, J. & Tversky, D. The language of online intercultural community formation. J. Comput. Mediat. Commun. 10, 1027 (2005). Google Scholar Danet, B. & Herring, S. C. Introduction: the multilingual internet. J. Comput. Mediat. Commun. 9, 9110 (2003). Google Scholar Kennedy, B. et al. Moral concerns are differentially observable in language. Cognition 212, 104696 (2021).Article PubMed Google Scholar Jackson, J. C., Gelfand, M., De, S. & Fox, A. The loosening of American culture over 200 years is associated with a creativity–order trade-off. Nat. Hum. Behav. 3, 244–250 (2019).Article PubMed Google Scholar Hofmann, V., Kalluri, P. R., Jurafsky, D. & King, S. AI generates covertly racist decisions about people based on their dialect. Nature 633, 147–154 (2024).Article CAS PubMed PubMed Central Google Scholar Richard, A. B., Lelandais, M., Reilly, K. T. & Jacquin-Courtois, S. Linguistic markers of subtle cognitive impairment in connected speech: a systematic review. J. Speech Lang. Hear. Res. 67, 4714–4733 (2024).Article PubMed Google Scholar Eyigoz, E., Mathur, S., Santamaria, M., Cecchi, G. & Naylor, M. Linguistic markers predict onset of Alzheimer’s disease. EClinicalMedicine 28, 100583 (2020).Roark, B., Mitchell, M., Hosom, J.-P., Hollingshead, K. & Kaye, J. Spoken language derived measures for detecting mild cognitive impairment. IEEE Trans. Audio Speech Lang. Process. 19, 2081–2090 (2011).Article PubMed PubMed Central Google Scholar Trifu, R. N. et al. Linguistic markers for major depressive disorder: a cross-sectional study using an automated procedure. Front. Psychol. 15, 1355734 (2024).Article PubMed PubMed Central Google Scholar Weerasinghe, J., Morales, K. & Greenstadt, R. "Because… i was told… so much”: linguistic indicators of mental health status on Twitter. Proc. Priv. Enhanc. Technol. 4, 152–171 (2019).Whorf, B. L. Language, Thought, and Reality: Selected Writings of Benjamin Lee Whorf (MIT Press, 2012).Eckert, P. Three waves of variation study: the emergence of meaning in the study of sociolinguistic variation. Annu. Rev. Anthropol. 41, 87–100 (2012).Article Google Scholar Hofstede, G. Culture’s Consequences: Comparing Values, Behaviors, Institutions and Organizations Across Nations 2nd edn (Sage, 2001).OpenAI. Introducing ChatGPT https://openai.com/blog/chatgpt (2022).Gemini Team et al. Gemini: a family of highly capable multimodal models. Preprint at https://doi.org/10.48550/arXiv.2312.11805 (2023).Bailyn, E. ChatGPT Usage Statistics: March 2026 https://firstpagesage.com/seo-blog/chatgpt-usage-statistics/ (FirstPageSage, 2026).Nearly 1 in 3 College Students Have Used ChatGPT on Written Assignments https://www.intelligent.com/nearly-1-in-3-college-students-have-used-chatgpt-on-written-assignments/ (Intelligent, 2024).McClain, C. Americans’ Use of ChatGPT is Ticking Up, But Few Trust Its Election Information https://www.pewresearch.org/short-reads/2024/03/26/americans-use-of-chatgpt-is-ticking-up-but-few-trust-its-election-information/ (Pew Research Center, 2024).Handa, K. et al. Which economic tasks are performed with AI? Evidence from millions of Claude conversations. Preprint at https://doi.org/10.48550/arXiv.2503.04761 (2025).Mizrahi, M. et al. State of what art? A call for multi-prompt LLM evaluation. Trans. Assoc. Comput. Linguist. 12, 933–949 (2024).Article Google Scholar Serapio-García, G. et al. A psychometric framework for evaluating and shaping personality traits in large language models. Nat. Mach. Intell. 7, 1954–1968 (2025).Article PubMed PubMed Central Google Scholar Ghosh, S. et al. A closer look at the limitations of instruction tuning. In Proc. 41st International Conference on Machine Learning 624 (JMLR, 2024).Santurkar, S. et al. Whose opinions do language models reflect? In International Conference on Machine Learning 29971–30004 (JMLR, 2023).Ireland, M.E. & Mehl, M. R. in The Oxford Handbook of Language and Social Psychology (ed. Holtgraves, T. M.) 201–218 https://doi.org/10.1093/oxfordhb/9780199838639.013.034 (Oxford Univ. Press, 2014).Corona Hernández, H. et al. Natural language processing markers for psychosis and other psychiatric disorders: emerging themes and research agenda from a cross-linguistic workshop. Schizophr. Bull. 49, 86–92 (2023).Article Google Scholar Rude, S., Gortner, E.-M. & Pennebaker, J. Language use of depressed and depression-vulnerable college students. Cogn. Emot. 18, 1121–1133 (2004).Article Google Scholar Coppersmith, G., Leary, R., Crutchley, P. & Fine, A. Natural language processing of social media as screening for suicide risk. Biomed. Inform. Insights 10, 1178222618792860 (2018).Article PubMed PubMed Central Google Scholar Matz, S. C. & Netzer, O. Using big data as a window into consumers’ psychology. Curr. Opin. Behav. Sci. 18, 7–12 (2017).Article Google Scholar Winter, S., Maslowska, E. & Vos, A. L. The effects of trait-based personalization in social media advertising. Comput. Hum. Behav. 114, 106525 (2021).Article Google Scholar Sundar, S. S. & Marathe, S. S. Personalization versus customization: the importance of agency, privacy, and power usage. Hum. Commun. Res. 36, 298–322 (2010).Article Google Scholar Ryan, M. J., Held, W. & Yang, D. Unintended impacts of LLM alignment on global representation. In Proc. 62nd Annual Meeting of the Association for Computational Linguistics1, 16121–16140 (2024).Navigli, R., Conia, S. & Ross, B. Biases in large language models: origins, inventory, and discussion. ACMJ. Data Inf. Qual. 15, 1–21 (2023).Article Google Scholar Atari, M., Xue, M. J., Park, P. S., Blasi, D. & Henrich, J. Which humans? Preprint at PsyArXiv https://doi.org/10.31234/osf.io/5b26t (2023).Wang, A., Morgenstern, J. & Dickerson, J. P. Large language models that replace human participants can harmfully misportray and flatten identity groups. Nat. Mach. Intell. 7, 400–411 (2025).Article Google Scholar Rozado, D. The political preferences of LLMs. PLoS ONE 19, 0306621 (2024).Article Google Scholar Pan, K. & Zeng, Y. Do LLMs possess a personality? Making the MBTI test an amazing evaluation for large language models. Preprint at https://doi.org/10.48550/arXiv.2307.16180 (2023).Abdurahman, S. et al. Perils and opportunities in using large language models in psychological research. PNAS Nexus 3, 245 (2024).Article Google Scholar Kobak, D., González-Márquez, R., Horvát, E. -Á & Lause, J. Delving into LLM-assisted writing in biomedical publications through excess vocabulary. Sci. Adv. 11, 3813 (2025).Article Google Scholar Liang, W. et al. Quantifying large language model usage in scientific papers. Nat. Hum. Behav. 9, 2599–2609 (2025).Article PubMed Google Scholar Liang, W. et al. The widespread adoption of large language model-assisted writing across society. Patterns 6, 101366 (2025).Bao, T., Zhao, Y., Mao, J. & Zhang, C. Examining linguistic shifts in academic writing before and after the launch of chatGPT: a study on preprint papers. Scientometrics 130, 3597–3627 (2025).Article Google Scholar