Title Corpus-Based study of lexical bundles in human-produced and artificial intelligence generated written and spoken english academic texts
Translation of Title Tekstynais paremtas leksinių samplaikų tyrimas žmogaus ir dirbtinio intelekto sukurtuose anglų kalbos rašytiniuose ir sakytiniuose akademiniuose tekstuose.
Authors Rudžionytė, Eva
Full Text Download
Pages 40
Keywords [eng] leksinės samplaikos, diskurso funkcijos, tekstynų lingvistika, akademinis diskursas, DI sugeneruoti tekstai, ChatGPT, lexical bundles, discourse functions, corpus linguistics, academic discourse, AI-generated texts
Abstract [eng] Lexical bundles have been explored in a wide range of linguistic studies, addressing questions of their frequency, structure, and discourse functions. However, the recent emergence of large lan-guage models (LLMs) such as ChatGPT, which was created by artificial intelligence research company OpenAI, has broadened the scope of linguistic research, providing opportunities for comparative analyses of artificial intelligence (AI)-generated texts and human-produced language. Although a few studies investigating lexical bundles in ChatGPT (versions 3.5; 4.0) generated texts exist, a more recent ChatGPT version (e.g., GPT-5.2) remains mostly unexplored. This study analyses and compares the use of lexical bundles and their discourse functions between ChatGPT-5.2-generated texts and human-produced academic texts of written registers and transcriptions of lectures found in spoken registers. The aim of the study is to determine whether there are any sim-ilarities or differences in the use of lexical bundles and their discourse functions between ChatGPT-generated and human-produced academic texts. The research material for the study includes the British Academic Written English (BAWE) cor-pus as well as the British Academic Spoken English (BASE) corpus, which have been selected to represent human-produced academic texts. To represent AI-generated texts, two prompts were created for ChatGPT-5.2 to generate 200 undergraduate-level essays and 200 lecture transcrip-tions. All generated texts were compiled into two different corpora: ChatGPT Written and ChatGPT Spoken. The study, therefore, compares lexical bundle usage across four corpora: BAWE, BASE, ChatGPT Written, and ChatGPT Spoken. All corpora were analysed using the cor-pus analysis platform Sketch Engine, and four-word lexical bundles were extracted using the n-gram function. A combination of quantitative and qualitative analyses was conducted to determine similarities and differences in lexical bundle use across the corpora. The extracted lexical bundles were coded according to the taxonomy of discourse functions developed by Biber et al. (2004). The findings of the study indicate that ChatGPT-generated texts generally contain a higher num-ber of lexical bundles per essay compared to human-written essays. In terms of functional distri-bution, no statistically significant difference was found between BAWE and ChatGPT Written texts, suggesting that overall functional patterns may be more closely related to academic register. However, qualitative analysis shows that ChatGPT tends to rely more heavily on discourse organ-isers, whereas human-produced texts display a more balanced distribution of lexical bundle func-tions. Furthermore, the results suggest that ChatGPT-generated texts align more closely with the written academic register than the spoken academic register, as lecture transcriptions generated by ChatGPT reproduce patterns more typical of written academic discourse.
Dissertation Institution Vilniaus universitetas.
Type Master thesis
Language English
Publication date 2026