Title Didelio kalbos modelio generuojamų kodo testų su struktūrine validacija kokybės vertinimo metodas
Translation of Title A method for quality evaluation of code tests generated by a large language model with structural validation.
Authors Zapkus, Dovydas Marius
Full Text Download
Pages 71
Abstract [eng] This paper examines the evaluation of the quality of unit tests generated by large language models, using a strict prompt structure and structural validation of the results. The literature review presents the principles of unit testing, metrics for evaluating the quality of unit tests, the operational characteristics of large language models and the processes of compilation, execution, and mutation testing of unit tests. The subsequent section of the paper presents a comparison of seven selected large language models based on an experiment using a C# dataset, using PydanticAI library for result validation, dotnet-coverage tool for code coverage assessment, and Stryker.NET tool for mutation testing. Following the experiment, the paper presents the compilation rates of generated tests, execution performance, code coverage, and mutation testing. The models are discussed with a focus on their ability to generate syntactically and semantically correct unit tests that are resistant to code faults. Based on the established comparison, the results are compared with studies conducted by other authors. The study reveals that the application of strict prompt structure and result validation can improve the compilation performance of generated unit tests, code coverage, and mutation coverage metrics. The grok-4 and gemini-2.5-flash models are evaluated as the most advanced models in the field of unit test generation.
Dissertation Institution Vilniaus universitetas.
Type Master thesis
Language Lithuanian
Publication date 2026