Original title:
Generování strukturovaných testovacích dat pomocí LLM
Authors:
Hajdúková, Jasmína ; Janoušek, Vladimír (referee) ; Smrčka, Aleš (advisor) Document type: Bachelor's theses
Year:
2026
Language:
slo Publisher:
Vysoké učení technické v Brně. Fakulta informačních technologií Abstract:
[slo][eng]
Práca sa zameriava na návrh a implementáciu nástroja pre automatické generovanie YAML konfigurácie (receptov) pre generátor testovacích dát Dataster. Manuálne vytváranie takýchto receptov je časovo namáhavé a vyžaduje rozsiahle znalosti o schéme databázy, vzťahoch medzi tabuľkami a charaktere hodnôt. Pomocou navrhnutého nástroja je možné plne automatizovať tento proces, pričom sa nástroj pripojí k relačnej databáze z ktorej získa schému spolu so vzorkami dát. Následne s pomocou lokálneho veľkého jazykového modelu (LLM) cez platformu Ollama vygeneruje recept. Samotné generovanie je rozložené do štyroch fáz, ktoré sú doplnené o validačnú a opravnú vrstvu. Vyhodnotenie úspešnosti je vykonané na troch úrovniach, pri modeli qwen2.5 na sade 31 testovacích databáz pri šiestich teplotách. Pri modeloch llama3.1 a mistral sú spúšťané testy nad 19 databázami pri jednej teplote (0.0). Výsledky ukazujú na to, že zmiešaná konfigurácia modelov qwen2.5 dosahuje kvalitu porovnateľnú s väčším modelom.
This thesis focuses on the design and implementation of a tool for the automatic generation of YAML configurations (recipes) for the Dataster test data generator. Manually creating such recipes is time-consuming and requires extensive knowledge of the database schema, relationships between tables, and the nature of the values. Using the proposed tool, this process can be fully automated; the tool connects to a relational database from which it retrieves the schema along with sample data. It then generates a recipe using a large language model (LLM) via the Ollama platform. The generation process itself is divided into four phases, supplemented by a validation and correction layer. Performance evaluation is conducted at three levels using the qwen2.5 model on a set of 31 test databases at six temperatures. For the llama3.1 and mistral models, tests are run on 19 databases at a single temperature (0.0). The results indicate that a mixed configuration of qwen2.5 models achieves quality comparable to that of a larger model.
Keywords:
data-driven testing; Dataster; large language models; LLM; Ollama; relational databases; software testing; test data generation; YAML
Institution: Brno University of Technology
(web)
Document availability information: Fulltext is available in the Brno University of Technology Digital Library. Original record: http://hdl.handle.net/11012/260633