Original title:
Vývoj toolkitu pro red-teaming velkých jazykových modelů (LLM): návrh, implementace a vyhodnocení
Translated title:
Development of an LLM Red-Teaming Toolkit: Design, Implementation, and Evaluation
Authors:
Veselý, Adam ; Firc, Anton (referee) ; Reš, Jakub (advisor) Document type: Bachelor's theses
Year:
2026
Language:
eng Publisher:
Vysoké učení technické v Brně. Fakulta informačních technologií Abstract:
[eng][cze]
Sentinel je modulárny open-source toolkit založený na manifestoch na systematický red-teaming veľkých jazykových modelov. Architektúra je postavená na malej sade stabilných pluginových rozhraní pre generátory útokov, modelové adaptéry a sudcov, ktoré umožňujú nasadiť ľubovoľné vlastné implementácie bez zásahu do jadra, pričom referenčné ukážky pokrývajú lokálne aj API-kompatibilné backendy a hybridný hodnotiaci reťazec kombinujúci heuristických aj LLM sudcov. Experimenty sa opisujú deklaratívne pomocou manifestov YAML alebo JSON, čo umožňuje reprodukovateľné behy, troma vrstvami nastaviteľnej paralelizácie (na úrovni kombinácií, promptov a sudcov), obmedzenie rýchlosti a logovanie do JSONL. Implementácia zahŕňa aj offline analýzu, rozhranie v príkazovom riadku a deterministickú testovaciu sadu. Práca sumarizuje bezpečnostné riziká LLM, prehľad metód a nástrojov red-teamingu, navrhuje architektúru toolkitu, realizuje funkčný prototyp a overuje ho na lokálnych aj vzdialených modeloch. Výsledkom je ľahký framework, ktorý znižuje bariéru pre opakovateľné testovanie bezpečnosti LLM.
Sentinel is a modular, manifest-driven open-source toolkit for systematic red-teaming of large language models. The architecture is built around a small set of stable plugin interfaces for attack generators, model adapters, and judges that let a deployment supply its own implementations without touching the core, while reference examples cover local and API-compatible backends and a hybrid judging pipeline that combines verdicts from heuristic and LLM-based judges by weighted vote. Experiments are described declaratively in YAML or JSON manifests, enabling reproducible runs, three layers of configurable parallelism (combo, prompt and judge), rate limiting, and JSONL logging. The implementation also includes offline analysis of logs, a command-line interface, and a deterministic test suite. This thesis surveys LLM safety risks, reviews red-teaming methods and tools, proposes the toolkit architecture, implements a working prototype, and validates it on local and remote model configurations. The result is a lightweight framework that lowers the barrier to repeatable LLM safety evaluation.
Keywords:
bezpečnosť; jailbreaky; prompt injection; red-teaming; toolkit; veľké jazykové modely; viackolové útoky; jailbreaks; large language models; multi-turn attacks; prompt injection; red-teaming; safety; toolkit
Institution: Brno University of Technology
(web)
Document availability information: Fulltext is available in the Brno University of Technology Digital Library. Original record: http://hdl.handle.net/11012/259305