Original title:
Monitorování a vizualizace metrik výkonu a zdrojů v HPC/AI klastrech
Translated title:
Monitoring and Visualization of Performance and Resource Metrics in HPC/AI Clusters
Authors:
Poláček, Stanislav ; Šrejber, Martin (referee) ; Ondrák, Viktor (advisor) Document type: Bachelor's theses
Year:
2026
Language:
cze Publisher:
Vysoké učení technické v Brně. Fakulta podnikatelská Abstract:
[cze][eng]
Tato bakalářská práce se zabývá návrhem a implementací moderního monitorovacího systému pro výpočetní klastr. Analyzuje současnou hardwarovou i síťovou infrastrukturu s ohledem na požadavky sledování systémových prostředků. Poskytuje teoretický přehled dostupných technologií pro sběr a ukládání metrik do databází časových řad. Hlavní část práce je zaměřena na nasazení platformy Prometheus, integraci s plánovačem úloh a vytvoření vizualizačních panelů v systému Grafana pro celkové zefektivnění správy daného klastru.
This bachelor's thesis deals with the design and implementation of a modern monitoring system for a computing cluster. It analyzes the current hardware and network infrastructure with regard to the requirements for monitoring system resources. It provides a theoretical overview of available technologies for collecting and storing metrics in time-series databases. The main part of the thesis focuses on the deployment of the Prometheus platform, integration with the job scheduler, and the creation of visualization dashboards in the Grafana system to streamline the overall management of the cluster.
Keywords:
Computer cluster; Containerization; Docker; Grafana; HPC; Monitoring system; Prometheus; Sun Grid Engine; Docker; Grafana; HPC; Kontejnerizace; Monitorovací systém; Počítačový klastr; Prometheus; Sun Grid Engine
Institution: Brno University of Technology
(web)
Document availability information: Fulltext is available in the Brno University of Technology Digital Library. Original record: http://hdl.handle.net/11012/259348