Original title:
Efektivní vícevektorové vyhledávání dokumentů s adaptivní velikostí jejich reprezentace
Translated title:
Efficient Multi-Vector Document Retrieval with Adaptive Representation Size
Authors:
Štětina, Jakub ; Beneš, Karel (referee) ; Fajčík, Martin (advisor) Document type: Bachelor's theses
Year:
2025
Language:
eng Publisher:
Vysoké učení technické v Brně. Fakulta informačních technologií Abstract:
[eng][cze]
Tato práce se zabývá problémy škálovatelnosti multivektorových modelů z oblasti information retrieval (IR), konkrétně se zaměřuje na omezení velikosti indexů tvořených modelem ColBERTv2 PLAID. K překonání těchto limitací je zavedena nová sparsifikační metoda---ColBERT-Sparse---která adaptivně volí velikost reprezentace každého indexovaného dokumentu z korpusu textu s minimálními degradací kvality vyhledávání. Jsou zkoumány dva mechanismy sparsifikace modelu: mechanismus se sigmoid vrstvou s prahováním a mechanismus s gumbel-softmax vrstvou založený na autonomním výběru modelem, oba integrované do původní architektury ColBERTv2. Experimenty jsou prováděny s dvěma základními base modely a dvěma návrhy architektury. Navrhovaná metoda dosahuje přes 80\,\% redukce velikosti indexu při zachování kvality vyhledávání porovnatelně s původním modelem. Další analýzy se zaměřují na naučené způsoby sparsifikace trénovanými modely s ohledem na délky dokumentů nebo zachování jednotlivých slovních druhů. Tato práce demosntruje úspěšnost navrhovaného řešení adaptivní sparsifikace textových reprezentací v dense IR modelech a navrhuje směry pro další zlepšení.
This thesis addresses the scalability challenges of multi-vector information retrieval (IR) systems, specifically focusing on the memory footprint limitations posed by the ColBERTv2 PLAID. To overcome this limitation, a new sparsification method---ColBERT-Sparse---is introduced, which adaptively chooses the representation size of each indexed document with a minimal sacrifice in retrieval quality. Two sparsity mechanisms are explored: a deterministic Sigmoid-based gating mechanism and a stochastic Gumbel-Softmax-based selection layer, both integrated into the original ColBERTv2 architecture. Experiments are conducted across multiple base models and architecture designs, with hyperparameter tuning and threshold selection. The proposed method achieves over 80\,\% index size reduction while maintaining near-original retrieval performance, with retrieval effectiveness degrading by only approximately 2 percentage points in Recall@10 compared to the original ColBERT baseline. Additional analyses reveal structured sparsity patterns across document lengths and part-of-speech categories. This work demonstrates the feasibility of end-to-end trainable adaptive sparsification in dense IR models and highlights directions for improving this approach in future retrieval systems.
Keywords:
BERT; ColBERT; Dense Retrieval; Information Retrieval; Komprese indexu; L1 regularizace; Late Interaction; Sparsity; Token Pruning; BERT; ColBERT; Dense Retrieval; Index Compression; Information Retrieval; L1 Regularization; Late Interaction; Sparsity; Token Pruning
Institution: Brno University of Technology
(web)
Document availability information: Fulltext is available in the Brno University of Technology Digital Library. Original record: http://hdl.handle.net/11012/253723