2026
Personalized kNN Query Execution in Vector Databases
ŠIKYŇA, Matúš a Pavel ZEZULAZákladní údaje
Originální název
Personalized kNN Query Execution in Vector Databases
Autoři
Vydání
New York, NY, USA, Proceedings of the 35th ACM International Conference on Information and Knowledge Management (CIKM '26), 11 s. 2026
Nakladatel
Association for Computing Machinery
Další údaje
Jazyk
angličtina
Typ výsledku
Stať ve sborníku
Obor
10200 1.2 Computer and information sciences
Stát vydavatele
Spojené státy
Utajení
není předmětem státního či obchodního tajemství
Forma vydání
elektronická verze "online"
Označené pro přenos do RIV
Ne
Organizační jednotka
Fakulta informatiky
Klíčová slova anglicky
Personalized Search; Metric Learning; Vector Databases
Štítky
Příznaky
Mezinárodní význam, Recenzováno
Změněno: 15. 9. 2026 06:50, RNDr. Mgr. Matúš Šikyňa
Anotace
V originále
Modern vector databases enable search over high-dimensional embeddings, but query execution engines rely on a specific similarity metric that cannot reflect how individual users perceive similarity. We advocate a metric-level perspective on personalization, keeping the shared embedding space and the underlying index unchanged, and representing each user's specific search needs by a small matrix for Mahalanobis distance learned using relevance feedback and metric learning. Using a lower bounding relationship between Mahalanobis and Euclidean distances, we propose a query execution strategy for personalized k-nearest neighbour search that retrieves an enlarged candidate set using a standard Euclidean (or inner product) index and refines it under the user’s subjective metric. Experiments on CLIP embeddings from Profiset and LAION-2B-en datasets using FAISS-based indices show that this approach can be layered on top of existing vector databases with predictable behaviour, high overlap with precise Mahalanobis kNN search results, and only a small overhead over standard Euclidean search. When the Euclidean–Mahalanobis discrepancy becomes too large for efficient refinement, we propose approximation techniques and show that a Cholesky-based whitening transformation can reorganize embeddings so that personalized search reduces to Euclidean indexing again.
Návaznosti
| MUNI/A/1860/2025, interní kód MU |
| ||
| MUNI/A/1873/2025, interní kód MU |
| ||
| VK01010147, projekt VaV |
|