Optimizing Data Processing with Learned Models: Techniques for Indexing and Cardinality Estimation
- In this thesis, we study the problem of optimizing data processing in database systems using learned models. As data volumes increase and datasets become structurally complex, classical database components struggle to exploit underlying data distributions, resulting in suboptimal performance and excessive memory usage due to their rigid designs. Advances in machine learning offer an opportunity to rethink the database operations as predictive tasks that capture complex correlations and provide compact representations. However, integrating them into database systems poses significant challenges, including reasoning over complex data, controlling model size and inference, and generating representative training and evaluation workloads. To address the challenges, we first introduce a memory-efficient learned index for metric data that supports point, range, and nearest-neighbor queries. The index projects points into a learned one-dimensional ordering using distance-based transformations. For high-dimensional categorical inputs where learned models struggle to remain compact, we propose a compressed learned Bloom filter. By partitioning attributes based on value distributions, the compression reduces embedding size and memory usage while preserving accuracy and provides a foundation for compressing other approaches proposed in this thesis. We further propose supervised and unsupervised estimators for cardinality estimation in knowledge graphs to handle complex subgraph patterns and their correlations, enabling accurate predictions with low memory overhead. For collections of variable-size sets, we design permutation-invariant models that support indexing, membership, and cardinality estimation by combining compression with the hybrid integration of learned and classical components. Finally, to address the scarcity of representative training workloads, we explore generative models for automated workload synthesis and compensate for their limitations in producing selectivity-constrained queries. These models produce diverse, semantically meaningful query workloads, reducing bias and mitigating cold-start issues.
| Author: | Angjela Davitkova Gjurovska |
|---|---|
| URN: | urn:nbn:de:hbz:386-kluedo-132646 |
| DOI: | https://doi.org/10.26204/KLUEDO/13264 |
| Advisor: | Sebastian Michel |
| Document Type: | Doctoral Thesis |
| Cumulative document: | No |
| Language of publication: | English |
| Date of Publication (online): | 2026/06/25 |
| Year of first Publication: | 2026 |
| Publishing Institution: | Rheinland-Pfälzische Technische Universität Kaiserslautern-Landau |
| Granting Institution: | Rheinland-Pfälzische Technische Universität Kaiserslautern-Landau |
| Acceptance Date of the Thesis: | 2026/04/24 |
| Date of the Publication (Server): | 2026/06/26 |
| Page Number: | 140 |
| Faculties / Organisational entities: | Kaiserslautern - Fachbereich Informatik |
| CCS-Classification (computer science): | H. Information Systems |
| DDC-Cassification: | 0 Allgemeines, Informatik, Informationswissenschaft / 004 Informatik |
| Licence (German): |
