Show/Hide Menu
Hide/Show Apps
Logout
Türkçe
Türkçe
Search
Search
Login
Login
OpenMETU
OpenMETU
About
About
Open Science Policy
Open Science Policy
Open Access Guideline
Open Access Guideline
Postgraduate Thesis Guideline
Postgraduate Thesis Guideline
Communities & Collections
Communities & Collections
Help
Help
Frequently Asked Questions
Frequently Asked Questions
Guides
Guides
Thesis submission
Thesis submission
MS without thesis term project submission
MS without thesis term project submission
Publication submission with DOI
Publication submission with DOI
Publication submission
Publication submission
Supporting Information
Supporting Information
General Information
General Information
Copyright, Embargo and License
Copyright, Embargo and License
Contact us
Contact us
AN LLM-POWERED CONVERSATIONAL ANALYTIC SYSTEM FOR INTELLIGENT DATA DISCOVERY ACROSS MESH-FABRIC DATA ENVIRONMENTS
Download
thesis_v18.pdf
Elif Beril Şayli_Tez Teslim Belgeleri.pdf
Date
2026-6-18
Author
Şayli, Elif Beril
Metadata
Show full item record
This work is licensed under a
Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License
.
Item Usage Stats
57
views
0
downloads
Cite This
Today, modern organizations operate in increasingly complex data environments where traditional centralized architectures struggle to meet demands for scalability and agility. A persistent challenge is the effective management of metadata and intuitive data discovery, particularly when natural-language questions must be translated into executable queries over raw lakehouse storage where foreign-key constraints are not explicitly declared. This thesis presents an approach that uses LLM-based metadata agents to support data discovery, enrichment, cataloging, and structuring across domains, and translates natural-language questions into SQL queries grounded in inferred metadata. The goal is to reduce the manual effort required for cataloging and schema exploration within an architecture informed by Data Mesh and Data Fabric principles. The system provides LLM-assisted metadata extraction and relationship inference to generate structured artifacts stored in versioned machine-readable formats so they can be inspected, reused, and updated as datasets evolve. This approach is useful in environments that require both decentralized ownership and cross-system interoperability while supporting consistency and reproducibility. The proposed system gathers schema and table metadata through the catalog and query layers and uses it for metadata enrichment and LLM-assisted SQL generation. The approach is evaluated on a controlled multi-domain benchmark by comparing configurations with and without inferred relationship metadata. The relation-aware configuration achieves a statistically significant correctness improvement on a specific set of analytical query patterns. The results show that versioned, inspectable relationship metadata can support NL to SQL generation in lakehouse environments under the tested conditions.
Subject Keywords
Data Lakehouse
,
Data Mesh
,
Data Fabric
,
Large Language Models
,
Natural Language to SQL
URI
https://hdl.handle.net/11511/119728
Collections
Graduate School of Informatics, Thesis
Citation Formats
IEEE
ACM
APA
CHICAGO
MLA
BibTeX
E. B. Şayli, “AN LLM-POWERED CONVERSATIONAL ANALYTIC SYSTEM FOR INTELLIGENT DATA DISCOVERY ACROSS MESH-FABRIC DATA ENVIRONMENTS,” M.S. - Master of Science, Middle East Technical University, 2026.