Show/Hide Menu
Hide/Show Apps
Logout
Türkçe
Türkçe
Search
Search
Login
Login
OpenMETU
OpenMETU
About
About
Open Science Policy
Open Science Policy
Open Access Guideline
Open Access Guideline
Postgraduate Thesis Guideline
Postgraduate Thesis Guideline
Communities & Collections
Communities & Collections
Help
Help
Frequently Asked Questions
Frequently Asked Questions
Guides
Guides
Thesis submission
Thesis submission
MS without thesis term project submission
MS without thesis term project submission
Publication submission with DOI
Publication submission with DOI
Publication submission
Publication submission
Supporting Information
Supporting Information
General Information
General Information
Copyright, Embargo and License
Copyright, Embargo and License
Contact us
Contact us
SALDIRAY: Scalable and Adaptive Language Diagnostics for Remediation of Anomalies and Typos
Date
2026-01-01
Author
Külah, Emre
Çetinkaya, Yusuf Mucahit
Alemdar, Hande
Metadata
Show full item record
This work is licensed under a
Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License
.
Item Usage Stats
50
views
0
downloads
Cite This
Natural language processing (NLP) systems working with real-world user data often encounter noisy, informal text, including typos, spelling variations, slang, and homophone substitutions. This noise can cause a severe drop in the performance of downstream models. This paper presents SALDIRAY (Scalable and Adaptive Language Diagnostics for Remediation of Anomalies and Typos), an integrated, domain-independent framework for text normalization that aims to boost the resilience of NLP systems across various model architectures and application types. It converts short, poorly formed text into standardized language through a synthetic noise modeling and standardization approach, instantiated in practice as a noise generation and normalization pipeline. This approach is enriched by incorporating curated slang and homophone lists to ensure linguistic realism. The process creates large amounts of paired noisy–clean data, which are used to fine-tune robust normalization models. Our detailed experiments consistently show that SALDIRAYimproves the accuracy and stability of systems across a wide range of tasks, including traditional machine learning–based classification, sentiment analysis, and instruction-tuned large language model tasks. The framework requires no task-specific supervision, operates as a modular preprocessing layer, and can be integrated into existing NLP workflows. The experiments confirm that the proposed framework offers a scalable and adaptive method for improving robustness to noisy text, particularly under controlled corruption settings, while showing promising transfer to real-world noisy inputs.
URI
https://doi.org/10.1109/ACCESS.2026.3693355
https://hdl.handle.net/11511/119212
Journal
IEEE ACCESS
DOI
https://doi.org/10.1109/access.2026.3693355
Collections
Department of Computer Engineering, Article
Citation Formats
IEEE
ACM
APA
CHICAGO
MLA
BibTeX
E. Külah, Y. M. Çetinkaya, and H. Alemdar, “SALDIRAY: Scalable and Adaptive Language Diagnostics for Remediation of Anomalies and Typos,”
IEEE ACCESS
, vol. 14, pp. 1–20, 2026, Accessed: 00, 2026. [Online]. Available: https://doi.org/10.1109/ACCESS.2026.3693355.