Document Type

Article

Publication Date

2026

Abstract

Arabic SMS-based e-commerce platforms pose unique challenges due to the spontaneous and noisy nature of user-generated text (e.g., abbreviations, dialectal Arabic, or “Arabizi” transliterations). In this paper, we present Classified Ads Text Service (CATS) 2.0, an improved classified ads system that combines probabilistic large language models (LLMs) with deterministic graph-based knowledge representations to achieve robust understanding and matching of Arabic SMS content. Building on earlier work that emphasized the importance of integrating sublanguage analysis with content-oriented methods, our approach uses a hybrid pipeline: an LLM interprets free-text messages and extracts structured information, which is then inserted into a Neo4j graph database representing the domain knowledge. This graph-based representation enables precise semantic matching of “selling” and “looking for” posts and supports reasoning over the ads network. We evaluate the system on real-world Arabic SMS e-commerce data. Experimental results show that the hybrid CATS 2.0 system achieves high accuracy in content extraction (improving the F-measure over the original system’s ~90%) and successfully handles multilingual and transliterated inputs. The proposed approach demonstrates how coupling an LLM’s flexibility with a knowledge graph’s rigor can substantially enhance the robustness and extensibility of e-commerce text processing in Arabic.

Share

COinS