Enhancing Resemblance Matching using Structural Awareness for Hierarchical LLM Caching
Authors
Information Technology, Mahatma Gandhi Institute of Technology (MGIT), Hyderabad, India. (India)
Information Technology, Mahatma Gandhi Institute of Technology (MGIT), Hyderabad, India. (India)
Information Technology, Mahatma Gandhi Institute of Technology (MGIT), Hyderabad, India. (India)
Article Information
DOI: 10.51584/IJRIAS.2026.11050049
Subject Category: Artificial Intelligence
Volume/Issue: 11/5 | Page No: 571-581
Publication Timeline
Submitted: 2026-05-03
Accepted: 2026-05-08
Published: 2026-05-27
Abstract
Large Language Models (LLMs) have become an integral part of our daily lives; they are used for tasks such as chatbots in customer services and require a lot of computing power. If the user base is large, generating different responses to similar queries results in slower performance and increased computational latency. Hence hierarchical caching systems like GPTCache and MinCache were introduced to reduce redundant inference using exact matching, resemblance matching and semantic matching of the prompts with stored queries to reuse LLM responses for similar queries. However, Unigram-based resemblance caching mechanisms are susceptible to adversarial lexical reordering leading to excessive false positive cache hits. The proposed research introduced structural-aware resemblance matching to improve the robustness of the system without violating the ideology of MinCache by using lightweight and fast similarity caching mechanisms. It has achieved 7.39x safer cache reuse compared to the standard 1-g Minhash while preserving 79.5% of resemblance layer throughput and maintained overall accuracy.
Keywords
Structural Shingling, Skip-Gram Similarity, Cache Eviction Policies, Large Language Models, Low-Latency Inference, Paraphrase Identification, Adversarial Text Reordering.
Downloads
References
1. Brown T, Mann B, Ryder N, Subbiah M, Kaplan JD, Dhariwal P, et al. (2020). Language models are few-shot learners. Adv. Neural Inf. Process. Syst. 33: 1877–1901. [Google Scholar] [Crossref]
2. Bommasani R, Hudson DA, Adeli E, Altman R, Arora S, von Arx S, et al. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258. [Google Scholar] [Crossref]
3. Zhou Z, Ning X, Hong K, Fu T, Xu J, Li S, et al. (2024). A survey on efficient inference for large language models. arXiv preprint arXiv:2404.14294. [Google Scholar] [Crossref]
4. Yuan Z, Shang Y, Zhou Y, Dong Z, Zhou Z, Xue C, et al. (2024). LLM inference unveiled: Survey and roofline model insights. arXiv preprint arXiv:2402.16363. [Google Scholar] [Crossref]
5. Xie Y, O’Hallaron D (2001). Locality in search engine queries and its implications for caching. CMU Technical Report CMU-CS-01-128. [Google Scholar] [Crossref]
6. Mookerjee VS, Tan Y (2002). Analysis of a least recently used cache management policy for web browsers. Oper. Res. 50(2): 345–357. [Google Scholar] [Crossref]
7. Markatos EP (2001). On caching search engine query results. Comput. Commun. 24(2): 137–143. [Google Scholar] [Crossref]
8. Miao X, Oliaro G, Zhang Z, Cheng X, Wang Z, Zhang Z, et al. (2024). SpecInfer: Accelerating large language model serving with tree-based speculative inference and verification. Proc. ACM ASPLOS: 932–949. [Google Scholar] [Crossref]
9. Ramírez G, Lindemann M, Birch A, Titov I (2024). Cache & distil: Optimising API calls to large language models. Findings ACL: 11838–11853. [Google Scholar] [Crossref]
10. Zhang Z, Sheng Y, Zhou T, Chen T, Zheng L, Cai R, et al. (2023). H2O: Heavy-hitter oracle for efficient generative inference of large language models. Adv. Neural Inf. Process. Syst. 36: 34661–34710. [Google Scholar] [Crossref]
11. Li J, Xu C, Wang F, von Riedemann IM, Zhang C, Liu J (2024). SCALM: Towards semantic caching for automated chat services with large language models. Proc. IEEE/ACM IWQoS: 1–10. [Google Scholar] [Crossref]
12. Bang F (2023). GPTCache: An open-source semantic cache for LLM applications enabling faster answers and cost savings. Proc. NLP-OSS Workshop: 212–218. [Google Scholar] [Crossref]
13. Haqiq K, Jahan MV, Farimani SA, Masoom SMF (2025). MinCache: A hybrid cache system for efficient chatbots with hierarchical embedding matching and LLM. Future Gener. Comput. Syst. 170: 107822. [Google Scholar] [Crossref]
14. Wu H, Pan J, Yang R, Zhang H, Shi G, Jiang Z, Liu Q (2025). PAWS: Passive concept drift adaptation based on instance weighting and subspace alignment in data stream. Proc. Int. Conf. Intelligent Computing: 236–247. [Google Scholar] [Crossref]
15. Quora Question Pairs Dataset (2017). Kaggle. [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- The Role of Artificial Intelligence in Revolutionizing Library Services in Nairobi: Ethical Implications and Future Trends in User Interaction
- ESPYREAL: A Mobile Based Multi-Currency Identifier for Visually Impaired Individuals Using Convolutional Neural Network
- Comparative Analysis of AI-Driven IoT-Based Smart Agriculture Platforms with Blockchain-Enabled Marketplaces
- AI-Based Dish Recommender System for Reducing Fruit Waste through Spoilage Detection and Ripeness Assessment
- SEA-TALK: An AI-Powered Voice Translator and Southeast Asian Dialects Recognition