Enhancing Resemblance Matching using Structural Awareness for Hierarchical LLM Caching

Authors

Dr. Chaitanya Udatha

Information Technology, Mahatma Gandhi Institute of Technology (MGIT), Hyderabad, India. (India)

Krithi Chippada

Information Technology, Mahatma Gandhi Institute of Technology (MGIT), Hyderabad, India. (India)

Satvik Dabbara

Information Technology, Mahatma Gandhi Institute of Technology (MGIT), Hyderabad, India. (India)

Article Information

DOI: 10.51584/IJRIAS.2026.11050049

Subject Category: Artificial Intelligence

Volume/Issue: 11/5 | Page No: 571-581

Publication Timeline

Submitted: 2026-05-03

Accepted: 2026-05-08

Published: 2026-05-27

Abstract

Large Language Models (LLMs) have become an integral part of our daily lives; they are used for tasks such as chatbots in customer services and require a lot of computing power. If the user base is large, generating different responses to similar queries results in slower performance and increased computational latency. Hence hierarchical caching systems like GPTCache and MinCache were introduced to reduce redundant inference using exact matching, resemblance matching and semantic matching of the prompts with stored queries to reuse LLM responses for similar queries. However, Unigram-based resemblance caching mechanisms are susceptible to adversarial lexical reordering leading to excessive false positive cache hits. The proposed research introduced structural-aware resemblance matching to improve the robustness of the system without violating the ideology of MinCache by using lightweight and fast similarity caching mechanisms. It has achieved 7.39x safer cache reuse compared to the standard 1-g Minhash while preserving 79.5% of resemblance layer throughput and maintained overall accuracy.

Keywords

Structural Shingling, Skip-Gram Similarity, Cache Eviction Policies, Large Language Models, Low-Latency Inference, Paraphrase Identification, Adversarial Text Reordering.

Downloads

References

1. Brown T, Mann B, Ryder N, Subbiah M, Kaplan JD, Dhariwal P, et al. (2020). Language models are few-shot learners. Adv. Neural Inf. Process. Syst. 33: 1877–1901. [Google Scholar] [Crossref]

2. Bommasani R, Hudson DA, Adeli E, Altman R, Arora S, von Arx S, et al. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258. [Google Scholar] [Crossref]

3. Zhou Z, Ning X, Hong K, Fu T, Xu J, Li S, et al. (2024). A survey on efficient inference for large language models. arXiv preprint arXiv:2404.14294. [Google Scholar] [Crossref]

4. Yuan Z, Shang Y, Zhou Y, Dong Z, Zhou Z, Xue C, et al. (2024). LLM inference unveiled: Survey and roofline model insights. arXiv preprint arXiv:2402.16363. [Google Scholar] [Crossref]

5. Xie Y, O’Hallaron D (2001). Locality in search engine queries and its implications for caching. CMU Technical Report CMU-CS-01-128. [Google Scholar] [Crossref]

6. Mookerjee VS, Tan Y (2002). Analysis of a least recently used cache management policy for web browsers. Oper. Res. 50(2): 345–357. [Google Scholar] [Crossref]

7. Markatos EP (2001). On caching search engine query results. Comput. Commun. 24(2): 137–143. [Google Scholar] [Crossref]

8. Miao X, Oliaro G, Zhang Z, Cheng X, Wang Z, Zhang Z, et al. (2024). SpecInfer: Accelerating large language model serving with tree-based speculative inference and verification. Proc. ACM ASPLOS: 932–949. [Google Scholar] [Crossref]

9. Ramírez G, Lindemann M, Birch A, Titov I (2024). Cache & distil: Optimising API calls to large language models. Findings ACL: 11838–11853. [Google Scholar] [Crossref]

10. Zhang Z, Sheng Y, Zhou T, Chen T, Zheng L, Cai R, et al. (2023). H2O: Heavy-hitter oracle for efficient generative inference of large language models. Adv. Neural Inf. Process. Syst. 36: 34661–34710. [Google Scholar] [Crossref]

11. Li J, Xu C, Wang F, von Riedemann IM, Zhang C, Liu J (2024). SCALM: Towards semantic caching for automated chat services with large language models. Proc. IEEE/ACM IWQoS: 1–10. [Google Scholar] [Crossref]

12. Bang F (2023). GPTCache: An open-source semantic cache for LLM applications enabling faster answers and cost savings. Proc. NLP-OSS Workshop: 212–218. [Google Scholar] [Crossref]

13. Haqiq K, Jahan MV, Farimani SA, Masoom SMF (2025). MinCache: A hybrid cache system for efficient chatbots with hierarchical embedding matching and LLM. Future Gener. Comput. Syst. 170: 107822. [Google Scholar] [Crossref]

14. Wu H, Pan J, Yang R, Zhang H, Shi G, Jiang Z, Liu Q (2025). PAWS: Passive concept drift adaptation based on instance weighting and subspace alignment in data stream. Proc. Int. Conf. Intelligent Computing: 236–247. [Google Scholar] [Crossref]

15. Quora Question Pairs Dataset (2017). Kaggle. [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles