HPCM: A Hybrid Multi-Layered Machine Learning Pipeline for Plagiarism Content Matching with Dynamic Threshold Calibration
Authors
Computer Science, Pune, Maharashtra, India (India)
Computer Science, Pune, Maharashtra, India (India)
Computer Science, Pune, Maharashtra, India (India)
Computer Science, Pune, Maharashtra, India (India)
Article Information
DOI: 10.51584/IJRIAS.2026.11050027
Subject Category: Machine Learning
Volume/Issue: 11/5 | Page No: 309-323
Publication Timeline
Submitted: 2026-04-24
Accepted: 2026-04-30
Published: 2026-05-23
Abstract
Conventional approaches to detecting plagiarism involve mainly string-matching and n-gram fingerprinting methods, which can detect plagiarised documents involving verbatim plagiarism, but they cannot catch paraphrasing, synonym substitutions, or imitations of writing styles. Such shortcomings have now gained importance due to developments of sophisticated intelligent paraphrasing and the use of advanced large language models, which help evade detection by conventional approaches. In this research, we present HPCM, an end-to-end plagiarism detection system that utilises a nine-module machine-learning-based pipeline combining three analysis components: the first is the cosine similarity of terms using the TF-IDF method, secondly, embedding-based semantic similarity using the all-miniLM-L6-v2 model, and thirdly, stylistic similarity based on the analysis of POS Distribution, Type Token Ratio, and Sentence length statistics. These results are combined through the application of a weighted sum fusion function that gives greater emphasis to the semantic similarity score. Additionally, a novel Dynamic Similarity Calibration (DSC) module adjusts the plagiarism score per pair based on the relative length of documents, their vocabulary richness, and topic similarity. Experiments conducted over four different categories of plagiarism reveal that HPCM scores 69.0% in detecting paraphrases compared to 24.9% by conventional approaches, showing a remarkable 44.1 percentage point improvement. It is implemented as a microservices system on Vercel, Render, Hugging Face Spaces, and MongoDB Atlas, proving the practicality of using multilayered neural models for detecting plagiarism even with only free-tier cloud resources. The source code, along with the testing data, is publicly available.
Keywords
Plagiarism detection, natural language processing, Sentence-BERT, TF-IDF, stylometry, semantic similarity, dynamic threshold calibration, machine learning pipeline, microservice architecture.
Downloads
References
1. A. Si, H. V. Leong, and R. W. H. Lau, “CHECK: A document plagiarism detection system,” in Proc. ACM Symp. Applied Computing, 1997, pp. 70–77. [Google Scholar] [Crossref]
2. S. Schleimer, D. S. Wilkerson, and A. Aiken, “Winnowing: Local algorithms for document fingerprinting,” in Proc. ACM SIGMOD Int. Conf. Management of Data, 2003, pp. 76–85. [Google Scholar] [Crossref]
3. G. Salton, A. Wong, and C. S. Yang, “A vector space model for automatic indexing,” Commun. ACM, vol. 18, no. 11, pp. 613–620, Nov. 1975. [Google Scholar] [Crossref]
4. D. M. Blei, A. Y. Ng, and M. I. Jordan, “Latent Dirichlet Allocation,” J. Mach. Learn. Res., vol. 3, pp. 993–1022, Jan. 2003. [Google Scholar] [Crossref]
5. N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using Siamese BERT-networks,” in Proc. Conf. Empirical Methods in Natural Language Processing (EMNLP), 2019, pp. 3982–3992. [Google Scholar] [Crossref]
6. A. Aiken, “MOSS: A system for detecting software plagiarism,” Stanford University, Tech. Rep., 1994. [Online]. Available: https://theory.stanford.edu/~aiken/moss/ [Google Scholar] [Crossref]
7. M. Potthast, B. Stein, A. Barrón-Cedeño, and P. Rosso, “An evaluation framework for plagiarism detection,” in Proc. Int. Conf. Computational Linguistics (COLING), 2010, pp. 997–1005. [Google Scholar] [Crossref]
8. S. Deerwester, S. T. Dumais, G. W. Furnas, T. K. Landauer, and R. Harshman, “Indexing by latent semantic analysis,” J. Amer. Soc. Inf. Sci., vol. 41, no. 6, pp. 391–407, 1990. [Google Scholar] [Crossref]
9. T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean, “Distributed representations of words and phrases and their compositionality,” in Adv. Neural Inf. Process. Syst. (NeurIPS), 2013, pp. 3111–3119. [Google Scholar] [Crossref]
10. J. Pennington, R. Socher, and C. D. Manning, “GloVe: Global vectors for word representation,” in Proc. EMNLP, 2014, pp. 1532–1543. A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Adv. Neural Inf. Process. Syst. (NeurIPS), 2017, pp. 5998–6008. [Google Scholar] [Crossref]
11. J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. NAACL-HLT, 2019, pp. 4171–4186. [Google Scholar] [Crossref]
12. E. Stamatatos, “A survey of modern authorship attribution methods,” J. Amer. Soc. Inf. Sci. Technol., vol. 60, no. 3, pp. 538–556, Mar. 2009. S. M. Alzahrani, N. Salim, and A. Abraham, “Understanding plagiarism linguistic patterns, textual features, and detection methods,” IEEE Trans. Syst., Man, Cybern. C, vol. 42, no. 2, pp. 133–149, Mar. 2012. [Google Scholar] [Crossref]
13. A. Barrón-Cedeño and P. Rosso, “On automatic plagiarism detection based on n-grams comparison,” in Proc. European Conf. Information Retrieval (ECIR), 2009, pp. 696–700. [Google Scholar] [Crossref]
14. M. Potthast, T. Gollub, M. Hagen, J. Grabegger, J. Kiesel, M. Michel, A. Oberländer, M. Tippmann, A. Barrón-Cedeño, P. Gupta, P. Rosso, and B. Stein, “Overview of the 4th international competition on plagiarism detection,” in CLEF Conf. on Multilingual and Multimodal Information Access Evaluation, 2012, pp. 17–19. W. Daelemans, “Explanation in computational stylometry,” in Proc. Int. Conf. Computational Linguistics and Intelligent Text Processing, 2013, pp. 451–462. [Google Scholar] [Crossref]
15. H. Maurer, F. Kappe, and B. Zaka, “Plagiarism — A survey,” J. Universal Comput. Sci., vol. 12, no. 8, pp. 1050–1084, 2006. [Google Scholar] [Crossref]
16. P. Clough and M. Stevenson, “Developing a corpus of plagiarised short answers,” Language Resources and Evaluation, vol. 45, no. 1, pp. 5–24, Mar. 2011. A. Paszke et al., “PyTorch: An imperative style, high-performance deep learning library,” in Adv. Neural Inf. Process. Syst. (NeurIPS), 2019, pp. 8024–8035. [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- A Machine Learning Model for Predicting the Risk of Developing Diabetes - T2DM Using Real-World Data from Kilifi, Kenya
- AI-Powered Facial Recognition Attendance System Using Deep Learning and Computer Vision
- A Comprehensive Review on Brain Tumour Segmentation Using Deep Learning Approach
- A Scalable Retrieval-Augmented Generation Pipeline for Domain-Specific Knowledge Applications
- Predictive Maintenance in Semiconductor Manufacturing Using Machine Learning on Imbalanced Dataset