HPCM: A Hybrid Multi-Layered Machine Learning Pipeline for Plagiarism Content Matching with Dynamic Threshold Calibration

Authors

Piyush Chavan

Computer Science, Pune, Maharashtra, India (India)

Prof.Moushmee Kuri

Computer Science, Pune, Maharashtra, India (India)

Tanvi Bokade

Computer Science, Pune, Maharashtra, India (India)

Pushkar Thombare

Computer Science, Pune, Maharashtra, India (India)

Article Information

DOI: 10.51584/IJRIAS.2026.11050027

Subject Category: Machine Learning

Volume/Issue: 11/5 | Page No: 309-323

Publication Timeline

Submitted: 2026-04-24

Accepted: 2026-04-30

Published: 2026-05-23

Abstract

Conventional approaches to detecting plagiarism involve mainly string-matching and n-gram fingerprinting methods, which can detect plagiarised documents involving verbatim plagiarism, but they cannot catch paraphrasing, synonym substitutions, or imitations of writing styles. Such shortcomings have now gained importance due to developments of sophisticated intelligent paraphrasing and the use of advanced large language models, which help evade detection by conventional approaches. In this research, we present HPCM, an end-to-end plagiarism detection system that utilises a nine-module machine-learning-based pipeline combining three analysis components: the first is the cosine similarity of terms using the TF-IDF method, secondly, embedding-based semantic similarity using the all-miniLM-L6-v2 model, and thirdly, stylistic similarity based on the analysis of POS Distribution, Type Token Ratio, and Sentence length statistics. These results are combined through the application of a weighted sum fusion function that gives greater emphasis to the semantic similarity score. Additionally, a novel Dynamic Similarity Calibration (DSC) module adjusts the plagiarism score per pair based on the relative length of documents, their vocabulary richness, and topic similarity. Experiments conducted over four different categories of plagiarism reveal that HPCM scores 69.0% in detecting paraphrases compared to 24.9% by conventional approaches, showing a remarkable 44.1 percentage point improvement. It is implemented as a microservices system on Vercel, Render, Hugging Face Spaces, and MongoDB Atlas, proving the practicality of using multilayered neural models for detecting plagiarism even with only free-tier cloud resources. The source code, along with the testing data, is publicly available.

Keywords

Plagiarism detection, natural language processing, Sentence-BERT, TF-IDF, stylometry, semantic similarity, dynamic threshold calibration, machine learning pipeline, microservice architecture.

Downloads

References

1. A. Si, H. V. Leong, and R. W. H. Lau, “CHECK: A document plagiarism detection system,” in Proc. ACM Symp. Applied Computing, 1997, pp. 70–77. [Google Scholar] [Crossref]

2. S. Schleimer, D. S. Wilkerson, and A. Aiken, “Winnowing: Local algorithms for document fingerprinting,” in Proc. ACM SIGMOD Int. Conf. Management of Data, 2003, pp. 76–85. [Google Scholar] [Crossref]

3. G. Salton, A. Wong, and C. S. Yang, “A vector space model for automatic indexing,” Commun. ACM, vol. 18, no. 11, pp. 613–620, Nov. 1975. [Google Scholar] [Crossref]

4. D. M. Blei, A. Y. Ng, and M. I. Jordan, “Latent Dirichlet Allocation,” J. Mach. Learn. Res., vol. 3, pp. 993–1022, Jan. 2003. [Google Scholar] [Crossref]

5. N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using Siamese BERT-networks,” in Proc. Conf. Empirical Methods in Natural Language Processing (EMNLP), 2019, pp. 3982–3992. [Google Scholar] [Crossref]

6. A. Aiken, “MOSS: A system for detecting software plagiarism,” Stanford University, Tech. Rep., 1994. [Online]. Available: https://theory.stanford.edu/~aiken/moss/ [Google Scholar] [Crossref]

7. M. Potthast, B. Stein, A. Barrón-Cedeño, and P. Rosso, “An evaluation framework for plagiarism detection,” in Proc. Int. Conf. Computational Linguistics (COLING), 2010, pp. 997–1005. [Google Scholar] [Crossref]

8. S. Deerwester, S. T. Dumais, G. W. Furnas, T. K. Landauer, and R. Harshman, “Indexing by latent semantic analysis,” J. Amer. Soc. Inf. Sci., vol. 41, no. 6, pp. 391–407, 1990. [Google Scholar] [Crossref]

9. T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean, “Distributed representations of words and phrases and their compositionality,” in Adv. Neural Inf. Process. Syst. (NeurIPS), 2013, pp. 3111–3119. [Google Scholar] [Crossref]

10. J. Pennington, R. Socher, and C. D. Manning, “GloVe: Global vectors for word representation,” in Proc. EMNLP, 2014, pp. 1532–1543. A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Adv. Neural Inf. Process. Syst. (NeurIPS), 2017, pp. 5998–6008. [Google Scholar] [Crossref]

11. J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. NAACL-HLT, 2019, pp. 4171–4186. [Google Scholar] [Crossref]

12. E. Stamatatos, “A survey of modern authorship attribution methods,” J. Amer. Soc. Inf. Sci. Technol., vol. 60, no. 3, pp. 538–556, Mar. 2009. S. M. Alzahrani, N. Salim, and A. Abraham, “Understanding plagiarism linguistic patterns, textual features, and detection methods,” IEEE Trans. Syst., Man, Cybern. C, vol. 42, no. 2, pp. 133–149, Mar. 2012. [Google Scholar] [Crossref]

13. A. Barrón-Cedeño and P. Rosso, “On automatic plagiarism detection based on n-grams comparison,” in Proc. European Conf. Information Retrieval (ECIR), 2009, pp. 696–700. [Google Scholar] [Crossref]

14. M. Potthast, T. Gollub, M. Hagen, J. Grabegger, J. Kiesel, M. Michel, A. Oberländer, M. Tippmann, A. Barrón-Cedeño, P. Gupta, P. Rosso, and B. Stein, “Overview of the 4th international competition on plagiarism detection,” in CLEF Conf. on Multilingual and Multimodal Information Access Evaluation, 2012, pp. 17–19. W. Daelemans, “Explanation in computational stylometry,” in Proc. Int. Conf. Computational Linguistics and Intelligent Text Processing, 2013, pp. 451–462. [Google Scholar] [Crossref]

15. H. Maurer, F. Kappe, and B. Zaka, “Plagiarism — A survey,” J. Universal Comput. Sci., vol. 12, no. 8, pp. 1050–1084, 2006. [Google Scholar] [Crossref]

16. P. Clough and M. Stevenson, “Developing a corpus of plagiarised short answers,” Language Resources and Evaluation, vol. 45, no. 1, pp. 5–24, Mar. 2011. A. Paszke et al., “PyTorch: An imperative style, high-performance deep learning library,” in Adv. Neural Inf. Process. Syst. (NeurIPS), 2019, pp. 8024–8035. [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles