When AI Starts Training on AI: Model Collapse, AI Slop, and the Emerging Crisis of Data Provenance
Authors
Ph.D Scholar at UPES, Dehradun (India)
Article Information
DOI: 10.51244/IJRSI.2026.1308000186
Subject Category: Artificial Intelligence
Volume/Issue: 13/8 | Page No: 3181-3185
Publication Timeline
Submitted: 2026-08-24
Accepted: 2026-09-02
Published: 2026-09-14
Abstract
Generative artificial intelligence systems no longer merely consume the human-authored internet; they now produce a substantial share of it. As synthetic text, images and data circulate back into the corpora used to train successive model generations, researchers have identified a degenerative process termed "model collapse," in which recursive training on machine-generated outputs causes models to progressively lose information about the rare, minority and low-frequency features of the original data distribution. This paper distinguishes model collapse, a training-dynamics phenomenon, from the related but distinct problem of "AI slop," a content-quality phenomenon, and argues that their interaction produces a more consequential legal concern: the erosion of data provenance. Drawing on the foundational 2024 Nature study and its subsequent refinements, together with regulatory developments such as the EU AI Act's training-data transparency obligations and technical standards including the Coalition for Content Provenance and Authenticity framework, the paper develops the concept of "epistemic due diligence" as an emerging obligation for AI developers and regulators alike. It concludes that the central resource constraint on AI development is shifting from computational power and raw data volume toward the scarcer commodity of verifiably authentic, traceable and diverse information, and that law is only beginning to develop the doctrinal tools required to govern this shift.
Keywords
Model Collapse; Synthetic Data; Data Provenance; AI Governance; EU AI Act; Epistemic Due Diligence
Downloads
References
1. Coalition for Content Provenance and Authenticity (C2PA). (2024). Content Credentials and AI/ML provenance specifications. C2PA. [Google Scholar] [Crossref]
2. European Commission. (2025). General-purpose AI training content transparency template (implementing Article 53 of the EU AI Act). European Commission. [Google Scholar] [Crossref]
3. Gerstgrasser, M., Schaeffer, R., Dey, A., Rafailov, R., Sleight, H., Hughes, J., Korbak, T., Agrawal, R., Pai, D., Gromov, A., Roberts, D. A., Yang, D., Donoho, D. L., & Koyejo, S. (2024). Is model collapse inevitable? Breaking the curse of recursion by accumulating real and synthetic data. arXiv preprint. [Google Scholar] [Crossref]
4. Jangjoo, N., et al. (2026). Lost in retraining: Closed-loop learning and model collapse in exponential families. Physical Review Letters. [Google Scholar] [Crossref]
5. Kang, S., et al. (2025). Demystifying synthetic data in LLM pre-training. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP). [Google Scholar] [Crossref]
6. Kazdan, J., Schaeffer, R., Dey, A., Gerstgrasser, M., Rafailov, R., Donoho, D. L., & Koyejo, S. (2025). Collapse or thrive: Perils and promises of synthetic data in a self-generating world. Proceedings of the 42nd International Conference on Machine Learning (ICML/PMLR). [Google Scholar] [Crossref]
7. National Institute of Standards and Technology (NIST). (2024). Reducing risks posed by synthetic content: An overview of technical approaches to digital content transparency. U.S. Department of Commerce. [Google Scholar] [Crossref]
8. Shabgahi, S., et al. (2026). ForTIFAI: Fending off recursive training induced failure for AI model collapse. npj Artificial Intelligence. [Google Scholar] [Crossref]
9. Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature, 631, 755–759. [Google Scholar] [Crossref]
10. U.S. Copyright Office. (2025). Copyright and artificial intelligence, Part 3: Generative AI training (pre-publication version). Library of Congress. [Google Scholar] [Crossref]
11. Villalobos, P., Ho, A., Sevilla, J., Besiroglu, T., Heim, L., & Hobbhahn, M. (2024). Will we run out of data? Limits of LLM scaling based on human-generated data. Proceedings of the 41st International Conference on Machine Learning (ICML/PMLR). [Google Scholar] [Crossref]
12. Ahrefs. (2025). Analysis of AI-generated content prevalence in newly created web pages, April 2025. Ahrefs. [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- The Role of Artificial Intelligence in Revolutionizing Library Services in Nairobi: Ethical Implications and Future Trends in User Interaction
- ESPYREAL: A Mobile Based Multi-Currency Identifier for Visually Impaired Individuals Using Convolutional Neural Network
- Comparative Analysis of AI-Driven IoT-Based Smart Agriculture Platforms with Blockchain-Enabled Marketplaces
- AI-Based Dish Recommender System for Reducing Fruit Waste through Spoilage Detection and Ripeness Assessment
- SEA-TALK: An AI-Powered Voice Translator and Southeast Asian Dialects Recognition