When AI Starts Training on AI: Model Collapse, AI Slop, and the Emerging Crisis of Data Provenance

Authors

Snigdha Singh Dewal

Ph.D Scholar at UPES, Dehradun (India)

Article Information

DOI: 10.51244/IJRSI.2026.1308000186

Subject Category: Artificial Intelligence

Volume/Issue: 13/8 | Page No: 3181-3185

Publication Timeline

Submitted: 2026-08-24

Accepted: 2026-09-02

Published: 2026-09-14

Abstract

Generative artificial intelligence systems no longer merely consume the human-authored internet; they now produce a substantial share of it. As synthetic text, images and data circulate back into the corpora used to train successive model generations, researchers have identified a degenerative process termed "model collapse," in which recursive training on machine-generated outputs causes models to progressively lose information about the rare, minority and low-frequency features of the original data distribution. This paper distinguishes model collapse, a training-dynamics phenomenon, from the related but distinct problem of "AI slop," a content-quality phenomenon, and argues that their interaction produces a more consequential legal concern: the erosion of data provenance. Drawing on the foundational 2024 Nature study and its subsequent refinements, together with regulatory developments such as the EU AI Act's training-data transparency obligations and technical standards including the Coalition for Content Provenance and Authenticity framework, the paper develops the concept of "epistemic due diligence" as an emerging obligation for AI developers and regulators alike. It concludes that the central resource constraint on AI development is shifting from computational power and raw data volume toward the scarcer commodity of verifiably authentic, traceable and diverse information, and that law is only beginning to develop the doctrinal tools required to govern this shift.

Keywords

Model Collapse; Synthetic Data; Data Provenance; AI Governance; EU AI Act; Epistemic Due Diligence

Downloads

References

1. Coalition for Content Provenance and Authenticity (C2PA). (2024). Content Credentials and AI/ML provenance specifications. C2PA. [Google Scholar] [Crossref]

2. European Commission. (2025). General-purpose AI training content transparency template (implementing Article 53 of the EU AI Act). European Commission. [Google Scholar] [Crossref]

3. Gerstgrasser, M., Schaeffer, R., Dey, A., Rafailov, R., Sleight, H., Hughes, J., Korbak, T., Agrawal, R., Pai, D., Gromov, A., Roberts, D. A., Yang, D., Donoho, D. L., & Koyejo, S. (2024). Is model collapse inevitable? Breaking the curse of recursion by accumulating real and synthetic data. arXiv preprint. [Google Scholar] [Crossref]

4. Jangjoo, N., et al. (2026). Lost in retraining: Closed-loop learning and model collapse in exponential families. Physical Review Letters. [Google Scholar] [Crossref]

5. Kang, S., et al. (2025). Demystifying synthetic data in LLM pre-training. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP). [Google Scholar] [Crossref]

6. Kazdan, J., Schaeffer, R., Dey, A., Gerstgrasser, M., Rafailov, R., Donoho, D. L., & Koyejo, S. (2025). Collapse or thrive: Perils and promises of synthetic data in a self-generating world. Proceedings of the 42nd International Conference on Machine Learning (ICML/PMLR). [Google Scholar] [Crossref]

7. National Institute of Standards and Technology (NIST). (2024). Reducing risks posed by synthetic content: An overview of technical approaches to digital content transparency. U.S. Department of Commerce. [Google Scholar] [Crossref]

8. Shabgahi, S., et al. (2026). ForTIFAI: Fending off recursive training induced failure for AI model collapse. npj Artificial Intelligence. [Google Scholar] [Crossref]

9. Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature, 631, 755–759. [Google Scholar] [Crossref]

10. U.S. Copyright Office. (2025). Copyright and artificial intelligence, Part 3: Generative AI training (pre-publication version). Library of Congress. [Google Scholar] [Crossref]

11. Villalobos, P., Ho, A., Sevilla, J., Besiroglu, T., Heim, L., & Hobbhahn, M. (2024). Will we run out of data? Limits of LLM scaling based on human-generated data. Proceedings of the 41st International Conference on Machine Learning (ICML/PMLR). [Google Scholar] [Crossref]

12. Ahrefs. (2025). Analysis of AI-generated content prevalence in newly created web pages, April 2025. Ahrefs. [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles