A SYSTEMATIC REVIEW OF PRIVACY RISKS AND DEFENSES IN RETRIEVAL-AUGMENTED AND AGENTIC LARGE LANGUAGE MODEL SYSTEMS
Keywords:
Large language models · Retrieval augmented generation · LLM agents · Privacy leakage · Membership inference · Differential privacy · Prompt injection · integrityAbstract
Large language models are no longer used as stand‑alone text generators. In real‑world situations, they become part of systems that pull documents from private data sources, track past chats, and use external tools to act on a user’ s behalf. This shift changes where personal information is stored and how it might leak. Previous privacy studies looked mainly at what a model remembers during training; less focus has been on leaks from search index embedding storage, agent memory, and tool connections. This review fills that gap. We group research published from 2023 to 2026 into a three-level system that splits risks into model data (L1), the search component (L2), and acting agents (L3). For each level, we look at how an attack works, what access the attacker needs, and how strong the proof is, using examples like copying search content at a 90 % success rate for adding information, five fake passages recovering 92 % of short texts from their embeddings, and models sharing too much context in 25–57 % of test cases. We then show which protections, like privacy in training and search, deleting data, fake data sets, limiting access, keeping control and data separate, and using hardware, match the threats they handle. We measure the cost to performance when numbers are available. The study shows that no single method covers all risks and that the weakest‑protected risks appear when the whole system works together. We ended with a look at how to evaluate these systems, how to link them to rules, and a research plan for keeping privacy safe in complex AI systems.


