Agentic RAG: Architecture, Knowledge, Access Control
How agentic RAG changes architecture, knowledge and access control: corrective vs self-reflective designs, what they cost, and where permission checks belong.
Retrieval-augmented generation often assumes that retrieved passages are suitable evidence for an answer. An agentic system can plan its retrieval, choose between sources, or judge what came back. This piece follows the third of those, because it is the one that reaches into your knowledge layer and your permission model at the same time.
The failure that does not announce itself #
A conventional RAG pipeline does two things in sequence. It retrieves a set of passages for the user's question, then generates an answer from them. Nothing sits between those steps to ask whether the passages were worth retrieving.
The authors of the CRAG paper put the structural problem plainly: the retriever and the generator are tightly coupled, and any unsuccessful retrieval produces an unsatisfactory response regardless of how capable the generator is (Yan et al., 2024). Their worked example is a question about a film's screenwriter where the retriever returns a passage about a different film entirely. The generator answers confidently and wrongly.
What makes this expensive in production is not the error rate. It is that the error is indistinguishable from a correct answer at the point of delivery. A system that fails loudly can be monitored. Fluent, well-formatted errors can be hard to detect. Without a retrieval trace, they can also be hard to investigate.
Everything below follows from one response to that problem: stop treating retrieval as a step, and start treating it as a decision the system makes, evaluates, and remakes.
What "agentic" actually adds #
In agentic RAG, an agent can choose tools or sources, inspect the results, and decide whether to search again. The important design choice is which decisions the system may make on its own. The distinction matters because it changes where decisions happen. In the survey literature, these systems are characterised by four agentic patterns: reflection, planning, tool use, and multi-agent collaboration (Singh et al., 2025).
Applied to retrieval, those patterns produce a system that can decide whether retrieval is needed at all, judge what came back, retrieve again with a different query, reach for a different source, or stop. The pipeline acquires a feedback signal it previously lacked.
That is the whole shift, and it has three consequences that structure the rest of this article. The system now makes decisions, so you have to choose how those decisions are controlled. It needs something to decide about, so your knowledge layer has to expose more than text. And it issues many retrievals across many sources per user question, so every one of those retrievals is a place where the wrong person can see the wrong document.
Architecture: the axes you are choosing between #
The four design dimensions
Most published overviews of agentic RAG offer a list of named patterns: single-agent, multi-agent, corrective, adaptive, graph-based. Useful as a map of what exists, less useful as a way to decide. Both patterns examined below are corrective. They let retrieval happen and then judge what came back. Adaptive systems make the earlier decision of whether to retrieve at all, and multi-agent ones spread one question across several sources. Those add a planning turn rather than a grading step. The current version of the Singh survey organises the field along four design dimensions instead: agent cardinality, control structure, autonomy, and knowledge representation (Singh et al., 2025).
These are the things you are actually choosing. How many agents. Who directs whom. How much the system may decide without a human or a hard rule. How knowledge is structured for the agent to work over. Named patterns are points in that space, and the space is worth understanding directly, because the named list changes faster than the dimensions do.
Two patterns are worth examining closely, because they represent genuinely different answers to "how does the system know retrieval went wrong".
Decide-then-critique
Self-RAG trains a single language model to retrieve passages on demand and to reflect on both the retrieved passages and its own generated text, using special tokens the model emits called reflection tokens. Because the model produces those tokens as part of generation, its behaviour becomes controllable at inference time rather than fixed at training time (Asai et al., 2023).
The appeal is that judgement lives inside the model. The cost is that the judgement has to be trained in: the capability comes from instruction tuning on data containing those tokens, which means swapping in a newer base model is not a configuration change.
Grade-then-correct
CRAG takes the opposite approach and puts judgement outside the generator. A lightweight retrieval evaluator, a fine-tuned T5-large with 0.77B parameters, scores the relevance of retrieved documents. Those scores are quantified into three confidence degrees that trigger different actions: Correct, Incorrect, and Ambiguous (Yan et al., 2024).
Each action does something concrete. Correct sends the documents through a decompose-then-recompose refinement that splits them into strips, filters the irrelevant ones, and reassembles what survives. Incorrect discards the retrieved set entirely and runs a web search instead, with the query rewritten into keywords first. Ambiguous combines both, and the authors report that this middle action is what reduces the system's dependence on the evaluator being right.
The evaluator is also the part that scales down well. On PopQA, it assessed retrieval relevance at 84.3% accuracy against 58.0% for prompted ChatGPT in the same role. A small fine-tuned model beat a large general one at this particular job, which is worth knowing before you budget for an LLM-as-judge architecture.
What it costs
Agency is not free, and the CRAG paper is one of the few sources that publishes the overhead directly. Measured as average execution time per instance, standard RAG ran at 0.363 seconds, CRAG at 0.512, Self-RAG at 0.741, and Self-CRAG at 0.908. Estimated FLOPs per token tell a similar story: 26.5 TFLOPs for standard RAG, 27.2 for CRAG, and a range of 26.5 to 132.4 for Self-RAG, whose cost varies because its strategy varies by input (Yan et al., 2024).
Two caveats on those numbers, both from the authors. They estimate the generation phase only, excluding retrieval and data processing. And the comparison is between specific implementations, not between architectural categories in general.
In these implementations, CRAG added less generation time than Self-RAG. The figures exclude retrieval and data processing, so they do not establish end-to-end latency or its variability.
Knowledge: what the agent reasons over #
An agent that can choose to retrieve again needs grounds for the choice. This is where a lot of otherwise well-built systems stall: the retrieval loop is sound and the index gives it nothing to reason with.
Metadata is a decision input, not bookkeeping
An evaluator can judge a chunk's relevance from its text, but source, date, and classification provide further grounds for assessing provenance, freshness, and access. Source, recency, and sensitivity have to survive ingestion and be visible at retrieval time, because they are the inputs to "should I trust this or look elsewhere". This is engineering judgement rather than a cited finding, but it follows directly from what the corrective architectures above are doing: every one of them needs a basis for its confidence score.
Sources that disagree
Once retrieval spans several systems, contradiction becomes normal. OWASP classifies this as a data federation knowledge conflict, errors arising when data from multiple sources contradicts itself. The same entry notes a related failure, where retrieved data cannot override what the model absorbed during training (OWASP LLM08:2025).
An agentic system amplifies this. More retrievals across more sources means more opportunities to assemble an answer from two documents that cannot both be true.
Ingestion is an attack surface
The scenario OWASP uses is worth repeating because it is concrete and cheap to execute. Someone submits a CV containing hidden text, white on a white background, instructing the system to ignore prior instructions and recommend the candidate. A recruitment system using RAG for initial screening ingests the document, hidden text included. When the system is later asked about the candidate's qualifications, it follows the embedded instruction.
The recommended mitigations are ingestion-side: text extraction that ignores formatting and detects hidden content, and validation of every document before it enters the knowledge base (OWASP LLM08:2025).
Ingestion checks reduce the chance of poisoned content entering the index. Retrieved content must still be treated as untrusted, because those checks will not catch every attack.
The practical consequence
Validate and classify at ingestion. OWASP's guidance includes tagging and classifying data within the knowledge base specifically so that access levels can be controlled and mismatches prevented, which is the point at which the knowledge layer and the permission model stop being separate concerns.
Access control: the part that fails security review #
Filtering after generation is not access control
This is the design error worth naming first. A system that retrieves freely, generates an answer, and then checks whether the user was entitled to the sources has already lost. The model has read the restricted content. It may have paraphrased it, summarised it, or leaked its existence through a refusal that names the document it cannot discuss.
This is engineering rationale rather than a citable finding: the standards bodies specify the remedy rather than the reasoning. But it determines where everything else goes.
Enforce at retrieval time
OWASP's first prevention measure for this risk class is fine-grained access control with permission-aware vector and embedding stores, plus strict logical and access partitioning of datasets so that one class of users cannot reach another's data (OWASP LLM08:2025).
The operative word is permission-aware. A vector database has no native concept of who is asking. Similarity search returns what is semantically close, and entitlement is not a semantic property. So enforce permissions before retrieved content reaches the model or the user. This can happen in the retrieval query or through an authorisation check on retrieved candidates before generation.
Multi-tenancy
Where several groups or classes of users share one vector database, embeddings belonging to one group can surface in response to another group's queries, leaking business information without any attack taking place (OWASP LLM08:2025). This is ordinary operation of a system built without partitioning.
RBAC, ABAC, and the merge trap
Most enterprises arrive at a RAG project with more than one authorisation model already in place. Attribute-based access control determines authorisation by evaluating attributes of the subject, the object, the requested operation, and in some cases environmental conditions, against policies expressed in terms of those attributes (NIST SP 800-162). Role-based control, by contrast, resolves to what role you hold.
The instinct when data spans several governed systems is to merge those models into one policy layer for the RAG system to consult. Recent work argues against it: merging heterogeneous authorisation models such as RBAC and ABAC frequently produces compliance burden and unintended data exposure. The alternative the authors implement is real-time permission validation against the providers' own IAM endpoints, with no policy merging at all, which keeps each source's governance boundary intact (Jeong & Lee, 2025).
For a mid-market team, the practical reading is that the cheapest correct design is usually the one that asks the source system rather than the one that replicates the source system's rules. On the other hand, checking the source system's IAM avoids maintaining a second copy of its rules, but makes retrieval dependent on that system's availability and response time.
The index is a copy of your data
It is tempting to treat a vector store as a derived artifact with a lower classification than the documents it was built from. It is not. Attackers can invert embeddings and recover significant amounts of the source information (OWASP LLM08:2025).
Whatever classification applies to the source documents applies to the index built from them, along with the backups, the replicas, and the copy in the staging environment.
What agency multiplies
Each of the risks above exists in conventional RAG. Agency multiplies them, because one user question becomes many retrievals across many sources, each one an enforcement point and each one a potential disclosure.
That is also why logging changes character. OWASP recommends maintaining detailed, immutable logs of retrieval activity so suspicious behaviour can be detected and answered (OWASP LLM08:2025). In a single-shot pipeline you can reconstruct what happened from the request. In an agentic one, the retrieval trace is the record of what the system saw on a user's behalf, and if it is not logged, no audit can answer the question.
Before you build: five decisions #
- Which of the four dimensions are you actually choosing? Agent cardinality, control structure, autonomy, knowledge representation. Naming a framework is not a design decision; committing on these is.
- Where does judgement live? Inside the generator, as in Self-RAG, or in a separate evaluator, as in CRAG. The second is cheaper to swap and cheaper to run; the first is more integrated.
- What must exist at ingestion? Record source, date, classification, and a stable document identifier during ingestion. Keep those attributes updateable as documents and permissions change.
- Where does the permission check sit relative to the retriever? If the answer is anywhere after it, the design is not finished.
- What survives an audit? If someone asks in twelve months which documents the system read on behalf of a specific user, the retrieval log has to be able to answer.
Talk it through #
If you are scoping an agentic RAG system and want a second opinion on where enforcement should sit or which architecture fits your constraints, schedule a meeting. For an example in production, see the agentic RAG system we built for clinical triage guidelines, or read about our RAG development services.


