Enterprise retrieval-augmented generation (RAG) architecture is the governed path from an authenticated user request to authorized evidence, an evaluated answer or bounded action, and a traceable fallback. A production system must manage identity, permissions, source freshness, retrieval quality, citations, model behavior, tools, latency, cost, and incidents as separate concerns. A model plus a vector database is only a prototype. RAG can reduce unsupported answers when retrieval and abstention are measured, but it cannot make probabilistic output error-free or make an agent safe by default.
Why a PDF chatbot is not a production RAG system
The foundational RAG paper combined a language model's parametric memory with retrieved non-parametric memory for knowledge-intensive tasks. Enterprise architecture adds concerns that the experiment did not claim to solve: changing permissions, source ownership, deletions, operational reliability, and business actions. The original research is useful context, not a production guarantee.
A prototype typically proves that a model can retrieve a passage and answer a sample question. A production system must also prove that it retrieves the right authorized passage, declines when evidence is insufficient, survives dependency failures, and records enough evidence to investigate an incident.
Microsoft's current RAG design and evaluation guide separates ingestion, chunking, enrichment, embedding, search, and response evaluation. That separation matters because a fluent answer can hide a retrieval failure, and a relevant passage can still be unauthorized or stale.
The enterprise RAG request path
Treat the request path as a chain of trust boundaries, not as one model call:
authenticated principal
↓
server-derived tenant, roles, purpose, and policy context
↓
query validation and routing
↓
authorized retrieval gateway
↓
lexical/vector search + metadata filters + reranking
↓
context policy: provenance, freshness, conflict, sensitivity
↓
answer with citations OR bounded tool proposal
↓
verification, trace, timeout, abstention, or human escalationEach boundary needs an owner and a failure rule. The identity provider authenticates the principal. The authorization service decides what that principal may access or do. The retrieval layer must enforce those decisions before evidence reaches the model. The model may propose an answer or action, but deterministic application code must validate privileged operations.
This distinction prevents a common design error: treating the prompt as the security boundary. Prompts influence behavior; they do not replace server-side authorization.
Standard RAG versus agentic RAG
Use the least dynamic architecture that can solve the approved task. Microsoft's agentic RAG guidance recommends standard RAG for a straightforward search against one index and agentic retrieval for cases such as query decomposition, dynamic source selection, iterative retrieval, or retrieval combined with actions.
| Decision factor | Standard RAG | Agentic RAG | Qualification question |
|---|---|---|---|
| Query shape | One bounded retrieval path | Multi-step or decomposed investigation | Does the next query depend on an earlier result? |
| Source selection | Sources and indexes are fixed by design | A runtime planner selects among approved tools | Must the system choose a source dynamically? |
| Tool use | None, or a deterministic post-answer workflow | Retrieval and business tools can be invoked in a loop | Does tool choice require model reasoning? |
| Latency and cost | Easier to budget and cache | Each reasoning and tool step adds variable work | Is the extra flexibility valuable enough to measure? |
| Reliability | Fewer states and simpler fallback | More failure states, loops, and partial results | Can the workflow stop safely at every step? |
| Observability | Trace one fixed sequence | Trace plans, tool choices, iterations, and budgets | Can reviewers reconstruct why each tool ran? |
Agentic RAG is not an upgrade tier. It is a different risk and operating model. If a fixed pipeline answers the use case, its predictability is an advantage. If an agent is justified, bound its available tools, iteration budget, time budget, data scope, and approval requirements.
Build a governed knowledge lifecycle
Retrieval quality starts before query time. Every indexed unit should preserve its source, owner, version, effective date, classification, permission reference, and deletion state. The ingestion pipeline then needs explicit behavior for:
- Parsing and segmentation: preserve headings, tables, identifiers, and surrounding context needed to interpret a passage.
- Provenance: keep a resolvable link from each retrieved unit to the authoritative source and version.
- Permission propagation: update retrieval controls when groups, document ACLs, or tenant relationships change.
- Freshness: define how source changes trigger reprocessing and how delayed or failed updates become visible.
- Deletion: remove or tombstone deleted material from every index, cache, and derived store covered by policy.
- Reconciliation and rollback: compare the index with its sources, report drift, and restore a known version after a bad ingestion.
A knowledge graph can help when entity resolution, provenance relationships, or traversal across connected records improves the task. It is not mandatory for every RAG system. Our guide to semantic entity graphs explains when a governed relationship model adds value beyond vector similarity.
Enforce permissions before retrieval reaches the model
Five controls are related but not interchangeable:
- Authentication establishes who or what made the request.
- Authorization decides which data and actions that principal may access.
- Tenant isolation defines whether tenants use separate stores, shared stores, or a combination.
- Retrieval filtering applies the derived policy to every search request.
- Encryption protects data in transit or at rest; it does not decide whether a decrypted record is authorized.
Microsoft's secure multitenant RAG architecture describes both store-per-tenant and shared-store approaches. In either topology, identity and authorization context must flow through the request chain, and shared-store retrieval needs a tenant discriminator plus any applicable user-level filters before grounding data is passed to the model. A namespace or metadata field alone is not a cryptographic boundary.
The server, not the caller, must derive retrieval scope from a validated principal:
async def retrieve_authorized_context(http_request, query: str):
principal = await identity_provider.verify(http_request.authorization)
policy = await authorization_service.resolve(
subject=principal.subject,
tenant=principal.tenant,
roles=principal.roles,
purpose=application_policy.purpose_for(http_request.route),
)
if not policy.may_search:
raise AccessDenied()
# The client cannot supply or widen these filters.
scope = policy.to_retrieval_scope()
candidates = await retrieval_gateway.search(
query=validate_query(query),
authorized_scope=scope,
retrieval_profile=policy.retrieval_profile,
)
permitted = [item for item in candidates if policy.allows(item)]
return reranker.rank(query=query, documents=permitted)This pseudocode illustrates the trust boundary, not a complete security implementation. Production code still needs token validation, policy-version handling, failure-closed behavior, audit controls, and tests against the actual identity and data platforms.
Design retrieval as an evidence pipeline
A practical retrieval pipeline usually combines several signals:
- validate and normalize the query without discarding meaningful identifiers;
- apply tenant, user, region, language, date, and classification filters derived from policy;
- combine lexical search for exact codes and names with vector search for semantic similarity where evaluation supports it;
- rerank only the authorized candidate set;
- assemble enough evidence to answer without flooding the model with irrelevant context;
- preserve source IDs, versions, dates, and scores for citations and diagnostics.
OpenAI's current Retrieval API guide documents file-attribute filtering and configurable hybrid ranking as one managed implementation. These features show what a retrieval component can expose; they do not prove that a default configuration fits a particular corpus. Test query rewriting, filters, ranking, and context size against a versioned evaluation set.
Knowledge conflicts also need policy. If two approved sources disagree, prefer an explicitly authoritative source, show the conflict, or escalate. Do not let the model quietly choose the most fluent passage.
Separate answers from actions
An answer path should cite evidence the current user can open. When the evidence is missing, stale, conflicting, or below the approved retrieval gate, the correct result may be a clarifying question, an abstention, or escalation instead of a completed sentence.
An action path needs more controls. Treat retrieved text as untrusted data, validate typed tool inputs, re-check authorization at execution time, use least-privilege service identities, make retryable writes idempotent, and require approval for high-impact operations. OWASP's prompt-injection guidance states that RAG does not fully mitigate prompt injection and recommends least privilege plus human approval for high-risk actions.
The model can propose “update this CRM record.” Application code should decide whether the action is allowed, show the proposed change when review is required, execute it once, and record the result. Our ERP and CRM integration service covers the deterministic middleware boundary behind those tools.
Use a production contract, not a generic layer count
The production contract below maps each quality domain to ownership, a measurable gate, failure behavior, and evidence. Thresholds are deliberately project-specific: set them from the use case, risk, evaluation set, and operating budget rather than copying a universal number.
| Gate | Accountable owner | Measurable acceptance gate | Failure behavior | Evidence artifact |
|---|---|---|---|---|
| Retrieval | Search/data engineering | Required evidence retrieval, ranking quality, irrelevant-context rate, freshness, and coverage meet approved thresholds by query class | Retry an approved fallback, ask for clarification, or return no evidence | Versioned query set, relevance judgments, index/config version, retrieval report |
| Answer | AI/application engineering with domain reviewer | Groundedness, completeness, citation resolution, abstention, and task correctness meet approved thresholds | Decline, expose uncertainty, or escalate; never invent missing support | Frozen answer set, reviewer rubric, prompt/model version, adjudication log |
| Authorization | Identity/security owner | Approved positive cases pass and no unauthorized evidence appears in the adversarial access suite | Fail closed, log the policy decision, and alert on suspected leakage | Permission matrix, policy version, access-test results, audit record |
| Operations | Platform/SRE owner | Freshness, latency, availability, cost, timeout, and recovery behavior meet workflow SLOs | Degrade to search, queue safely, stop the action, or route to a human | Trace schema, dashboards, runbooks, incident test, rollback record |
Evaluate the gates separately and then test the complete user journey. NIST's Generative AI Profile frames risk work across design, development, use, and evaluation; it is a voluntary risk-management resource, not a certification of a specific RAG implementation.
Define failure behavior before launch
| Failure mode | Safe default | What to verify |
|---|---|---|
| No evidence | Abstain or ask a clarifying question | No unsupported answer is presented as sourced |
| Conflicting evidence | Show the conflict or prefer a governed authoritative source | Source precedence and dates are visible |
| Stale content | Warn, block the workflow, or use an approved current source | Freshness state follows the source SLA |
| Unauthorized content | Exclude before generation and fail closed | Adversarial tenant, role, and document tests pass |
| Tool failure | Do not claim success; return a recoverable status | Retries are bounded and side effects are idempotent |
| Timeout | Stop remaining work and use the approved fallback | Partial results are not presented as complete |
| Budget exhaustion | End the reasoning loop and escalate or simplify | Token, tool-call, and spend limits are enforced |
OWASP's vector and embedding risk guidance highlights access leakage, cross-context leakage, poisoning, and the need for permission-aware stores, source validation, monitoring, and logs. Use that guidance to seed threat scenarios, then adapt the tests to the actual corpus and tenancy model.
Operate the system as a changing product
Trace enough to reconstruct a response without copying sensitive content into every log. Useful trace fields include policy and index versions, selected source IDs, retrieval and reranking decisions, model and prompt versions, tool names, approvals, timings, costs, fallback state, and user feedback. Apply access controls, retention, redaction, and integrity protections to telemetry itself.
Re-run evaluation after changes to sources, parsers, chunking, embeddings, search settings, prompts, models, tools, permissions, or policies. Monitor results by query class, language, tenant pattern, and source, not only as one average that can hide a severe subgroup failure.
For EU deployments, regulatory and data-protection obligations depend on the system, its role, and its use. Treat the official EU AI Act text as an architecture input and obtain legal and security review; this guide is not legal advice.
When RAG is the wrong architecture
Do not add RAG when an authorized database query, rules engine, or ordinary search interface can produce the required result more reliably. RAG is also a poor fit when nobody owns source quality, permissions cannot be propagated, the task has no testable acceptance criteria, or the workflow cannot tolerate probabilistic output.
Long context may be enough for a small, stable evidence set. Fine-tuning may help change style or task behavior, but it is not a substitute for current, attributable enterprise facts or server-side access control. A full RAG-versus-fine-tuning decision deserves its own evaluation rather than a slogan.
If the architecture choice is still open, use the enterprise RAG system selection guide. If the main uncertainty is delivery scope, the fixed-price engineering guide explains why experimental AI work often needs a bounded discovery or evaluation phase first.
The practical next step
Before choosing a vector database or agent framework, write the production contract for one use case. Name the principal, permitted sources, expected evidence, evaluation queries, unsafe outcomes, action boundary, fallback, owners, and proof required at acceptance. Then prototype the smallest architecture that can pass those gates.
AppWebSeo's enterprise RAG architecture and pipeline work connects retrieval, authorization, evaluation, observability, and bounded integrations in one reviewable scope. The useful deliverable is not an “autonomous” demo; it is evidence that the approved workflow behaves correctly, including when it cannot complete the task.
Sources and verification notes
Sources were rechecked on 2026-08-28 (Europe/Vienna). Product documentation describes current platform patterns and features, not universal performance. Security and risk frameworks provide design inputs, not proof that an implementation is secure or compliant.
- Lewis et al.: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks: foundational research; its experimental results do not establish enterprise production behavior.
- Microsoft: Design and develop a RAG solution: data/application flows and stage-specific evaluation.
- Microsoft: Develop an agentic RAG solution: standard-versus-agentic fit, tool loops, latency, cost, reliability, and evaluation dimensions.
- Microsoft: Secure multitenant RAG: identity propagation, storage topology, security trimming, and the retrieval API boundary.
- OpenAI: Retrieval API guide: current managed retrieval filters and ranking controls.
- OWASP: LLM01 Prompt Injection: prompt-injection risk and impact-reduction controls.
- OWASP: LLM08 Vector and Embedding Weaknesses: access, leakage, poisoning, validation, and monitoring risks.
- NIST AI 600-1: Generative AI Profile: voluntary lifecycle risk-management guidance.
- EUR-Lex: Regulation (EU) 2024/1689: official regulation; interpretation remains subject to qualified legal review.