AI & Automation13 min read

Enterprise RAG Architecture: From Retrieval to Governed AI Agents

A production enterprise RAG architecture connects authenticated identity, authorized retrieval, evaluated answers, bounded actions, and observable failure behavior, not only a model and vector database.

Too technical? Pick your depth.

Same topic, explained for where you are — from a first-timer to a working specialist.

Enterprise retrieval-augmented generation (RAG) architecture is the governed path from an authenticated user request to authorized evidence, an evaluated answer or bounded action, and a traceable fallback. A production system must manage identity, permissions, source freshness, retrieval quality, citations, model behavior, tools, latency, cost, and incidents as separate concerns. A model plus a vector database is only a prototype. RAG can reduce unsupported answers when retrieval and abstention are measured, but it cannot make probabilistic output error-free or make an agent safe by default.

Why a PDF chatbot is not a production RAG system

The foundational RAG paper combined a language model's parametric memory with retrieved non-parametric memory for knowledge-intensive tasks. Enterprise architecture adds concerns that the experiment did not claim to solve: changing permissions, source ownership, deletions, operational reliability, and business actions. The original research is useful context, not a production guarantee.

A prototype typically proves that a model can retrieve a passage and answer a sample question. A production system must also prove that it retrieves the right authorized passage, declines when evidence is insufficient, survives dependency failures, and records enough evidence to investigate an incident.

Microsoft's current RAG design and evaluation guide separates ingestion, chunking, enrichment, embedding, search, and response evaluation. That separation matters because a fluent answer can hide a retrieval failure, and a relevant passage can still be unauthorized or stale.

The enterprise RAG request path

Treat the request path as a chain of trust boundaries, not as one model call:

TEXT
authenticated principal
        ↓
server-derived tenant, roles, purpose, and policy context
        ↓
query validation and routing
        ↓
authorized retrieval gateway
        ↓
lexical/vector search + metadata filters + reranking
        ↓
context policy: provenance, freshness, conflict, sensitivity
        ↓
answer with citations OR bounded tool proposal
        ↓
verification, trace, timeout, abstention, or human escalation

Each boundary needs an owner and a failure rule. The identity provider authenticates the principal. The authorization service decides what that principal may access or do. The retrieval layer must enforce those decisions before evidence reaches the model. The model may propose an answer or action, but deterministic application code must validate privileged operations.

This distinction prevents a common design error: treating the prompt as the security boundary. Prompts influence behavior; they do not replace server-side authorization.

Standard RAG versus agentic RAG

Use the least dynamic architecture that can solve the approved task. Microsoft's agentic RAG guidance recommends standard RAG for a straightforward search against one index and agentic retrieval for cases such as query decomposition, dynamic source selection, iterative retrieval, or retrieval combined with actions.

Decision factorStandard RAGAgentic RAGQualification question
Query shapeOne bounded retrieval pathMulti-step or decomposed investigationDoes the next query depend on an earlier result?
Source selectionSources and indexes are fixed by designA runtime planner selects among approved toolsMust the system choose a source dynamically?
Tool useNone, or a deterministic post-answer workflowRetrieval and business tools can be invoked in a loopDoes tool choice require model reasoning?
Latency and costEasier to budget and cacheEach reasoning and tool step adds variable workIs the extra flexibility valuable enough to measure?
ReliabilityFewer states and simpler fallbackMore failure states, loops, and partial resultsCan the workflow stop safely at every step?
ObservabilityTrace one fixed sequenceTrace plans, tool choices, iterations, and budgetsCan reviewers reconstruct why each tool ran?

Agentic RAG is not an upgrade tier. It is a different risk and operating model. If a fixed pipeline answers the use case, its predictability is an advantage. If an agent is justified, bound its available tools, iteration budget, time budget, data scope, and approval requirements.

Build a governed knowledge lifecycle

Retrieval quality starts before query time. Every indexed unit should preserve its source, owner, version, effective date, classification, permission reference, and deletion state. The ingestion pipeline then needs explicit behavior for:

  1. Parsing and segmentation: preserve headings, tables, identifiers, and surrounding context needed to interpret a passage.
  2. Provenance: keep a resolvable link from each retrieved unit to the authoritative source and version.
  3. Permission propagation: update retrieval controls when groups, document ACLs, or tenant relationships change.
  4. Freshness: define how source changes trigger reprocessing and how delayed or failed updates become visible.
  5. Deletion: remove or tombstone deleted material from every index, cache, and derived store covered by policy.
  6. Reconciliation and rollback: compare the index with its sources, report drift, and restore a known version after a bad ingestion.

A knowledge graph can help when entity resolution, provenance relationships, or traversal across connected records improves the task. It is not mandatory for every RAG system. Our guide to semantic entity graphs explains when a governed relationship model adds value beyond vector similarity.

Enforce permissions before retrieval reaches the model

Five controls are related but not interchangeable:

  • Authentication establishes who or what made the request.
  • Authorization decides which data and actions that principal may access.
  • Tenant isolation defines whether tenants use separate stores, shared stores, or a combination.
  • Retrieval filtering applies the derived policy to every search request.
  • Encryption protects data in transit or at rest; it does not decide whether a decrypted record is authorized.

Microsoft's secure multitenant RAG architecture describes both store-per-tenant and shared-store approaches. In either topology, identity and authorization context must flow through the request chain, and shared-store retrieval needs a tenant discriminator plus any applicable user-level filters before grounding data is passed to the model. A namespace or metadata field alone is not a cryptographic boundary.

The server, not the caller, must derive retrieval scope from a validated principal:

PYTHON
async def retrieve_authorized_context(http_request, query: str):
    principal = await identity_provider.verify(http_request.authorization)
    policy = await authorization_service.resolve(
        subject=principal.subject,
        tenant=principal.tenant,
        roles=principal.roles,
        purpose=application_policy.purpose_for(http_request.route),
    )

    if not policy.may_search:
        raise AccessDenied()

    # The client cannot supply or widen these filters.
    scope = policy.to_retrieval_scope()
    candidates = await retrieval_gateway.search(
        query=validate_query(query),
        authorized_scope=scope,
        retrieval_profile=policy.retrieval_profile,
    )

    permitted = [item for item in candidates if policy.allows(item)]
    return reranker.rank(query=query, documents=permitted)

This pseudocode illustrates the trust boundary, not a complete security implementation. Production code still needs token validation, policy-version handling, failure-closed behavior, audit controls, and tests against the actual identity and data platforms.

Design retrieval as an evidence pipeline

A practical retrieval pipeline usually combines several signals:

  • validate and normalize the query without discarding meaningful identifiers;
  • apply tenant, user, region, language, date, and classification filters derived from policy;
  • combine lexical search for exact codes and names with vector search for semantic similarity where evaluation supports it;
  • rerank only the authorized candidate set;
  • assemble enough evidence to answer without flooding the model with irrelevant context;
  • preserve source IDs, versions, dates, and scores for citations and diagnostics.

OpenAI's current Retrieval API guide documents file-attribute filtering and configurable hybrid ranking as one managed implementation. These features show what a retrieval component can expose; they do not prove that a default configuration fits a particular corpus. Test query rewriting, filters, ranking, and context size against a versioned evaluation set.

Knowledge conflicts also need policy. If two approved sources disagree, prefer an explicitly authoritative source, show the conflict, or escalate. Do not let the model quietly choose the most fluent passage.

Separate answers from actions

An answer path should cite evidence the current user can open. When the evidence is missing, stale, conflicting, or below the approved retrieval gate, the correct result may be a clarifying question, an abstention, or escalation instead of a completed sentence.

An action path needs more controls. Treat retrieved text as untrusted data, validate typed tool inputs, re-check authorization at execution time, use least-privilege service identities, make retryable writes idempotent, and require approval for high-impact operations. OWASP's prompt-injection guidance states that RAG does not fully mitigate prompt injection and recommends least privilege plus human approval for high-risk actions.

The model can propose “update this CRM record.” Application code should decide whether the action is allowed, show the proposed change when review is required, execute it once, and record the result. Our ERP and CRM integration service covers the deterministic middleware boundary behind those tools.

Use a production contract, not a generic layer count

The production contract below maps each quality domain to ownership, a measurable gate, failure behavior, and evidence. Thresholds are deliberately project-specific: set them from the use case, risk, evaluation set, and operating budget rather than copying a universal number.

GateAccountable ownerMeasurable acceptance gateFailure behaviorEvidence artifact
RetrievalSearch/data engineeringRequired evidence retrieval, ranking quality, irrelevant-context rate, freshness, and coverage meet approved thresholds by query classRetry an approved fallback, ask for clarification, or return no evidenceVersioned query set, relevance judgments, index/config version, retrieval report
AnswerAI/application engineering with domain reviewerGroundedness, completeness, citation resolution, abstention, and task correctness meet approved thresholdsDecline, expose uncertainty, or escalate; never invent missing supportFrozen answer set, reviewer rubric, prompt/model version, adjudication log
AuthorizationIdentity/security ownerApproved positive cases pass and no unauthorized evidence appears in the adversarial access suiteFail closed, log the policy decision, and alert on suspected leakagePermission matrix, policy version, access-test results, audit record
OperationsPlatform/SRE ownerFreshness, latency, availability, cost, timeout, and recovery behavior meet workflow SLOsDegrade to search, queue safely, stop the action, or route to a humanTrace schema, dashboards, runbooks, incident test, rollback record

Evaluate the gates separately and then test the complete user journey. NIST's Generative AI Profile frames risk work across design, development, use, and evaluation; it is a voluntary risk-management resource, not a certification of a specific RAG implementation.

Define failure behavior before launch

Failure modeSafe defaultWhat to verify
No evidenceAbstain or ask a clarifying questionNo unsupported answer is presented as sourced
Conflicting evidenceShow the conflict or prefer a governed authoritative sourceSource precedence and dates are visible
Stale contentWarn, block the workflow, or use an approved current sourceFreshness state follows the source SLA
Unauthorized contentExclude before generation and fail closedAdversarial tenant, role, and document tests pass
Tool failureDo not claim success; return a recoverable statusRetries are bounded and side effects are idempotent
TimeoutStop remaining work and use the approved fallbackPartial results are not presented as complete
Budget exhaustionEnd the reasoning loop and escalate or simplifyToken, tool-call, and spend limits are enforced

OWASP's vector and embedding risk guidance highlights access leakage, cross-context leakage, poisoning, and the need for permission-aware stores, source validation, monitoring, and logs. Use that guidance to seed threat scenarios, then adapt the tests to the actual corpus and tenancy model.

Operate the system as a changing product

Trace enough to reconstruct a response without copying sensitive content into every log. Useful trace fields include policy and index versions, selected source IDs, retrieval and reranking decisions, model and prompt versions, tool names, approvals, timings, costs, fallback state, and user feedback. Apply access controls, retention, redaction, and integrity protections to telemetry itself.

Re-run evaluation after changes to sources, parsers, chunking, embeddings, search settings, prompts, models, tools, permissions, or policies. Monitor results by query class, language, tenant pattern, and source, not only as one average that can hide a severe subgroup failure.

For EU deployments, regulatory and data-protection obligations depend on the system, its role, and its use. Treat the official EU AI Act text as an architecture input and obtain legal and security review; this guide is not legal advice.

When RAG is the wrong architecture

Do not add RAG when an authorized database query, rules engine, or ordinary search interface can produce the required result more reliably. RAG is also a poor fit when nobody owns source quality, permissions cannot be propagated, the task has no testable acceptance criteria, or the workflow cannot tolerate probabilistic output.

Long context may be enough for a small, stable evidence set. Fine-tuning may help change style or task behavior, but it is not a substitute for current, attributable enterprise facts or server-side access control. A full RAG-versus-fine-tuning decision deserves its own evaluation rather than a slogan.

If the architecture choice is still open, use the enterprise RAG system selection guide. If the main uncertainty is delivery scope, the fixed-price engineering guide explains why experimental AI work often needs a bounded discovery or evaluation phase first.

The practical next step

Before choosing a vector database or agent framework, write the production contract for one use case. Name the principal, permitted sources, expected evidence, evaluation queries, unsafe outcomes, action boundary, fallback, owners, and proof required at acceptance. Then prototype the smallest architecture that can pass those gates.

AppWebSeo's enterprise RAG architecture and pipeline work connects retrieval, authorization, evaluation, observability, and bounded integrations in one reviewable scope. The useful deliverable is not an “autonomous” demo; it is evidence that the approved workflow behaves correctly, including when it cannot complete the task.

Sources and verification notes

Sources were rechecked on 2026-08-28 (Europe/Vienna). Product documentation describes current platform patterns and features, not universal performance. Security and risk frameworks provide design inputs, not proof that an implementation is secure or compliant.

A

AppWebSeo

SEO & Engineering Editorial Team

Specializing in high-performance web systems, Generative Engine Optimization, and enterprise AI architecture at AppWebSeo.

Share this Technical Breakdown

Forward this architecture guide to your team, colleagues, or engineering network.

Transform These Insights into Production Architecture

Schedule a technical architecture review with our senior engineering team.

All Topics