GUIDE WHAT IS

What is RAG (retrieval-augmented generation)?

RAG connects AI-generated answers to external knowledge without retraining the model. Explore its architecture, implementation choices, security risks, and evaluation criteria.

If you are asking “what is rag (retrieval-augmented generation)?”, the practical answer is: RAG is a way to give an AI model relevant reference material before it answers a question. Instead of relying only on knowledge learned during training, the system retrieves information from sources such as product documentation, internal policies, or research papers and includes it in the model’s input.

For decision-makers, RAG offers a path to AI assistants that use current, organization-specific information. For practitioners, it is an application architecture combining search, data processing, access controls, and language-model generation. Its value depends on how reliably those components work together—not simply on which model you choose.

What RAG means—and what it does not

Retrieval-augmented generation has two core operations:

  • Retrieval: Find information relevant to the user’s request.
  • Generation: Ask a language model to produce an answer using that information.

Imagine an engineer asking, “How do we rotate production database credentials?” A conventional chatbot might describe a generic procedure. A RAG assistant searches approved operational runbooks, retrieves the relevant credential-rotation steps, and generates an answer grounded in those steps.

The model’s parameters do not need to change when the runbook changes. Instead, the retrieval system must pick up the updated content.

RAG is not a particular database, model, or vendor product. It also does not guarantee factual accuracy. A system can retrieve an outdated document, miss an important exception, or generate a claim unsupported by its sources.

Grounding means supplying evidence; it does not mean guaranteeing that the model uses that evidence correctly.

How retrieval-augmented generation works

Most RAG implementations have two connected workflows: preparing knowledge and answering requests.

Preparing the knowledge base

The preparation workflow makes source material searchable:

  • Ingest content from repositories, content-management systems, file stores, or databases.
  • Parse and normalize it, preserving headings, tables, code blocks, and useful structure.
  • Split it into retrievable units, commonly called chunks.
  • Attach metadata, such as document owner, source URL, version, access permissions, and update date.
  • Build indexes for keyword search, vector search, or both.

For vector search, an embedding model converts content into numerical representations. Content with similar meaning should have representations that are close under the index’s similarity measure.

A vector database can store and search these representations, but RAG does not require a dedicated vector database. Existing search engines, relational databases with vector extensions, and other retrieval systems can support it.

Answering a user’s question

At query time, a typical RAG pipeline:

  1. Receives the question and identifies the user.
  2. Rewrites or expands the query if needed.
  3. Searches authorized content.
  4. Reranks candidate passages for relevance.
  5. Packs selected passages into the model’s context.
  6. Generates an answer with source references.
  7. Checks the output or declines when evidence is insufficient.

Consider a support assistant answering whether a product feature works on a particular software version. Good retrieval should locate the compatibility matrix for that version—not merely a marketing page that mentions the feature.

The generated answer is only the visible end of this pipeline. Many apparent “model mistakes” begin with missing, poorly parsed, or incorrectly selected evidence.

These approaches solve overlapping but different problems.

ApproachBest suited toMain limitation
RAGAnswering from changing, private, or attributable knowledgeDepends on retrieval quality and source governance
Fine-tuningShaping behavior, style, formatting, or task performanceNot a dependable replacement for a current knowledge store
Long-context promptingWorking directly with a bounded collection of supplied materialLarger inputs can increase cost and latency; relevant details may still be missed
Traditional searchFinding documents for people to inspectUsers must interpret and combine results themselves

RAG and fine-tuning are complementary. A model might be fine-tuned to follow a structured response format while RAG supplies the facts.

Long-context models can reduce the need for elaborate retrieval when the document collection is small enough. However, putting every available document into each request can be wasteful, and access control still matters.

Traditional search may be the better product when users mainly want original documents. Adding generation is worthwhile when synthesis, comparison, or conversational explanation materially improves the task.

When RAG is a good fit

RAG is especially useful when answers depend on information that is private, changes regularly, or needs attribution.

Common applications include:

  • Internal knowledge assistants: Policies, onboarding material, and operational procedures.
  • Customer support: Product documentation, troubleshooting guides, and release notes.
  • Developer tools: Repository documentation, architecture decisions, and service runbooks.
  • Research workflows: Discovering and comparing evidence across an approved collection.
  • Sales enablement: Approved product positioning and current technical specifications.

A promising use case has a clear audience, an authoritative corpus, and questions that the corpus can actually answer.

RAG alone is a weaker fit for exact calculations, transaction execution, or rapidly changing operational state. For “What is this customer’s current account balance?”, an authorized database query or API call is usually more appropriate than searching embedded document snapshots.

A hybrid assistant can retrieve policy explanations while calling a live service for account data.

Key RAG architecture decisions

Keyword, vector, or hybrid retrieval

Keyword search is strong for exact identifiers: error codes, product names, legal clauses, and function names.

Vector search helps when the question and source use different wording. “Leaving the company” may match an “employee offboarding” policy even without shared keywords.

Hybrid search combines these signals. It is often a sensible starting point for enterprise content containing both natural language and precise terminology.

A reranker can improve the order of retrieved candidates, but adds computation and potentially another service dependency. Measure whether it improves answer quality enough to justify its latency and cost.

Chunking and context assembly

Chunk boundaries should follow meaning where possible. A paragraph detached from its heading may lose scope; half a table may become misleading.

Evaluate chunks against concrete questions:

  • Does each chunk preserve the information needed to interpret it?
  • Are exceptions separated from the rules they qualify?
  • Can the system retrieve adjacent material when necessary?
  • Are source identifiers stable enough for reliable citations?

Small chunks improve retrieval granularity but may lack context. Large chunks preserve context but can dilute relevance.

There is no universally correct chunk size. Test alternatives on your documents rather than copying a default configuration.

Tools and frameworks

Named options include:

  • PostgreSQL with pgvector: Useful when the team already operates PostgreSQL and wants relational data and vector search together.
  • Elasticsearch and Azure AI Search: Search platforms supporting lexical and vector retrieval capabilities.
  • Pinecone, Weaviate, and Qdrant: Products commonly used for vector-oriented retrieval.
  • LangChain and LlamaIndex: Frameworks for assembling ingestion, retrieval, and generation workflows.
  • Amazon Bedrock Knowledge Bases: A managed option for integrating retrieval with supported AWS AI workflows.

The LangChain retrieval documentation explains common retrieval building blocks. Microsoft’s Azure AI Search RAG overview describes search’s role in a RAG architecture.

Choose based on filtering, tenancy, update behavior, operational skills, and evaluation results—not a feature checklist alone.

How to build a RAG system step by step

1. Define the task and acceptable failure behavior

Start with a narrow objective, such as answering administrator questions about one product.

Define whether the assistant should provide instructions, summarize evidence, or locate documents. Specify when it must abstain or escalate to a person.

For consequential tasks, “No supported answer found” is often better than a plausible guess.

2. Audit sources and permissions

Identify authoritative sources, document owners, update frequency, and conflicting versions. Exclude material that should not influence answers.

Preserve access-control information during ingestion. Plan for deletions and permission changes, not just new uploads.

3. Create a representative evaluation set

Collect realistic questions with expected supporting sources and acceptable answers. Include:

  • Exact-name and paraphrased questions.
  • Questions requiring multiple documents.
  • Outdated or contradictory content.
  • Questions with no answer in the corpus.
  • Requests for information the user cannot access.

Keep a held-out set for regression testing rather than repeatedly optimizing against every example.

4. Build a simple baseline

Begin with straightforward ingestion, retrieval, and generation. Require answers to cite supplied source identifiers.

Record retrieved passages and the final context, with appropriate privacy controls. Without these traces, diagnosing failures becomes guesswork.

5. Improve the weakest component

If the right document is missing, improve ingestion or retrieval. If it is found but excluded from context, adjust ranking or context assembly. If the evidence is present but the answer is wrong, revise generation instructions or model selection.

Avoid adding query rewriting, reranking, and agentic loops simultaneously. Otherwise, attribution of improvements becomes difficult.

6. Pilot and operate it as a maintained service

Pilot with a defined user group. Track unresolved questions, unsupported claims, stale sources, and latency.

Assign ownership for the corpus, retrieval pipeline, evaluation suite, and incident response. RAG quality degrades when source maintenance has no owner.

How to evaluate RAG quality

A fluent answer is not sufficient evidence of success. Evaluate retrieval and generation separately.

CriterionWhat to inspectWhy it matters
Retrieval recallWhether expected evidence appears among retrieved resultsMissing evidence limits answer quality
Context relevanceHow much selected context actually helps answerIrrelevant passages waste context and can distract
Answer correctnessWhether the response satisfies the question accuratelyGrounded wording can still produce a wrong conclusion
FaithfulnessWhether claims are supported by supplied evidenceDetects unsupported additions
Citation qualityWhether each citation supports its associated claimA link alone is not proof
Abstention behaviorWhether unanswerable questions trigger a safe responseReduces confident guessing
Access isolationWhether unauthorized content stays inaccessiblePrevents information disclosure
FreshnessWhether updates and deletions become effective as requiredKeeps answers aligned with current sources

Also measure end-to-end latency, including slower-tail requests, and cost per successfully resolved task.

Automated evaluators can accelerate testing, but model-based judges have biases and can miss errors. Use human review for consequential cases and validate automated scoring against expert judgments.

Security, cost, and operational trade-offs

Treat retrieved content as untrusted input

Documents can contain malicious instructions, whether deliberately planted or accidentally copied. A retrieved page saying “ignore previous instructions” must remain source content, not become an instruction to the assistant.

Enforce authorization before content enters the model context. Do not retrieve confidential passages and rely on prompting to hide them.

Use tenant isolation, permission-aware retrieval, restricted tool access, and tests for cross-user leakage. Check logs, caches, and citation previews too: these can expose information even when the final answer does not.

The OWASP guidance on prompt injection explains why instruction-bearing external content creates a distinct application-security risk.

Budget for the entire pipeline

RAG costs include parsing, embeddings, indexing, storage, retrieval, reranking, model tokens, evaluation, and operations.

More retrieved passages can increase input-token costs without improving answers. More sophisticated retrieval can improve accuracy while increasing latency.

Caching can reduce repeated work, but cached results must respect permissions and freshness. Re-embedding a large corpus after changing embedding models can also become a substantial migration task.

Common RAG mistakes

  • Treating all documents as equally authoritative: Establish precedence for approved policies, current versions, and historical material.
  • Assuming more context is always better: Select evidence deliberately instead of flooding the prompt.
  • Using vector search for every query: Exact identifiers often benefit from lexical retrieval.
  • Confusing citations with verification: Check that cited passages actually support the claims.
  • Ignoring deletion workflows: Removed or restricted documents must stop appearing in indexes and caches.
  • Using prompts as access controls: Enforce permissions in application and retrieval infrastructure.
  • Skipping no-answer tests: An assistant needs to recognize when its sources cannot resolve a question.
  • Adding agents too early: Multi-step search can help complex questions but also introduces more failure paths.

Frequently asked questions

Does RAG eliminate hallucinations?

No. RAG can reduce unsupported answers by supplying relevant evidence, but retrieval and generation can both fail. Reliable systems combine source governance, evaluation, citation checks, and explicit abstention behavior.

Does RAG require a vector database?

No. RAG can use keyword search, SQL queries, APIs, vector retrieval, or a combination. A vector database is useful for some workloads, but the requirement is relevant evidence—not a specific storage technology.

Is RAG better than fine-tuning?

Neither is universally better. Use RAG to supply accessible, current knowledge; consider fine-tuning to improve behavior or specialized task performance. Some applications benefit from both, while others need neither.

How current is a RAG system’s knowledge?

Only as current as its sources and retrieval pipeline. Freshness depends on ingestion schedules, indexing delays, cache invalidation, and deletion handling. Set a freshness requirement and test how quickly actual source changes affect answers.

The practical takeaway

RAG turns a language model into the answer-generation layer of a knowledge system. Its strongest implementations combine trustworthy sources, effective retrieval, enforced permissions, and measurable answer quality.

Start with a bounded use case and a testable baseline. Expand only when evidence shows that the system retrieves the right material, answers faithfully, and handles uncertainty appropriately.

For related explanations of AI, software, and cloud concepts, browse more What is topics.

Have a question about this topic?

Ask the community and get answers from practitioners.

Start a discussion