GUIDE TECH STACK

Tech stack for a RAG application

A production RAG stack needs more than a vector database and an LLM. Learn how to choose ingestion, retrieval, generation, and evaluation tools around your data, security requirements, and operating constraints.

Start with the retrieval problem, not the framework

Choosing a tech stack for a rag application means deciding how your product will find trustworthy information, supply it to a language model, and verify that the resulting answer is useful. Retrieval-augmented generation, or RAG, connects a model to external knowledge at query time. Its quality depends as much on document processing, access control, and search relevance as on the model itself.

For decision-makers, the central question is not “Which vector database is best?” It is “What information must this application retrieve, for whom, under what operating constraints?” An internal policy assistant, a customer-facing support bot, and a technical research tool require different architectures.

A defensible stack starts with a representative evaluation set and the simplest infrastructure that meets its requirements. Add specialized components only when measured failures justify them.

The architecture of a production RAG application

Most production systems have two distinct paths: an ingestion path that prepares knowledge and a request path that answers users.

The ingestion path typically includes:

  • Source connectors for repositories, document stores, websites, or databases.
  • Parsing, OCR, normalization, and duplicate detection.
  • Chunking and metadata enrichment.
  • Embedding generation and indexing.
  • Update, deletion, and permission synchronization.

The request path usually includes:

  • Authentication and query interpretation.
  • Permission-aware retrieval.
  • Optional query rewriting and reranking.
  • Context assembly and model generation.
  • Citation rendering, response validation, and telemetry.

Keep ingestion asynchronous unless the product explicitly requires immediate processing. A document upload should create a tracked job rather than hold an HTTP request open through parsing, embedding, and indexing.

A practical component map looks like this:

LayerRepresentative toolsPrimary selection criterion
Application APIFastAPI, Django, NestJSTeam expertise and integration needs
Workflow coordinationPlain application code, LlamaIndex, LangChain, LangGraphWorkflow complexity and debugging visibility
Background processingCelery, BullMQ, TemporalRetry, scheduling, and durability requirements
ParsingUnstructured, Docling, Apache TikaFidelity on your actual documents
Search and storagePostgreSQL with pgvector, Qdrant, Pinecone, ElasticsearchRetrieval quality, filtering, and operational fit
EmbeddingsOpenAI, Cohere, BGE-family modelsDomain quality, deployment constraints, and cost
RerankingCohere Rerank, BGE rerankersRelevance improvement versus latency
GenerationOpenAI, Anthropic, Gemini, self-hosted open-weight modelsGrounded accuracy, latency, and governance
Evaluation and tracingLangfuse, Phoenix, Ragas, OpenTelemetryFailure diagnosis and reproducible testing

These are alternatives, not a shopping list. A first release rarely needs a separate vendor for every row.

Define requirements before selecting vendors

Write down acceptance criteria before running model or database comparisons.

Corpus characteristics

Determine whether your knowledge consists of prose, scanned PDFs, tables, code, or frequently changing records. A text splitter that works for help-center articles may destroy the relationships inside a financial table.

Capture document volume, expected growth, languages, update frequency, and average document complexity. Measure retrievable chunks, not just files: a long manual can produce far more index entries than hundreds of short FAQs.

Product and operating constraints

Specify:

  • Freshness: How quickly must updates and deletions become searchable?
  • Authorization: Are permissions global, tenant-specific, group-based, or document-specific?
  • Response behavior: Must answers cite sources, return structured data, or decline unsupported questions?
  • Latency: Does the product prioritize rapid interaction or deeper research?
  • Deployment: Can content leave your cloud account or geographic region?
  • Economics: What cost per successful task can the business support?

Avoid assuming that a large context window eliminates retrieval. Sending an entire corpus can increase cost and obscure relevant passages, while leaving permissions and freshness unresolved.

Choose the application and orchestration layer

FastAPI is a natural choice for Python teams working with document processing and machine-learning libraries. NestJS fits TypeScript teams that want a consistent language across application services and the frontend. Existing Django or other mature backends can also host RAG endpoints without a rewrite.

Keep retrieval logic behind explicit interfaces: retrieve, rerank, assemble_context, and generate. This makes components replaceable and tests easier to interpret.

When frameworks help

LlamaIndex provides ingestion and retrieval abstractions useful for knowledge-centric applications. LangChain provides model and tool integrations, while LangGraph is useful when workflows require branching, persistence, or human approval.

Frameworks reduce integration work, but abstractions can hide defaults, retries, and extra model calls. Pin versions and inspect the generated requests and retrieved context.

For a straightforward question-answering endpoint, ordinary application code may be clearer. Use stateful orchestration when the workflow actually needs it—not because the application uses an LLM.

Build ingestion around document fidelity

Document quality often sets the ceiling for answer quality.

Apache Tika is useful for extracting text and metadata from many file formats. Unstructured and Docling support richer document-processing workflows. For difficult scans, compare managed services such as Azure AI Document Intelligence, Amazon Textract, or Google Cloud Document AI against local OCR.

Evaluate parsers on representative failures: multi-column PDFs, repeated headers, nested lists, tables, and low-quality scans.

Chunking and metadata

Start with structure-aware chunks aligned to headings, paragraphs, or logical sections. Test modest overlap where boundaries split meaning. There is no universally correct chunk size: shorter chunks improve targeting, while longer chunks preserve context.

Store metadata alongside each chunk:

  • Stable document and chunk identifiers.
  • Source location, title, and section.
  • Document version and ingestion timestamp.
  • Tenant, authorization attributes, and sensitivity labels.
  • Parser and embedding versions.

Use object storage such as Amazon S3 or compatible storage for originals and parsed artifacts. Keep ingestion state in a transactional database.

Make jobs idempotent so retries do not create duplicate chunks. Updates must remove obsolete content, and deletions must propagate to indexes and caches.

Select embeddings and the retrieval engine together

Embedding models map queries and content into comparable vectors. Choose them using your own retrieval tests, especially when the corpus includes technical terminology, multiple languages, or specialized abbreviations.

Compare hosted embedding APIs with self-hosted models on relevance, throughput, data handling, and total operating cost. Follow each model’s required query/document formatting; some retrieval models use distinct instructions or prefixes.

Changing embedding models generally requires re-embedding the corpus. Record model identifiers and dimensions, and plan migrations with parallel indexes when uninterrupted service matters.

PostgreSQL or a dedicated search system?

PostgreSQL with pgvector is a strong starting point when your team already operates PostgreSQL and values keeping metadata, permissions, and vectors close together. It avoids adding a separate database before there is evidence you need one.

Approximate vector retrieval combined with selective filters needs careful testing. Index configuration, query planning, and scan behavior affect both recall and latency. The official pgvector documentation explains indexing options and filtering considerations.

Choose Qdrant, Pinecone, or another dedicated vector service when vector workloads, scaling requirements, or operational preferences justify it. Evaluate ingestion throughput, filtered retrieval, backup and recovery, and tenant isolation—not just an unfiltered nearest-neighbor benchmark.

Choose Elasticsearch or OpenSearch when lexical search, complex filters, aggregations, and established search operations are central to the product.

Hybrid retrieval and reranking

Semantic search helps with paraphrases. Keyword search helps with exact identifiers, product codes, error messages, and uncommon names.

Hybrid retrieval combines both, often using reciprocal rank fusion to merge rankings without assuming their scores are directly comparable.

A reranker can then reorder a candidate set using a more expensive relevance model. Add it when evaluation shows that relevant passages are retrieved but ranked too low. If the right passage never enters the candidate set, reranking cannot recover it.

Choose the generation model for grounded behavior

Test generation models against the same retrieved evidence. Otherwise, you may attribute a retrieval improvement to the model or blame the model for missing context.

Evaluate:

  • Whether claims are supported by retrieved passages.
  • Whether the model recognizes insufficient or conflicting evidence.
  • Citation correctness.
  • Instruction following and structured-output reliability.
  • Latency under realistic concurrency.
  • Input, output, and deployment costs.

Hosted models from OpenAI, Anthropic, and Google reduce infrastructure work. Self-hosting with tools such as vLLM provides more deployment control, but introduces GPU capacity planning, serving optimization, and model maintenance.

Use a prompt that clearly separates user requests, application instructions, and retrieved material. Require source identifiers for factual claims, then validate that returned identifiers belong to the supplied context.

Streaming improves perceived responsiveness, but it does not reduce retrieval latency or guarantee correctness. Decide whether answers can stream immediately or require validation before display.

Treat security as a retrieval requirement

Authorization must constrain retrieval before content reaches the model. Do not retrieve across tenants and ask the LLM to hide unauthorized information.

Propagate permissions from source systems and define how quickly revocations take effect. For high-sensitivity applications, consider stronger physical separation rather than relying entirely on metadata filters.

Retrieved documents are untrusted input. They may contain instructions designed to override the assistant or exfiltrate data. Delimiters and prompts help communicate boundaries, but they are not complete defenses. Restrict tool permissions, validate actions, and avoid exposing credentials or privileged capabilities to document-driven workflows.

The OWASP guidance for LLM applications is a useful starting point for threat modeling.

Also review provider retention policies, regional processing, encryption, and contractual requirements. Logs and traces can contain sensitive queries and documents; apply redaction, access controls, and retention limits.

Evaluate quality, latency, and cost as one system

Build an evaluation set from realistic questions, including unanswerable requests, ambiguous wording, outdated documents, and permission boundaries.

Separate retrieval evaluation from answer evaluation:

  • Retrieval: Did the evidence needed to answer appear in the candidate set and final context?
  • Generation: Was the answer correct, supported, and appropriately qualified?
  • Product: Did the user complete the intended task?
  • Security: Were unauthorized passages excluded?

Ragas can help implement automated checks; Langfuse and Phoenix can support tracing and evaluation workflows. Model-based judges are useful screening tools, not unquestionable ground truth. Calibrate them against human review.

Trace each stage, including parsing failures, retrieval scores, reranking time, context tokens, and model calls. Follow OpenTelemetry documentation to integrate telemetry with broader application monitoring.

Estimate cost per request as the sum of retrieval, reranking, generation, and allocated infrastructure. Track ingestion separately. Include failed requests and retries: cost per successful answer is more meaningful than cost per model call.

A step-by-step stack selection process

1. Assemble a representative test corpus

Include common content and difficult edge cases. Write questions with known supporting passages, plus questions that should produce no answer.

2. Build the smallest end-to-end baseline

Use your existing backend, one parser, one embedding model, PostgreSQL with pgvector or an existing search service, and one generation model. Implement citations and permission checks immediately.

3. Diagnose failures before adding components

Inspect retrieved passages. Distinguish parsing failures, missing content, poor ranking, insufficient context, and unsupported generation.

4. Run targeted comparisons

Compare chunking strategies first where extraction or context boundaries are failing. Test hybrid search, reranking, and alternative embeddings against the same evaluation set.

5. Test production conditions

Exercise concurrency, large documents, provider timeouts, queue backlogs, permission changes, and deletions. Verify retry limits and recovery procedures.

6. Establish release gates

Require acceptable quality, latency, cost, and authorization results before deployment. Version prompts, parsers, indexes, and models so regressions can be reproduced and rolled back.

Common mistakes and a sensible default stack

The most expensive mistakes are often architectural:

  • Starting with agents: Unnecessary model-driven loops add latency and make failures harder to reproduce.
  • Choosing by benchmark alone: Public leaderboards may not reflect your corpus or filters.
  • Ignoring deletion semantics: Stale embeddings can expose withdrawn or superseded content.
  • Caching without permission context: Cache keys must reflect tenant, access scope, and relevant data versions.
  • Equating citations with truth: A valid source link does not prove that the source supports the claim.
  • Skipping abstention tests: A reliable system must know when its evidence is insufficient.

For an internal knowledge assistant, a sensible baseline is FastAPI, an appropriate document parser, S3-compatible storage, PostgreSQL with pgvector, a hosted embedding API, a hosted generation model, and a background worker. Add tracing from the beginning.

Move to specialized search infrastructure or more elaborate orchestration only when evaluation and load testing identify a concrete need. For adjacent architecture decisions, browse more Tech stack topics.

Frequently asked questions

Do I need a vector database for RAG?

No. RAG can retrieve through keyword search, SQL, APIs, or a combination of methods. Vector search is useful for semantic matching, but structured records and exact identifiers may be better served by conventional queries.

Is RAG better than fine-tuning?

They solve different problems. RAG supplies external evidence that can be updated independently of the model. Fine-tuning can improve behavior, style, or task-specific performance. It is generally not a substitute for permission-aware access to changing documents.

Should I use LangChain or LlamaIndex?

Use the framework that simplifies your actual workflow. LlamaIndex is often convenient for ingestion and retrieval; LangChain offers broad integrations. Neither is mandatory. Compare debugging clarity, dependency stability, and team familiarity before committing.

What should I optimize first in a RAG application?

Start with evidence quality: extraction, permissions, chunking, and retrieval recall. Then improve ranking and generation. Once the system answers reliably, optimize latency and cost without weakening the evaluation results that establish its usefulness.

Have a question about this topic?

Ask the community and get answers from practitioners.

Start a discussion