GUIDE TECH STACK

Tech stack for building AI agents

Building reliable AI agents requires more than a model and a framework. This guide explains how to choose the runtime, tools, data, safeguards, and infrastructure that turn agent prototypes into dependable software.

Start with the work, not the framework

Choosing a tech stack for building ai agents starts with deciding what the system can do, what it may access, and how failures will be contained. A support assistant that drafts replies needs a different architecture from an operations agent that changes infrastructure or a purchasing agent that commits money.

For MyDiscussions readers evaluating real products, the central question is not which framework offers the most features. It is which combination of components makes the required behavior reliable, observable, and economically sustainable.

An agent typically combines a model with tools, state, and a control loop. Unlike a simple chatbot, it can choose intermediate actions and use their results to continue a task. That flexibility introduces architectural obligations: authorization, execution limits, recovery, and evidence that the system performs acceptably.

Define the operating constraints first

Write a short workload specification before selecting vendors. Include representative tasks, allowed actions, failure consequences, and measurable acceptance criteria.

Decide how much autonomy is necessary

Use the least autonomous design that solves the problem:

  • Fixed workflow: Application code determines every step; a model handles bounded tasks such as extraction or classification.
  • Bounded agent: The model chooses tools within a predefined workflow, permission scope, and execution budget.
  • Open-ended agent: The model plans and revises a longer sequence of actions, usually requiring stronger isolation and supervision.

A document-processing pipeline rarely needs unrestricted planning. An investigation assistant may benefit from iterative search, but still should not inherit unrestricted database access.

Turn product requirements into selection criteria

CriterionQuestion to answerArchitectural consequence
Action riskCan the agent spend money, delete data, or contact customers?Approval gates, scoped credentials, audit logs
Task durationDoes work finish immediately or wait for external events?Request handler versus durable workflow
Data sensitivityCan prompts contain regulated or confidential information?Provider review, regional controls, redaction
FreshnessMust answers reflect changing internal information?Live tools or retrieval with refresh policies
Load patternAre tasks interactive, scheduled, or bursty?Streaming, queues, concurrency limits
Failure toleranceCan a task be retried without harmful duplication?Idempotency keys and reconciliation
Quality targetWhat constitutes a correct completed task?Task-specific evaluation and release gates

These answers should determine the stack—not the popularity of a library.

Choose models through task-level evaluation

Model choice affects tool selection, argument accuracy, latency, context handling, and operating cost. Evaluate complete tasks rather than selecting solely from general benchmark rankings.

Hosted APIs versus self-hosted models

Hosted services from OpenAI, Anthropic, and Google reduce infrastructure work and provide access to capable models. Enterprise procurement may instead favor services such as Azure OpenAI or Amazon Bedrock because of existing identity, billing, and governance arrangements.

Compare:

  • Tool-calling reliability and adherence to output schemas.
  • Performance on your documents, terminology, and ambiguous requests.
  • Regional availability, retention terms, and contractual controls.
  • Rate limits, timeout behavior, and support arrangements.
  • Input, output, caching, and related tool charges.

Self-hosting open-weight models with vLLM or Hugging Face Text Generation Inference gives greater deployment control. It also requires capacity planning, GPU operations, model-serving security, and upgrade evaluation.

Self-hosting is not automatically cheaper or more private. Economics depend on utilization and operational overhead; privacy depends on the entire deployment, including logs and telemetry.

Use routing only when it earns its complexity

A smaller model can handle intent classification or routine extraction, while a stronger model handles difficult planning. Start with one model, however, unless evaluation shows a clear reason to split workloads.

Routing introduces another failure point: a request sent to an insufficient model may produce a plausible but incorrect result. Measure the router and downstream task together.

Read the official function-calling documentation when designing model-tool interactions. Structured schemas help with argument shape, but valid JSON does not establish authorization or correctness.

Separate agent orchestration from durable execution

Orchestration decides what happens next. Durable execution ensures work survives interruptions. Some products overlap, but these are distinct responsibilities.

Select the smallest useful orchestration layer

A straightforward service using an SDK and explicit application code often works best for agents with only a few tools. Teams retain direct control over prompts, retries, state, and tracing.

Consider frameworks when their abstractions solve concrete problems:

  • LangGraph: Useful for explicit state graphs, branching, checkpoints, and human intervention.
  • OpenAI Agents SDK: Relevant for teams wanting agent, tool, handoff, and tracing primitives within its ecosystem.
  • PydanticAI: Attractive to Python teams emphasizing typed dependencies and validated outputs.
  • LlamaIndex: Useful when retrieval and document-centered workflows dominate the application.

Evaluate framework escape hatches. Can you inspect state, override retries, export traces, and test individual steps without running an entire agent?

Avoid adopting multiple overlapping agent frameworks. Their different concepts of memory, execution, and tool registration can make debugging harder than plain code.

Add durable workflows for long-running work

Use Temporal, or an equivalent durable workflow system, when tasks must survive restarts, wait for approvals, or coordinate external operations over time.

A useful separation is:

  • The agent proposes the next action.
  • Application policy validates it.
  • The workflow system schedules and records execution.
  • Tool results return to the agent as observations.

Keep nondeterministic model calls inside activities or equivalent units appropriate to the workflow engine. Do not assume an agent checkpoint provides all the guarantees of durable business-process execution.

Design tools as secure application interfaces

Tools are where generated intentions become real-world effects. Their design often matters more than prompt sophistication.

Prefer narrow, typed tools

Expose functions such as lookup_order, draft_refund, and submit_refund_for_approval, rather than unrestricted SQL or a general-purpose shell.

Each tool should define:

  • A validated input schema and bounded output.
  • Server-side authentication and authorization.
  • Timeouts and explicit error categories.
  • Idempotency behavior for state-changing actions.
  • Audit metadata identifying the user, tenant, and task.

Keep business rules outside the model. A refund limit belongs in application code, even if the prompt also describes it.

Separate read, prepare, and commit operations. An agent may gather evidence and prepare an action automatically while a human approves the consequential step.

Treat MCP as connectivity, not a security boundary

The Model Context Protocol, or MCP, can standardize how compatible clients connect to tools and resources. It is useful when integrations need reuse across agent applications.

However, using MCP does not make a server trustworthy or its tools safe. Review server provenance, credential handling, accessible resources, and authorization behavior. The official MCP documentation explains the protocol’s roles and architecture.

Prefer direct API integrations when they are simpler and already satisfy your requirements. Standardization should reduce integration work, not add an unnecessary network layer.

Build state, memory, and retrieval deliberately

“Agent memory” can refer to several different systems. Mixing them creates correctness and privacy problems.

Keep four kinds of information separate

  • Run state: Current step, intermediate outputs, outstanding actions, and checkpoints.
  • Conversation history: Messages needed to interpret the user’s request.
  • Knowledge retrieval: Documents or records selected as evidence.
  • Long-term preferences: Explicitly retained facts about a user or organization.

PostgreSQL is a strong default for durable application state and metadata. Redis can support caching, rate limiting, and ephemeral coordination, but should not casually become the only source of truth for important work.

For retrieval, pgvector keeps vector search close to relational data. Dedicated systems such as Qdrant, Pinecone, or Weaviate may fit workloads requiring specialized retrieval features or independent scaling. Elasticsearch and OpenSearch are relevant when keyword search and filtering are central.

Choose retrieval by failure mode

Semantic search helps with paraphrases. Keyword search matters for exact identifiers, product codes, and error messages. Hybrid retrieval can address both; reranking may improve relevance at additional latency and cost.

Enforce tenant and document permissions during retrieval, not after confidential content has entered the model context. Store source identifiers and versions so answers can cite evidence and investigations can reconstruct what the agent saw.

Treat retrieved documents and tool responses as untrusted data. They may contain instructions attempting to redirect the agent or expose secrets.

Make evaluation, security, and observability core layers

A successful demo proves that one path can work. Production readiness requires evidence across ordinary, ambiguous, and adversarial cases.

Evaluate outcomes and trajectories

Build a versioned test set from representative tasks. Include missing information, unavailable tools, contradictory documents, and malicious instructions embedded in retrieved content.

Measure:

  • Task completion against explicit acceptance criteria.
  • Correct tool selection and argument values.
  • Unauthorized action attempts and policy violations.
  • Evidence quality and unsupported claims.
  • Human escalation frequency.
  • End-to-end latency and cost per completed task.

Use deterministic assertions where possible. Model-based judges can assess softer qualities, but calibrate them against human review.

Tools such as Braintrust, LangSmith, and Arize Phoenix can help organize traces and evaluations. Select based on your deployment constraints, retention needs, and ability to export data.

Trace behavior without creating a data leak

Record model versions, prompt versions, tool calls, errors, timings, and approval decisions. Use correlation IDs across the API, agent runtime, queue, and workers.

OpenTelemetry can provide a vendor-neutral foundation for service traces. Agent-specific instrumentation may still be needed for prompt and tool details.

Do not log every prompt and response indiscriminately. Redact secrets, restrict trace access, define retention, and distinguish production telemetry from evaluation datasets.

Deploy for bounded execution and controlled cost

A common starting architecture uses FastAPI or Node.js/TypeScript for the API, PostgreSQL for state, object storage for files, and workers for asynchronous execution.

Managed containers are often sufficient. Kubernetes becomes useful when existing platform capabilities or scheduling requirements justify its operational burden. Serverless functions suit short steps, but long agent loops may conflict with execution limits.

Set explicit budgets for:

  • Model calls and tool calls per task.
  • Elapsed time and concurrent work.
  • Retrieved context and generated output.
  • Sandbox CPU, memory, filesystem, and network access.

Generated code should run in an isolated environment, never inside the main application process with production credentials.

Track cost per successful task, including retries, retrieval, infrastructure, and human review. A cheaper model can increase total cost if it needs more attempts or escalations.

A step-by-step process for selecting the stack

  1. Choose one bounded use case. Define inputs, outputs, prohibited actions, and the conditions requiring human approval.
  2. Build a non-agent baseline. Try ordinary code, search, or a fixed model workflow. Establish whether autonomous tool selection adds value.
  3. Create the evaluation set. Include realistic edge cases before optimizing prompts or comparing frameworks.
  4. Benchmark candidate models. Run the same tasks with equivalent tools and record quality, latency, and total task cost.
  5. Implement narrow tools. Add authorization, schemas, idempotency, and audit events before granting write access.
  6. Add state and retrieval selectively. Introduce vector search, long-term memory, or durable workflows only for demonstrated needs.
  7. Test failure recovery. Interrupt workers, duplicate requests, expire credentials, and simulate provider failures.
  8. Roll out gradually. Start with read-only or draft mode, then expand permissions using measured results and rollback criteria.

For a support-resolution agent, this might produce a Python service, a hosted model, explicit orchestration or LangGraph, PostgreSQL with pgvector, and narrow help-desk tools. Temporal becomes useful if cases wait on approvals or external events; it is not automatically necessary for the first version.

Common mistakes that make agent stacks fragile

  • Starting with multi-agent coordination: Add separate agents only when role separation measurably improves results. Every handoff creates context and debugging overhead.
  • Treating conversation history as a database: Store authoritative business facts in application systems, not only in prompts.
  • Retrying writes blindly: After a timeout, reconcile external state before repeating an action that may already have succeeded.
  • Assuming provider portability is free: APIs may look similar while tool behavior, schemas, and refusal patterns differ.
  • Using prompts as permission controls: Enforce access and spending limits outside the model.
  • Shipping without versioned evaluations: Model, prompt, retrieval, and tool changes can all introduce regressions.

For adjacent architecture decisions, browse more Tech stack topics.

Frequently asked questions

What is the minimum tech stack for building AI agents?

Start with an application runtime, a model SDK, a few typed tools, durable task records, and basic tracing and evaluation. Add retrieval only when the agent needs external knowledge. A dedicated vector database and multi-agent framework are not mandatory.

Is Python or TypeScript better for agent development?

Python fits teams using data science, evaluation, and machine-learning libraries. TypeScript fits teams integrating agents into existing web products and Node.js services. Prefer the language your team can secure and operate; both support capable agent applications.

When should we use a vector database?

Use vector retrieval when semantic similarity helps locate relevant information across a substantial corpus. First test whether SQL, keyword search, or direct API queries already solve the problem. PostgreSQL with pgvector can avoid introducing another datastore prematurely.

Should we build one agent or several specialized agents?

Begin with one bounded agent or workflow. Introduce specialized agents when distinct permissions, expertise, or context requirements justify separation. Compare the improvement against added latency, coordination failures, and evaluation work before committing to a multi-agent architecture.

Have a question about this topic?

Ask the community and get answers from practitioners.

Start a discussion