How to integrate an LLM into an existing app
Adding an LLM to production software requires more than an API call. This guide explains how to choose an architecture, connect application data safely, evaluate quality, and roll out a reliable feature.
Start with an application workflow, not a model
Understanding how to integrate an llm into an existing app starts with identifying a workflow where language understanding or generation improves a measurable outcome. For an established product, the challenge is not getting a model to respond. It is fitting probabilistic behavior into existing authentication, data access, interfaces, and operational guarantees.
A useful first feature might summarize support tickets, extract fields from uploaded documents, or draft answers grounded in a knowledge base. “Add a chatbot” is less actionable because it leaves the task, permitted data, and success criteria undefined.
For decision-makers, the integration decision involves delivery effort, risk, and operating cost. For practitioners, it involves API contracts, retrieval, failure handling, and evaluation. Both groups should agree on one principle: the application remains the system of record and authority; the model provides bounded assistance.
Define the feature’s contract before selecting technology
Write a short feature specification that answers:
- Who uses it? Identify roles, tenant boundaries, and accessibility requirements.
- What does it receive? List user input, application records, documents, and conversation history.
- What can it return? Specify prose, structured fields, recommendations, or proposed actions.
- What must it never do? Include unauthorized disclosure, unsupported claims, and unapproved writes.
- How will success be measured? Define task completion, review effort, error tolerance, latency, and cost limits.
For example, a support application could generate a suggested response using the current ticket and approved documentation. The acceptance criteria might require valid source references, no unsupported refund promises, and agent approval before sending.
That is a testable integration contract. It also exposes whether an LLM is necessary: deterministic rules or conventional search may handle simpler requirements more reliably.
Choose an architecture that fits the existing app
Most applications should add a server-side AI integration layer, rather than call a model directly from browser or mobile code. This layer protects provider credentials and centralizes authorization, prompts, routing, retries, and usage accounting.
A typical request path is:
Client → application backend → authorization and context builder → model API → output validation → client
Retrieval systems and approved tools connect through the backend, not around it.
Compare the main deployment options
| Approach | Best fit | Main advantage | Important trade-off |
|---|---|---|---|
| Direct hosted API, such as OpenAI or Anthropic | Teams shipping an initial feature | Low infrastructure burden | Provider-specific behavior and contractual constraints |
| Managed cloud service, such as Amazon Bedrock, Azure OpenAI, or Vertex AI | Organizations with established cloud controls | Integration with cloud identity and procurement | Model availability, quotas, and features vary |
| Self-hosted open-weight model using vLLM or Hugging Face tooling | Specialized control or deployment requirements | Greater control over serving and data handling | GPU operations, capacity planning, and model maintenance |
Self-hosting is not automatically cheaper or safer. Compare total costs, including utilization, patching, scaling, and engineering ownership. Likewise, a managed service does not make an application compliant by itself.
For a first implementation, an official SDK and a small internal adapter are often enough. LangChain or LlamaIndex can help with retrieval and orchestration. LangGraph can support explicit, stateful workflows. Introduce frameworks when they reduce real complexity, not simply because the feature uses AI.
Keep your internal interface narrow enough to test independently, but avoid pretending all providers offer identical capabilities.
Step 1: Select a model using your own examples
Model leaderboards cannot tell you whether a model can correctly interpret your application’s records or follow your approval policy.
Build a representative evaluation set from approved, de-identified examples. Include routine cases, ambiguous requests, missing information, multilingual input where relevant, and adversarial instructions embedded in documents.
Assess candidate models against concrete criteria:
- Task quality: Does the answer satisfy the intended workflow?
- Grounding: Are factual claims supported by supplied evidence?
- Structured output: Does it follow the required schema consistently?
- Latency: Measure both first visible output and full completion.
- Context handling: Does performance hold up with realistic document volumes?
- Tool behavior: Does it select appropriate tools and supply valid arguments?
- Deployment suitability: Check region availability, retention controls, quotas, and contractual terms.
Test a stronger model and a less expensive alternative. A smaller model may handle extraction adequately while struggling with ambiguous synthesis. Routing tasks to different models can improve economics, but it adds evaluation and maintenance work.
Record exact model versions where available. A model change is a production dependency change, not merely a configuration edit.
Step 2: Implement a secure backend boundary
Create an application endpoint around a specific capability, such as drafting a ticket response. Do not expose an unrestricted proxy to the model provider.
The backend should:
- Authenticate the caller using existing application controls.
- Authorize access to the requested records.
- Load only the fields necessary for the task.
- Construct the model request from versioned instructions.
- Enforce timeouts, concurrency limits, and token budgets.
- Validate the response before returning or storing it.
Store provider keys in a secrets manager, such as AWS Secrets Manager, Azure Key Vault, or Google Cloud Secret Manager. Use separate credentials and budgets for development and production.
Treat user input and retrieved text as untrusted data, even when they come from internal systems. A document containing “ignore previous instructions” must not gain authority over application policy.
Retries also need boundaries. Retry transient failures with backoff and jitter, respecting provider guidance. Do not blindly retry invalid requests or completed tool actions: a repeated database write can be more damaging than a failed response.
Step 3: Connect application data with retrieval
A model does not automatically know your current product documentation, customer records, or internal policies. Provide the relevant context explicitly.
For a known ticket or account, ordinary authorized database queries may be sufficient. Do not introduce a vector database when the application already knows exactly which records it needs.
Use retrieval-augmented generation, or RAG, when the feature must locate relevant passages across a document collection.
Build a permission-aware retrieval pipeline
A practical RAG pipeline includes:
- Ingest approved documents and preserve source identifiers.
- Split documents into meaningful sections.
- Generate embeddings and store searchable metadata.
- Retrieve candidate passages using semantic, keyword, or hybrid search.
- Rerank candidates when relevance warrants the additional latency.
- Pass a limited evidence set to the model.
- Return source references alongside the answer.
PostgreSQL with pgvector may fit an application already running Postgres. Elasticsearch or OpenSearch can support hybrid search. Dedicated services such as Pinecone or Weaviate may simplify other workloads.
Filter by tenant and permissions before content reaches the model. Design for permission changes, document deletion, reindexing, and embedding-model migrations.
Measure retrieval separately from answer generation. If the correct policy never enters the context, a stronger generation model may only produce a more convincing wrong answer.
Step 4: Make outputs predictable and actions constrained
Free-form text is appropriate for drafts. Application state usually requires structured data.
Define a schema for responses that downstream code must consume. A ticket assistant might return:
draft_textsource_idsneeds_human_reviewreview_reasons
Use provider-supported structured outputs where available, then validate again in application code with tools such as Pydantic, Zod, or JSON Schema. The OpenAI structured outputs documentation explains supported approaches and limitations.
Schema compliance does not prove factual accuracy. Validate that referenced sources exist, identifiers belong to the current tenant, and field values satisfy business rules. Handle refusals, truncation, and invalid responses explicitly.
Separate suggestions from execution
If the model can invoke tools, expose narrow operations with typed arguments. Prefer a tool such as “retrieve current subscription” over arbitrary SQL execution.
For consequential actions:
- Recheck authorization at execution time.
- Enforce business rules outside the model.
- Require confirmation where appropriate.
- Use idempotency keys for retryable writes.
- Audit the requested action and its actual result.
A model-generated tool call is a proposal, not permission. Sending messages, modifying clinical records, or issuing refunds should follow established approval requirements.
Step 5: Address privacy and prompt injection explicitly
Before sending production data to a provider, review its applicable contract, retention settings, processing locations, subprocessors, and training-use terms. These may vary by product and configuration.
Minimize transmitted data. A request to summarize a technical incident usually does not need customer billing details.
For healthcare integrations, an LLM must not bypass established FHIR or HL7 access controls. Verify applicable legal requirements and contractual coverage, including a business associate agreement where required. Generated clinical text should enter an appropriately reviewed workflow, not silently become authoritative patient data.
Prompt injection can arrive through users, uploaded files, retrieved pages, or tool responses. Defensive instructions help, but prompts alone are not a security boundary.
Reduce the possible impact through restricted tools, least-privilege credentials, controlled network access, and independent authorization checks. The OWASP Top 10 for LLM applications provides a useful threat-modeling reference.
Treat model output as untrusted at rendering time too. Sanitize HTML, restrict unsafe Markdown behavior, and never execute generated code automatically.
Step 6: Test the complete workflow
Unit-test context construction, authorization filters, schema validation, and tool dispatch without making live model calls. Mock provider responses for failures, refusals, truncation, and malformed output.
Then run integration evaluations against real models.
Maintain separate measurements for:
- Answer quality: correctness, completeness, and usefulness.
- Retrieval quality: whether necessary evidence was found.
- Security: cross-tenant leakage and unauthorized action attempts.
- Reliability: timeouts, rate limits, and dependency failures.
- User experience: clear uncertainty, editable drafts, and recoverable errors.
Human review is particularly important for ambiguous or high-impact tasks. Automated model-based graders can help scale evaluation, but calibrate them against human judgments.
Version prompts, model configuration, retrieval settings, and evaluation datasets together. Run regression tests whenever any of them changes.
Step 7: Roll out with cost and reliability controls
Begin with internal users, then a limited opt-in cohort. Use feature flags and a kill switch. Shadow testing can assess outputs without exposing them to users, provided production data processing is authorized.
Track request latency, token usage, failures, retrieval time, validation failures, and tool execution outcomes. OpenTelemetry can connect AI requests to existing traces. Avoid collecting full prompts and responses indiscriminately; logs can contain sensitive data.
Estimate cost from the whole workflow:
Cost per completed task = model calls + retrieval and reranking + tool services + infrastructure + retries and review overhead
Check current rates on the applicable provider’s pricing page, such as Anthropic’s API pricing documentation. Account for output tokens, repeated context, multi-step workflows, and failed attempts.
Set per-user or per-tenant limits. Cap output length and orchestration steps. Cache only when freshness and access-control semantics permit it.
Design a useful fallback: preserve manual editing, offer conventional search, or queue nonurgent work. Streaming improves perceived responsiveness, but it does not remove the need to validate output before consequential use.
Common mistakes that undermine integration
- Choosing chat by default: An inline summary or structured extraction may fit the workflow better.
- Sending entire records: Extra context increases exposure and can dilute relevant evidence.
- Relying on prompt instructions for permissions: Authorization belongs in application code.
- Fine-tuning before diagnosing failures: Fine-tuning may improve behavior, but it does not automatically supply current facts or fix retrieval.
- Treating fluent answers as correct: Require evidence and task-specific evaluations.
- Adding autonomous loops too early: Start with explicit steps and bounded tool access.
- Ignoring dependency changes: Model updates, reindexing, and prompt edits can all cause regressions.
The strongest initial integration is usually narrow, observable, and reversible. Expand only after the feature demonstrates useful quality within acceptable risk and operating cost.
For related implementation guides, browse more Integration topics.
Frequently asked questions
Do I need to train an LLM for my existing app?
Usually not initially. Start with a hosted or deployed model, clear instructions, and authorized application context. Consider fine-tuning when evaluations show persistent behavioral problems that prompting and retrieval do not solve.
Should I use RAG or fine-tuning?
Use RAG for changing information that needs source attribution and access controls. Consider fine-tuning for learned task behavior or style. They can complement each other, but neither replaces authorization or factual evaluation.
Can an LLM safely update application data?
Yes, within a constrained workflow. Let the model propose typed actions, then enforce permissions, validation, confirmations, and audit logging in application code. Start with read-only capabilities before introducing writes.
How do I know the integration is ready for production?
Require task-specific quality thresholds, tested failure paths, privacy review, cost limits, and operational ownership. Demonstrate that unauthorized data stays inaccessible, consequential actions require proper approval, and users can recover when the model or provider fails.
Ask the community and get answers from practitioners.