GUIDE MIGRATION

Migrating from OpenAI API to open-source LLMs

Moving away from OpenAI requires more than replacing an API endpoint. This guide explains how to evaluate open models, choose a deployment architecture, preserve application behavior, and roll out safely.

What changes when you leave the OpenAI API?

For teams migrating from openai api to open-source llms, the central challenge is replacing a managed behavior contract—not simply choosing a model with downloadable weights. Your application depends on instruction following, tool calls, response schemas, latency, safety controls, and operational reliability. A successful migration preserves the behaviors that matter while changing who controls—and maintains—the underlying system.

The destination is not necessarily a GPU cluster you operate yourself. You can run open models through a managed inference provider, deploy them inside your cloud account, or host them on infrastructure you own. Each option shifts cost, control, and operational responsibility differently.

Treat this as a workload migration, not a model popularity contest. Start with a bounded application where success is measurable, then expand only after quality and economics hold up under production conditions.

Clarify what “open-source LLM” means for your project

Many models commonly described as open source are more accurately open-weight models. Downloadable weights do not automatically imply unrestricted commercial use, complete training transparency, or access to training data.

Model families such as Llama, Qwen, Mistral, and Gemma contain different releases, sizes, capabilities, and licensing terms. Evaluate the exact checkpoint rather than assuming every model in a family carries the same permissions.

Before approving a candidate, check:

  • Commercial permissions: Are your product, deployment scale, and use case permitted?
  • Redistribution obligations: What applies if you ship weights, adapters, or a quantized derivative?
  • Acceptable-use restrictions: Are there conditions relevant to your sector?
  • Artifact availability: Can you obtain the tokenizer, configuration, and required runtime components?
  • Supply-chain integrity: Can you pin a revision and verify the source of model artifacts?

If an organizational policy requires an actual open-source license, involve legal reviewers before benchmarking. Otherwise, you may optimize around a model you cannot deploy.

Decide which workloads should move

A mandate to replace every OpenAI call usually creates unnecessary risk. Inventory workloads individually and identify the reason for moving each one.

Good early candidates often include document classification, bounded extraction, internal summarization, and retrieval-grounded question answering. These have testable outputs and relatively clear failure conditions. Long-running agents, complex multimodal workflows, and high-stakes advice require broader evaluation.

Define concrete acceptance criteria

Translate business expectations into measurable gates:

DimensionWhat to measureExample acceptance rule
Task qualityCorrect classifications, supported answers, extraction accuracyMeets the approved baseline on representative cases
Structured outputSchema validity and field correctnessNo invalid payload reaches a downstream action
Tool useCorrect tool, arguments, and sequenceEvery privileged action passes independent authorization
LatencyTime to first token and end-to-end latencyMeets the application SLO at expected concurrency
ReliabilityTimeouts, overloads, retries, unavailable replicasRemains within the service error budget
EconomicsCost per successfully completed taskImproves cost without unacceptable quality loss
GovernanceData handling, licenses, artifact provenanceApproved by security and legal owners

Set thresholds before comparing candidates. Otherwise, teams tend to reinterpret disappointing results as acceptable after investing in infrastructure.

Use cost per accepted outcome, not token price alone. A cheaper model can cost more when it requires longer prompts, repeated attempts, or human correction.

Choose a deployment model

Managed open-model inference

Providers such as Hugging Face Inference Endpoints, Together AI, Fireworks AI, and cloud model platforms can reduce the work involved in serving open models.

This approach suits teams that want model choice without immediately owning GPU scheduling, runtime upgrades, and capacity management. However, verify exact checkpoint availability, deployment regions, retention policies, rate limits, and version-pinning guarantees.

Managed hosting does not automatically provide portability. Provider-specific tool parsers, adapters, and output constraints can become new dependencies.

Self-hosted inference

Self-hosting provides stronger control over model versions, network boundaries, and scheduling. Common infrastructure choices include NVIDIA GPU instances on AWS, Azure, and Google Cloud, plus specialist GPU providers.

For serving, vLLM and SGLang are relevant production candidates. llama.cpp is useful for local inference and supported quantized deployments. Kubernetes can support scheduling and recovery, but it does not eliminate GPU capacity planning.

Self-hosting is most attractive when workloads justify sustained capacity, privacy requirements demand tighter control, or customization needs exceed managed offerings. For small or highly bursty applications, idle capacity can undermine the economics.

Hybrid routing

A hybrid architecture routes selected tasks to open models while retaining OpenAI for workloads that have not passed migration gates.

This supports gradual adoption, but routing rules need explicit boundaries. Never silently send restricted data to an external fallback. If data residency motivated the migration, some failures must return a controlled error instead.

A step-by-step migration process

Step 1: Inventory your actual OpenAI dependencies

Search application code, gateways, background jobs, and notebooks for API usage. Record:

  • Endpoint, model, parameters, and prompt templates.
  • Input and output token distributions.
  • Tool definitions and structured-output schemas.
  • Streaming behavior, retries, timeouts, and cancellation.
  • Embeddings, retrieval stores, file processing, and multimodal inputs.
  • Provider-managed conversation state or orchestration.

Distinguish model inference from platform features. An OpenAI-compatible chat endpoint does not necessarily replace the Responses API, hosted tools, stored state, or file workflows.

If those features are involved, budget for application-side replacements. Frameworks such as LlamaIndex or LangChain may help, but evaluate whether a smaller explicit orchestration layer would be easier to maintain.

Step 2: Build a representative evaluation set

Create a versioned dataset from permitted, appropriately de-identified production examples. Include routine traffic, edge cases, multilingual requests, long inputs, ambiguous instructions, and adversarial content.

For each case, define what success means: a reference answer, required fields, allowed tool actions, supporting evidence, or an explicit abstention.

Useful tools include promptfoo for application-level comparisons, Inspect AI for structured evaluations, and lm-evaluation-harness for standardized model testing. General benchmarks help shortlist candidates; they do not substitute for testing your application.

Use human review on a stratified sample. Model-based judging can accelerate analysis, but calibrate it against expert judgments and avoid making the candidate model its own sole evaluator.

Step 3: Shortlist models and test the intended configuration

Compare a small set of plausible models rather than dozens of unrelated checkpoints. Consider task quality, language support, context needs, license, available hardware, and serving-engine compatibility.

Test the configuration you expect to deploy, including quantization. A full-precision evaluation does not establish that a lower-bit deployment preserves extraction accuracy or tool-call reliability.

Long-context claims also require application tests. Measure whether the model finds relevant evidence at different positions and whether latency remains acceptable.

Keep the checkpoint revision, tokenizer, chat template, runtime version, and decoding parameters pinned. Changing any of these can change behavior.

Step 4: Select the serving stack and establish API parity

For an existing OpenAI client integration, an OpenAI-compatible server can minimize transport changes. The vLLM OpenAI-compatible server documentation describes supported interfaces and configuration.

However, API compatibility is not behavioral compatibility. Test:

  • Roles, chat templates, and system-message handling.
  • Streaming chunks and completion termination.
  • Tool-call formatting and argument parsing.
  • JSON-schema support and unsupported schema features.
  • Stop sequences, token limits, and usage accounting.
  • Error responses, retries, and cancellation.

A gateway such as LiteLLM can centralize routing and request normalization. Keep provider-specific exceptions visible: an abstraction should not conceal unsupported capabilities.

Step 5: Adapt prompts, tools, and structured outputs

Begin with existing prompts to establish a baseline, then change one variable at a time.

Open models may need clearer task boundaries, shorter instructions, or a few carefully selected examples. Conversely, elaborate prompts can hurt smaller models by consuming attention and context.

For structured output, use constrained decoding when the model and serving stack support the required schema. Still validate outputs using tools such as Pydantic or JSON Schema. Syntactically valid JSON can contain incorrect business data.

For tool use, place authorization outside the model:

  1. Parse and validate the proposed call.
  2. Check user permissions and resource scope.
  3. Require confirmation for consequential actions.
  4. Execute only allowlisted operations.
  5. Return sanitized results to the model.

A generated tool call is a proposal, not permission.

Step 6: Migrate embeddings separately

Replacing the generation model does not require replacing the embedding model.

If you do change embeddings, re-embed the corpus and build a parallel index. Different models produce different vector spaces, even when their output dimensions match. Do not mix old document embeddings with new query embeddings.

Evaluate retrieval independently using relevance judgments and recall-oriented measures. Check normalization, distance metrics, chunking, and whether the model expects query or document prefixes.

Systems such as pgvector, Qdrant, Milvus, and Weaviate can support retrieval architectures, but storage choice cannot compensate for poorly evaluated embeddings.

Step 7: Load-test and model total cost

Measure under realistic request lengths, output lengths, and concurrency. Include burst traffic and cancellation—not just sequential benchmark requests.

Track time to first token, inter-token latency, queue time, tail latency, throughput, and memory use. Keep prefill-heavy document tasks distinct from generation-heavy workloads.

Model total cost as:

Compute + redundancy + storage + networking + operations + evaluation + failure recovery.

Compare this with your current bills and the official OpenAI API pricing, including any caching or batch-processing discounts your workload actually uses.

Weight memory is only part of GPU capacity. KV cache, runtime overhead, batching, and context length also matter. Lower-bit weights can reduce memory needs, but the resulting quality and speed depend on hardware, kernels, and workload.

Step 8: Roll out with reversible controls

Start with offline replay, then shadow traffic where data policies permit. Shadow responses should not reach users or execute tools.

Next, canary a controlled traffic segment. For stateful interactions, use sticky routing so a conversation does not unexpectedly switch behavior between turns.

Version the complete release: model, prompt, tokenizer, runtime, tool parser, retrieval configuration, and routing rules. Set automatic rollback conditions for quality signals, schema failures, latency, and availability.

Retain the previous path until the replacement survives representative traffic peaks and operational incidents.

Security and operations you now own

Running a model in your environment changes the trust boundary; it does not remove risk.

Review model artifacts and dependencies, restrict outbound network access, and avoid enabling arbitrary repository code without inspection. Authenticate inference endpoints, apply tenant-aware limits, and control access to logs and prompt traces.

Production controls should include:

  • Input limits: Bound context length, attachments, and output budgets.
  • Data protection: Define retention, redaction, encryption, and access policies.
  • Prompt-injection defenses: Treat retrieved documents and tool responses as untrusted.
  • Observability: Record operational metadata without unnecessarily retaining sensitive content.
  • Incident response: Support artifact rollback, endpoint isolation, and credential rotation.

The NIST AI Risk Management Framework offers a useful governance structure. Translate it into named owners and release gates rather than treating it as a checklist detached from engineering.

Common mistakes that derail migrations

  • Choosing by leaderboard rank: Public benchmarks rarely represent your prompts, documents, or tool workflows.
  • Assuming self-hosting is automatically cheaper: Low utilization, redundancy, and engineering time can outweigh token savings.
  • Migrating everything together: Simultaneous changes to generation, embeddings, retrieval, and orchestration make regressions difficult to diagnose.
  • Fine-tuning before diagnosing failures: Missing knowledge may require retrieval; invalid output may require constrained decoding.
  • Ignoring model lifecycle management: Unpinned revisions and runtime upgrades can introduce silent behavior changes.
  • Using uncontrolled fallback: Automatic external routing can violate privacy requirements and hide capacity problems.

The practical remedy is consistent: isolate changes, evaluate them against explicit criteria, and preserve reversibility.

Frequently asked questions

Can I keep using the OpenAI SDK?

Often, yes, if the serving platform exposes supported OpenAI-compatible endpoints. You may change the base URL, credentials, and model identifier. Still test endpoint coverage, streaming, tool calls, and errors; compatibility is rarely universal.

Do I need to fine-tune an open model?

Not necessarily. First test a suitable instruction-tuned model with task-specific prompts, retrieval, and output constraints. Fine-tuning becomes useful when repeated failures reflect a learnable behavior or domain pattern and you have sufficient high-quality examples.

Will an open model be cheaper than the OpenAI API?

It depends on utilization, model size, hardware, reliability requirements, and operational overhead. Compare equivalent quality and service levels. Include idle capacity and failed attempts rather than comparing GPU rental with nominal token prices alone.

Should we replace OpenAI completely?

Only if every relevant workload meets its acceptance gates. Partial migration can capture control or cost benefits while preserving stronger performance elsewhere. Maintain explicit routing policies, especially where privacy requirements prohibit external fallback.

Make the first migration deliberately narrow

Choose one measurable workload, establish its baseline, and test a pinned model configuration behind a reversible routing layer. Expand only when quality, security, reliability, and total cost are acceptable together.

That approach turns model choice into an engineering decision rather than a platform bet. For related planning guidance, browse more Migration topics.

Have a question about this topic?

Ask the community and get answers from practitioners.

Start a discussion