GUIDE INTEGRATION

AI integration into legacy systems

Add AI capabilities to established enterprise applications without replacing their core transaction engines. This guide covers architecture, data access, security, evaluation, and production rollout.

Integrate AI without destabilizing the systems that matter

Successful ai integration into legacy systems starts with a constraint: the existing application must keep doing its job. An insurer’s claims platform, a manufacturer’s ERP, or a hospital’s patient administration system cannot become unreliable because an AI service produces an unexpected response.

The practical goal is rarely to “make the mainframe intelligent.” It is to add a bounded capability—document extraction, knowledge retrieval, classification, or assisted decision-making—around established business processes while preserving transaction integrity, permissions, and operational ownership.

That makes this an integration problem before it becomes a model-selection problem. For decision-makers, the critical questions concern risk, measurable value, and maintenance cost. For practitioners, they concern interfaces, data contracts, failure isolation, and testing.

Decide whether AI belongs in the workflow

Not every legacy limitation requires AI. If a task follows stable rules and has structured inputs, SQL, a rules engine, or ordinary application code may be cheaper and more dependable.

AI becomes useful when the workflow contains ambiguity: inconsistent supplier descriptions, scanned forms, free-text case notes, or questions spanning a large documentation collection.

Apply concrete acceptance criteria

Before choosing a vendor, define a candidate use case against these criteria:

  • Business outcome: Identify the operational change, such as fewer manual document corrections or faster access to maintenance instructions.
  • Ground truth: Determine whether historical cases, approved answers, or expert reviewers can evaluate outputs.
  • Error impact: Separate reversible suggestions from actions affecting money, access, safety, or legal obligations.
  • Integration feasibility: Confirm that supported APIs, events, exports, or controlled batch interfaces exist.
  • Data eligibility: Establish whether the required data can leave its current environment.
  • Fallback: Specify how work continues if the model, network, or integration service fails.

A strong first use case is extracting proposed invoice fields into an accounts-payable review queue. A weak first use case is autonomously posting payments through undocumented database writes.

Bound the decision, not just the prompt. Define which decisions the model may support, which require approval, and which it must never make.

Choose an architecture that protects the transaction core

Legacy systems differ substantially. A COBOL application with nightly file exchanges needs a different approach from an older Java application with usable REST endpoints.

The safest common pattern is an AI sidecar: a separately deployed service that reads approved inputs, invokes a model, validates the result, and returns a proposal through a controlled interface.

PatternBest fitMain advantageMain trade-off
Synchronous API adapterInteractive assistance with supported APIsImmediate responsesAdds latency and runtime dependencies
Asynchronous event consumerClassification, enrichment, document processingIsolates failures and absorbs burstsRequires duplicate handling and status tracking
Batch enrichmentFile-based systems and scheduled workflowsFits existing operating proceduresResults are not immediate
Retrieval-augmented generationQuestions over manuals, policies, case knowledgeSupplies relevant source materialRetrieval and permissions require careful engineering
UI automation bridgeSystems without supported interfacesCan enable a limited pilotFragile when screens or workflows change

Prefer supported interfaces over direct database writes

Use business APIs whenever possible. SAP BAPIs or supported OData services, IBM z/OS Connect, and established enterprise service layers can expose functionality without bypassing application rules.

MuleSoft Anypoint Platform, Boomi, and Azure Logic Apps can simplify connector management and orchestration. Their value depends on connector coverage, licensing, throughput, and whether your team already operates them.

Read-only database access can be reasonable for an approved integration. Direct writes are different: they can bypass validation, audit logic, and transaction boundaries.

Where APIs are absent, controlled file exchange may be safer than robotic process automation. UiPath or Power Automate can bridge a UI, but session failures, layout changes, and partial completion must be treated as normal failure modes.

Keep slow AI work outside critical transactions

Do not hold a database transaction open while waiting for a model response. Commit the business event, process AI work separately, then submit the validated result through an approved command.

Apache Kafka, RabbitMQ, or managed queues support this separation. Use correlation identifiers to connect source records, model requests, approvals, and eventual writes.

Build a trustworthy data path

Most integration problems emerge before inference: inconsistent identifiers, undocumented fields, stale exports, duplicate records, and inaccessible attachments.

Create a data contract covering field definitions, null handling, timestamps, encoding, ownership, and permitted uses. Legacy dates and customer identifiers deserve particular attention; an apparently harmless conversion can change business meaning.

Retrieve context rather than copying everything

Retrieval-augmented generation, or RAG, can supply relevant documents at request time without training a model on the entire enterprise archive.

A typical pipeline:

  • Ingests approved documents and extracts text.
  • Splits content into meaningful sections.
  • Attaches provenance, version, access rules, and effective dates.
  • Builds keyword and vector indexes.
  • Retrieves authorized material for each request.
  • Returns answers with traceable source references.

PostgreSQL with pgvector can suit teams already operating PostgreSQL. Elasticsearch, OpenSearch, and Azure AI Search offer alternatives with different retrieval, filtering, and operational capabilities.

RAG is not an authorization layer. The retrieval service must enforce access before content reaches the model. Filtering only the final answer is too late.

Keep source identifiers and index freshness visible. For balances, eligibility, or other rapidly changing facts, prefer a live authorized lookup rather than a stale document index.

Treat healthcare interfaces as semantic contracts

Healthcare integration adds terminology, identity matching, and clinical context to the interface problem.

FHIR defines resources and exchange patterns, but a FHIR endpoint does not guarantee that two applications interpret every field identically. Specify supported versions, implementation guides, profiles, and terminology bindings using the official HL7 FHIR documentation.

For older HL7 v2 feeds, preserve message acknowledgments and patient-identity reconciliation. AI-generated summaries should not silently overwrite source clinical records; separate them from authoritative observations and route consequential changes through appropriate review.

Select models and platforms against operating constraints

Start with deployment constraints, then compare models using representative tasks. A general benchmark cannot tell you whether a model correctly extracts fields from your organization’s scanned purchase orders.

OptionAppropriate whenCosts and risks to examine
Hosted model APIFast deployment and limited infrastructure capacityData terms, network dependency, variable usage
Managed enterprise platformExisting cloud governance and identity controlsRegional availability, service limits, platform coupling
Self-hosted open-weight modelLocal processing or infrastructure control is essentialHardware, patching, scaling, evaluation, licensing
Specialized extraction or classification serviceTask scope is narrow and repeatableFormat coverage, customization limits, per-document cost

Azure OpenAI, Amazon Bedrock, and Google Cloud Vertex AI provide managed options. Their capabilities, model availability, and data-handling terms vary by service configuration and region; verify rather than assuming equivalence.

For self-hosting, vLLM is one serving option. Hugging Face Transformers provides tooling for working with compatible models. Neither removes the need for capacity planning or security operations.

Frameworks such as LangChain and LlamaIndex can accelerate retrieval and orchestration. For a single extraction call, however, a small explicit service may be easier to audit.

Calculate cost per completed business task, including retries, OCR, retrieval, human review, storage, and support—not just tokens. A cheaper model can cost more overall if it creates substantial correction work.

Implement the integration step by step

1. Map the existing workflow and ownership

Document the current sequence, including manual work, exception queues, scheduled jobs, and downstream consumers.

Identify who owns the source system, data, integration service, model evaluation, and business approval. Record peak workloads and change windows.

Deliverable: a workflow diagram with explicit read, proposal, approval, and write boundaries.

2. Establish a non-AI baseline

Measure the current process on a representative sample. For extraction, track field-level correctness and correction effort. For knowledge assistance, assess answer usefulness and source accuracy.

Include difficult cases, not just clean examples. Compare AI against simpler alternatives such as templates, search, and rules.

Deliverable: a baseline and acceptance criteria approved by the process owner.

3. Build a narrow adapter

Expose only the operations the use case requires. Translate legacy representations into versioned schemas and isolate model-specific code behind a replaceable interface.

For extraction, request structured output and validate it against JSON Schema. Parsing successfully is necessary, but not sufficient: apply domain checks for dates, identifiers, totals, and allowed values.

Deliverable: a contract-tested adapter that rejects malformed or unauthorized requests.

4. Implement identity and access controls

Use established identity infrastructure, such as Microsoft Entra ID or Keycloak, with short-lived credentials where supported.

Distinguish the user’s permissions from the service account’s permissions. A service account that can access every claim must not expose all claims to every employee.

Store secrets in a managed vault. Redact sensitive content from logs and define retention separately for prompts, responses, and audit events.

Deliverable: tested authorization boundaries and documented data handling.

5. Evaluate the entire pipeline

Test the combined system, not just model responses. Retrieval failures, outdated source records, and broken mappings can defeat an otherwise capable model.

Include:

  • Ambiguous documents and missing information.
  • OCR errors and unsupported languages.
  • Attempts to retrieve unauthorized records.
  • Instructions embedded in retrieved documents.
  • Provider timeouts and throttling.
  • Duplicate or out-of-order events.
  • Valid JSON containing incorrect business values.

Follow a lifecycle approach to risk identification and monitoring, such as the NIST AI Risk Management Framework.

Deliverable: a versioned test set with explicit release gates.

6. Run in shadow mode

Process real inputs without changing production records. Compare outputs with the existing workflow and collect reviewer feedback.

Shadow mode can expose actual document variability, burst patterns, and integration failures. Apply the same privacy controls as production; “read-only” does not mean risk-free.

Deliverable: evidence that the system meets acceptance criteria under realistic conditions.

7. Release gradually with a kill switch

Start with a limited workflow or user group. Make suggestions visibly distinct from committed records and display supporting evidence.

Increase automation only for task categories that pass agreed evaluation thresholds. Preserve a feature flag that disables AI while leaving the original process operational.

Deliverable: a staged release plan, rollback procedure, and named operational owner.

Engineer for unreliable outputs and unreliable networks

An AI integration has at least two failure classes: technical failure and plausible-but-wrong output. Each needs different controls.

Make execution deterministic even when generation is not

Set timeouts, bound retries, and use exponential backoff for retryable failures. Add circuit breakers to prevent an unavailable model provider from exhausting application resources.

Use idempotency keys for downstream writes. Queue redelivery must not create a second invoice, case, or appointment.

For partial failures, maintain explicit states such as pending_review, approved, applied, and failed. A durable workflow engine such as Temporal can help coordinate long-running operations.

Never treat generated text as an executable command. Map proposed actions to allowlisted operations, validate parameters, and recheck authorization immediately before execution.

Treat retrieved content as untrusted input

A support ticket or PDF may contain instructions designed to manipulate the model. System prompts alone are not a sufficient defense.

Limit tool permissions, separate retrieved content from trusted instructions, and require approval for consequential actions. The OWASP guidance for LLM applications provides useful threat categories for security review.

Avoid common implementation mistakes

  • Starting with unrestricted agents: Broad tool access makes failures harder to constrain. Begin with a fixed workflow and narrowly scoped tools.
  • Fine-tuning to compensate for missing facts: Frequently changing enterprise information generally belongs in retrieval or live lookups.
  • Trusting model-reported confidence: Validate any confidence signal against observed outcomes; prefer measurable checks and evidence requirements.
  • Ignoring deletion and permission changes: Indexes and caches need update, revocation, and deletion mechanisms.
  • Logging everything for debugging: Prompts and responses can contain sensitive records. Minimize collection and restrict access.
  • Skipping model-change testing: Pin versions where available and rerun regression tests before deployment changes.
  • Measuring output volume instead of value: Track accepted results, correction effort, exception rates, and completed-task cost.

Operational dashboards should distinguish model latency, retrieval latency, adapter failures, and downstream write failures. A single “AI success rate” hides the component that needs attention.

Frequently asked questions

Can AI be added without replacing the legacy application?

Yes. An external service can consume approved APIs, events, or exports and return suggestions through supported interfaces. Replacement becomes necessary only when the required access, controls, or performance cannot be achieved safely around the existing system.

Should we use RAG or fine-tuning?

Use RAG when answers require current, attributable enterprise knowledge. Consider fine-tuning when the challenge is consistent behavior on a stable task and suitable training examples exist. Neither approach substitutes for authorization, data quality, or end-to-end evaluation.

How should we decide whether to automate a write operation?

Consider reversibility, error impact, deterministic validation, and observed performance on representative cases. Begin with approval-gated writes. Expand automation only when failures can be detected, contained, and corrected without unacceptable business consequences.

What should the first production deployment include?

Choose one bounded use case, a supported integration interface, a representative evaluation set, access controls, monitoring, and a working fallback. Assign both technical and business owners. A small dependable capability provides a stronger foundation than a broad demonstration.

Build the integration as a maintainable capability

The strongest approach preserves the legacy system’s authority while placing AI behind explicit contracts. Keep transactions deterministic, retrieve only authorized context, evaluate business outcomes, and make disabling the AI component straightforward.

For related implementation patterns across enterprise APIs and data exchange, browse more Integration topics.

Have a question about this topic?

Ask the community and get answers from practitioners.

Start a discussion