LLM engineer interview questions
Evaluate LLM engineers on the decisions that determine production quality, not their recall of model terminology. This guide includes role-specific questions, strong-answer signals, a practical exercise, and a structured hiring scorecard.
What an LLM engineer interview should actually measure
The best llm engineer interview questions reveal whether a candidate can turn uncertain model behavior into a reliable product. Knowing transformer terminology helps, but hiring decisions should depend on how someone defines quality, selects architectures, diagnoses failures, manages costs, and protects data.
For MyDiscussions readers building technical teams, the central distinction is between making a convincing demo and owning a production system. A candidate may create an impressive chatbot without understanding retrieval permissions, evaluation leakage, or provider failures. Conversely, a strong engineer may choose a simple workflow over an elaborate agent because it better satisfies the requirements.
This guide focuses on applied LLM engineering: integrating models with software, data, tools, and operational controls. Research-heavy roles require additional assessment of model training, mathematics, and experimental design.
Define the role before choosing interview questions
“LLM engineer” can describe several different jobs. Establish the expected work before setting the interview bar.
| Role emphasis | Capabilities to prioritize | Useful interview evidence |
|---|---|---|
| LLM application engineer | APIs, structured outputs, workflow design, backend integration | Working implementation with validation and failure handling |
| RAG engineer | Ingestion, retrieval, ranking, permissions, grounding | Retrieval debugging and an evidence-based evaluation plan |
| Model adaptation engineer | Dataset design, supervised fine-tuning, LoRA, training evaluation | Controlled experiments and data leakage prevention |
| LLM platform engineer | Serving, observability, routing, reliability, cost | Capacity reasoning, deployment design, incident response |
Most positions combine these responsibilities, but not equally. Do not reject an excellent application engineer for lacking distributed training experience unless the job requires it.
Specify ownership boundaries: Will the hire control ingestion pipelines? Deploy open-weight models? Maintain customer-facing APIs? Handle regulated data? These answers should determine question selection and scoring weights.
Core LLM engineer interview questions and strong-answer signals
1. How would you decide whether a feature needs an LLM?
Present a concrete feature, such as classifying support tickets and drafting responses.
A strong candidate first asks about error tolerance, available labels, traffic, latency, and escalation requirements. They compare an LLM against rules, search, templates, and conventional classifiers.
Look for:
- A non-LLM baseline.
- Separate success criteria for classification and generation.
- Human review where mistakes create material harm.
- An experiment that could disprove the LLM approach.
Warning sign: Selecting a model before establishing the task and acceptance criteria.
2. How do tokenization and context limits affect system design?
Ask what happens when a multilingual conversation contains long documents and grows over time.
Strong answers distinguish tokens from characters and words. Token counts vary by model tokenizer, language, and content. Candidates should account for instructions, conversation history, retrieved passages, tool schemas, and output capacity.
Useful strategies include selective retrieval, conversation summaries, and explicit truncation policies. The trade-off matters: summaries may omit constraints, while a larger context window increases processing work and does not guarantee that every relevant detail will be used correctly.
A candidate should verify the selected model’s actual limits rather than quote a universal context size.
3. When would you use prompting, RAG, or fine-tuning?
Use a scenario involving company policies that change frequently and responses that must follow a consistent style.
Strong answers separate the objectives:
- Prompting: Establish instructions, examples, and output requirements.
- Retrieval-augmented generation: Supply current, authorized knowledge with traceable sources.
- Fine-tuning: Adapt recurring behavior or task performance when representative training data exists.
These approaches can be combined. Fine-tuning is not a dependable substitute for retrieving changing facts, and RAG does not automatically fix poor instruction-following.
Ask what evidence would justify fine-tuning. Good answers include persistent baseline failures, sufficient training examples, a held-out evaluation set, and an expected benefit that outweighs data preparation and maintenance costs.
4. Design a RAG system for an internal knowledge base
Require the candidate to cover the complete path:
- Parse documents while preserving useful structure.
- Attach document IDs, versions, timestamps, and access metadata.
- Choose chunk boundaries appropriate to the material.
- Index content for lexical, vector, or hybrid retrieval.
- Retrieve and optionally rerank candidates.
- Assemble context within a token budget.
- Generate an answer with source references.
- Evaluate retrieval and answer quality separately.
Tools such as PostgreSQL with pgvector, Elasticsearch, OpenSearch, and managed vector services can all be reasonable choices. LlamaIndex and LangChain can accelerate orchestration, but naming frameworks is not architectural reasoning.
Probe permissions directly. Unauthorized passages must not reach the model or leak through shared caches. The candidate should explain where access controls are enforced and tested.
5. How would you debug a plausible but incorrect RAG answer?
Strong engineers localize the failure instead of immediately changing the prompt.
Ask them to inspect:
- Whether the source contains the answer.
- Whether parsing or chunking damaged the evidence.
- Whether retrieval found the relevant passage.
- Whether ranking or context assembly excluded it.
- Whether the model misinterpreted or ignored it.
- Whether the cited source actually supports the claim.
Useful measures include recall at a chosen retrieval cutoff, ranking quality, answer correctness, and citation support. Candidates should understand that retrieval success and generation success are different outcomes.
A useful diagnostic experiment is to provide the correct evidence directly. If the model still fails, improving the vector index alone will not solve the problem.
6. How would you evaluate a model or prompt change?
Give the candidate a proposed upgrade that produces more fluent answers but may be less accurate.
Expect a versioned evaluation set containing representative requests, difficult cases, unanswerable questions, and relevant adversarial inputs. Strong answers segment results by task rather than reporting one aggregate score.
They should combine:
- Deterministic checks for schemas, required fields, and business rules.
- Reference-based or rubric-based quality assessments.
- Human review for ambiguous or high-impact cases.
- Operational measurements such as latency, tokens, and cost.
LLM judges can help scale evaluation, but their preferences need calibration against human judgments. Keep final test examples separate from prompt and training development. The OpenAI evaluation guide provides a concrete reference for evaluation workflows.
7. How would you make tool calling reliable?
Use an assistant that can inspect orders and request refunds.
A strong candidate treats a tool call as a model-generated proposal, not an authorization decision. Application code must validate arguments, authenticate the user, enforce permissions, and apply business rules.
Look for:
- Typed schemas and handling of missing or invalid arguments.
- Timeouts and bounded retries.
- Idempotency protection for side effects.
- Explicit confirmation for consequential actions.
- Audit records that connect requests, decisions, and results.
Probe ambiguous failures: if a refund request times out after the payment service accepts it, blindly retrying can duplicate the operation. This question distinguishes backend judgment from agent-framework familiarity.
8. How would you reduce inference cost without sacrificing quality?
Ask candidates to optimize a measured workload, not an abstract model bill.
Strong answers begin by identifying cost drivers: input tokens, output tokens, repeated context, retrieval services, tool calls, and infrastructure. They might propose shorter context, output limits, caching, batch processing, or routing simpler tasks to smaller models.
Each choice has a trade-off. Routing introduces classification errors; caching risks stale or cross-user results; aggressive context reduction can remove essential evidence.
For hosted APIs, candidates should consult current rates rather than memorize them. The Anthropic pricing documentation illustrates how model choice and features affect billing.
Evaluate cost per successful task, not just cost per request.
9. When would you self-host an open-weight model?
Good answers compare data constraints, customization, workload predictability, available hardware, and operational capacity.
Candidates should discuss GPU memory, weight precision, KV-cache growth, concurrency, batching, and context length. Quantization can reduce memory requirements, but its quality impact needs measurement on the actual task.
Serving frameworks such as vLLM offer capabilities including efficient memory management and batching; the official vLLM documentation is a useful technical reference.
Avoid rewarding the claim that self-hosting is automatically cheaper or more private. Total cost includes utilization, engineering, monitoring, security, and incident response. Privacy depends on the complete deployment and data-handling design.
10. How would you defend against prompt injection?
Describe a retrieved page that instructs the assistant to reveal secrets or invoke an administrative tool.
Strong answers recognize retrieved text and tool outputs as untrusted data. Defensive design should include least-privilege tools, authorization outside the model, constrained execution, controlled network access, and careful handling of secrets.
Instruction hierarchy and injection detection may help, but neither replaces security boundaries. Ask how the candidate would test indirect injection, data exfiltration attempts, and malicious tool responses.
Warning sign: “We put a stronger instruction in the system prompt, so the problem is solved.”
11. What would you monitor after deployment?
Expect more than uptime and token counts.
A useful monitoring plan includes request failures, latency distributions, retrieval failures, tool errors, refusals, escalation rates, and sampled quality reviews. Traces should identify model, prompt, retrieval configuration, and relevant data versions.
Good candidates balance debugging value against privacy: logging every prompt indefinitely can create a sensitive-data repository.
Ask for a rollback plan. Strong answers include feature flags, staged releases, evaluation gates, and compatible fallbacks. A fallback is not safe merely because it returns text; it must preserve permissions and critical task constraints.
A practical LLM engineer interview exercise
Give candidates a small document collection, a starter repository, and access to either a model endpoint or recorded responses. Ask them to build a question-answering service that returns an answer, supporting sources, and an explicit insufficient-evidence outcome.
Keep the task bounded. Do not require candidates to fund API usage or build a production platform over a weekend.
What to assess
- Correctness: Does the answer follow the supplied evidence?
- Retrieval: Can the candidate explain which passages were selected?
- Validation: Are malformed outputs and dependency failures handled?
- Testing: Are unanswerable and conflicting-evidence cases included?
- Judgment: Are limitations and next steps prioritized clearly?
Introduce one controlled change during discussion: conflicting document versions, tenant-specific access, or a provider timeout. Observe whether the candidate revisits assumptions instead of layering patches onto the original solution.
Allow documentation use. This exercise should measure implementation and reasoning, not memorized SDK syntax.
Step-by-step interview process and scoring
Step 1: Define outcomes and non-negotiables
Write down what the hire should deliver in the first several months. Identify essential capabilities and which skills can be learned on the job.
Step 2: Select a consistent question set
Use the same core scenario and follow-up prompts for comparable candidates. Choose questions aligned with the role rather than attempting to cover every LLM specialty.
Step 3: Collect implementation evidence
Use the practical exercise or a code-review task. For senior candidates, add a production incident or architecture trade-off discussion.
Step 4: Score independently before discussing
Use anchored ratings:
| Dimension | Weak evidence | Solid evidence | Strong evidence |
|---|---|---|---|
| Architecture | Picks tools without constraints | Designs a workable system | Compares alternatives and failure boundaries |
| Evaluation | Relies on a demo | Defines tests and metrics | Controls leakage and analyzes failure segments |
| Reliability | Assumes successful calls | Handles errors and timeouts | Addresses side effects, rollback, and recovery |
| Security | Trusts prompts for enforcement | Adds permissions and validation | Tests adversarial paths and limits impact |
| Economics | Considers only model price | Measures tokens and latency | Optimizes cost per successful outcome |
Step 5: Make an evidence-based decision
Record concrete observations, not impressions such as “seemed senior.” Treat unsafe authorization or side-effect handling as serious gaps for roles owning those systems. A high average score should not erase a critical weakness.
Common mistakes when hiring LLM engineers
- Overweighting transformer trivia: Deep theory matters for some roles, but cannot replace evidence of applied competence.
- Rewarding framework name-dropping: Ask what happens when retrieval, parsing, or tool execution fails.
- Accepting demo-only evaluation: Require tests that could contradict the candidate’s preferred design.
- Ignoring ordinary software engineering: APIs, databases, testing, concurrency, and access control remain central.
- Demanding one “correct” architecture: Several designs may satisfy the same constraints.
- Confusing fluency with ownership: Ask what the candidate personally implemented, measured, and changed.
For adjacent hiring guides, browse more Interview questions topics.
Frequently asked questions
How are LLM engineer interviews different from machine learning engineer interviews?
Applied LLM interviews emphasize pretrained-model integration, retrieval, generation evaluation, tool use, and runtime controls. Broader machine learning interviews may emphasize feature engineering, model training, and statistical validation. The overlap increases for fine-tuning and model-serving roles.
Should an LLM engineer know PyTorch?
PyTorch is important for training, adaptation, and many open-weight-model workflows. It is not necessarily essential for an API-focused application role. Assess the stack the person will own rather than treating one framework as a universal requirement.
What distinguishes a senior LLM engineer?
Senior candidates connect technical choices to business risk and operational outcomes. They define acceptance criteria, anticipate failure modes, design rollback paths, and explain when a simpler non-agentic solution is preferable.
How should candidates prepare for LLM engineer interview questions?
Build one small system and evaluate it rigorously. Practice explaining retrieval failures, model-selection trade-offs, token costs, tool authorization, and deployment monitoring. Bring examples of unsuccessful experiments and what they changed; those often demonstrate more judgment than a polished demo.
Ask the community and get answers from practitioners.