GUIDE TUTORIALS

Fine-tune LLaMA on custom data

Build a reproducible LLaMA fine-tuning workflow using QLoRA, Hugging Face, and task-specific evaluation. Learn when tuning is worthwhile, how to prepare data, and what to verify before deployment.

What fine-tuning LLaMA should accomplish

To fine-tune llama on custom data effectively, start with a measurable behavior you want to change—not simply a folder of documents. Fine-tuning can teach a model to produce consistent structured output, follow domain-specific workflows, classify specialized requests, or respond in an approved style. It is less suitable as the sole mechanism for maintaining a large, frequently changing knowledge base.

This MyDiscussions tutorial uses Meta Llama 3.1 8B Instruct, Hugging Face Transformers, TRL, PEFT, and bitsandbytes to demonstrate supervised fine-tuning with QLoRA. The workflow targets a single NVIDIA GPU and produces a lightweight adapter rather than a separate, fully trained foundation model.

For decision-makers, the central question is whether improved task performance justifies the ongoing cost of dataset maintenance, evaluation, and serving. For practitioners, success means a reproducible experiment that beats a credible baseline without introducing unacceptable regressions.

Decide whether fine-tuning is the right approach

Before allocating GPU time, compare three approaches against the same representative test cases.

ApproachBest fitMain trade-off
Prompt engineeringInstructions, formatting, and small behavior changesFast to implement, but complex prompts can be brittle
Retrieval-augmented generation, or RAGPrivate documents, citations, and changing informationRequires retrieval infrastructure and evidence-quality evaluation
Fine-tuningRepeated task patterns, specialized terminology, and consistent behaviorRequires curated examples and model lifecycle management
RAG plus fine-tuningSpecialized workflows grounded in current evidenceGreater flexibility, but more components to diagnose

Choose fine-tuning when you can define the desired output and supply representative examples of it. For example, a support organization might train a model to convert incident descriptions into validated JSON containing a product, severity, and routing team.

Prefer RAG when the primary requirement is answering questions from updated product manuals. Fine-tuning does not make memorized facts reliably current, attributable, or removable.

Establish acceptance criteria before training:

  • Schema validity on held-out requests.
  • Correct routing or classification under a documented rubric.
  • Appropriate abstention when essential information is missing.
  • No material regression on relevant general-purpose capabilities.
  • Acceptable latency and serving cost at expected traffic levels.

Use thresholds derived from business risk, not a universal benchmark score.

Select the model, hardware, and training method

Why use an instruct model with QLoRA?

An instruct checkpoint already understands conversational requests, making it a practical starting point for supervised task adaptation.

Full fine-tuning updates the base model’s parameters and requires substantial memory for gradients and optimizer state. LoRA instead trains small low-rank adapter matrices. QLoRA combines these adapters with a frozen, quantized base model, reducing training memory requirements further.

The trade-off is that adapter tuning may be insufficient for some deep behavioral changes. Nevertheless, it is a strong first experiment because it lowers iteration costs and simplifies versioning.

An 8B model with QLoRA can often be trained on a 24 GB GPU using short sequences, small microbatches, and gradient checkpointing. This is a starting estimate, not a guarantee: longer contexts and software differences can increase memory requirements substantially.

Check access and licensing

The example checkpoint is meta-llama/Llama-3.1-8B-Instruct. It requires appropriate access through Hugging Face and acceptance of applicable terms.

Review the official model card and license information before using it commercially or redistributing derived artifacts. Confirm that your organization also has permission to use the training data.

For infrastructure, an AWS EC2 instance with an NVIDIA A10G or another compatible GPU instance from Google Cloud, Azure, or a specialist GPU provider can work. Compare total experiment cost, including failed runs, storage, and data transfer—not just hourly GPU pricing.

Step 1: Define the task and build the dataset

For this walkthrough, the model converts support requests into structured routing decisions.

A training example should demonstrate the actual production contract:

```json

{

"messages": [

{

"role": "system",

"content": "Route support requests. Return only JSON with product, severity, and team. Use null for unknown values."

},

{

"role": "user",

"content": "Our production API returns 503 for every request. Dashboard login still works."

},

{

"role": "assistant",

"content": "{\"product\":\"api\",\"severity\":\"critical\",\"team\":\"platform\"}"

}

]

}

```

Store one object per line in JSONL files. Include more than straightforward examples:

  • Ambiguous requests requiring null values.
  • Messages mentioning multiple products.
  • Informal wording, misspellings, and realistic abbreviations.
  • Low-severity issues that sound urgent.
  • Requests containing instructions that conflict with the routing policy.
  • Legitimate edge cases from production.

A few hundred carefully curated examples can support an initial feasibility test. Broader tasks generally require more coverage; there is no universal minimum dataset size.

Do not manufacture volume by duplicating examples. Repetition can amplify narrow patterns without improving generalization. If using synthetic examples, have domain reviewers check labels and compare their distribution with real traffic.

Remove secrets, unnecessary personal information, and data without valid usage rights. Fine-tuning is not a privacy-preserving storage mechanism: models can reproduce training content.

Step 2: Split and validate before tokenization

Create train.jsonl, validation.jsonl, and test.jsonl. Use training data for optimization, validation data for model selection, and the test set only for the final comparison.

The split must reflect how information overlaps:

  • Keep messages from the same ticket in one split.
  • Group related customers, documents, or templates where needed.
  • Remove exact and near-duplicate examples across splits.
  • Use a later time period as a test set when temporal drift matters.

A random row split can create misleading results if paraphrases of the same incident appear in both training and testing.

Validate required roles, nonempty content, JSON syntax, allowed labels, and contradictory annotations. Inspect token-length distributions with the selected tokenizer.

Choose a sequence limit that preserves both the request and its answer. If truncation removes the target response, the example may contribute little useful training signal.

Step 3: Prepare a reproducible environment

Use a Linux environment with an NVIDIA driver and a CUDA-compatible PyTorch installation. Install PyTorch using the build appropriate for your machine, then add the training libraries:

```bash

python -m venv .venv

source .venv/bin/activate

pip install transformers datasets accelerate peft trl bitsandbytes huggingface_hub

hf auth login

```

The code below uses the modern TRL SFTConfig and SFTTrainer interfaces. Library APIs change, so verify your installed versions against the official TRL supervised fine-tuning documentation.

After a successful smoke test, freeze dependencies:

```bash

pip freeze > requirements.lock.txt

```

Record the GPU type, driver, PyTorch and CUDA versions, model revision, tokenizer revision, and dataset fingerprint. Pin model revisions to immutable commits in repeatable production experiments.

Authenticate through your normal secret-management process. Never commit Hugging Face tokens to a repository or embed them in training data.

Step 4: Load LLaMA with four-bit quantization

Create a training script with the following setup:

```python

import torch

from datasets import load_dataset

from transformers import (

AutoModelForCausalLM,

AutoTokenizer,

BitsAndBytesConfig,

)

from peft import LoraConfig

from trl import SFTConfig, SFTTrainer

model_id = "meta-llama/Llama-3.1-8B-Instruct"

if not torch.cuda.is_available():

raise RuntimeError("This example requires a compatible NVIDIA GPU.")

use_bf16 = torch.cuda.is_bf16_supported()

compute_dtype = torch.bfloat16 if use_bf16 else torch.float16

tokenizer = AutoTokenizer.from_pretrained(model_id)

if tokenizer.pad_token is None:

tokenizer.pad_token = tokenizer.eos_token

tokenizer.padding_side = "right"

quantization = BitsAndBytesConfig(

load_in_4bit=True,

bnb_4bit_quant_type="nf4",

bnb_4bit_use_double_quant=True,

bnb_4bit_compute_dtype=compute_dtype,

)

model = AutoModelForCausalLM.from_pretrained(

model_id,

quantization_config=quantization,

torch_dtype=compute_dtype,

device_map={"": 0},

)

model.config.use_cache = False

model.config.pad_token_id = tokenizer.pad_token_id

```

NF4 quantization reduces storage for frozen base-model weights. Computation still uses floating-point operations; “four-bit training” does not mean every tensor occupies four bits.

This example explicitly places the model on one GPU. Multi-GPU training requires an appropriate distributed configuration rather than simply changing the device map.

Step 5: Configure completion-only supervised training

Convert the conversational examples into prompt/completion records. This walkthrough assumes exactly one final assistant answer per record:

```python

dataset = load_dataset(

"json",

data_files={

"train": "train.jsonl",

"validation": "validation.jsonl",

},

)

def to_prompt_completion(row):

messages = row["messages"]

if len(messages) < 2 or messages[-1]["role"] != "assistant":

raise ValueError("Each example needs a final assistant answer.")

return {

"prompt": messages[:-1],

"completion": messages[-1:],

}

dataset = dataset.map(

to_prompt_completion,

remove_columns=dataset["train"].column_names,

)

lora = LoraConfig(

r=16,

lora_alpha=32,

lora_dropout=0.05,

bias="none",

task_type="CAUSAL_LM",

target_modules="all-linear",

)

config = SFTConfig(

output_dir="outputs/llama-support",

max_length=1024,

per_device_train_batch_size=1,

per_device_eval_batch_size=1,

gradient_accumulation_steps=16,

learning_rate=1e-4,

num_train_epochs=2,

warmup_ratio=0.03,

lr_scheduler_type="cosine",

gradient_checkpointing=True,

bf16=use_bf16,

fp16=not use_bf16,

logging_steps=10,

eval_strategy="epoch",

save_strategy="epoch",

load_best_model_at_end=True,

metric_for_best_model="eval_loss",

greater_is_better=False,

completion_only_loss=True,

packing=False,

report_to="none",

seed=42,

)

trainer = SFTTrainer(

model=model,

args=config,

train_dataset=dataset["train"],

eval_dataset=dataset["validation"],

processing_class=tokenizer,

peft_config=lora,

)

trainer.train()

trainer.save_model("artifacts/llama-support-adapter")

tokenizer.save_pretrained("artifacts/llama-support-adapter")

```

TRL applies the tokenizer’s chat template to conversational prompt/completion records. Completion-only loss focuses optimization on the desired answer rather than training the model to reproduce the incoming request.

Inspect a tokenized sample and its label mask before a full run. Confirm that prompt tokens are ignored and the assistant completion—including its termination marker—remains trainable.

The settings are starting points, not recommended optima. On one GPU, the effective batch size is 16 examples before accounting for variable sequence lengths. Compare a small learning-rate sweep and one-to-three-epoch runs instead of assuming longer training improves quality.

Use the official PEFT quantization guide when adapting this workflow to custom training loops, where explicit preparation for k-bit training may be necessary.

Step 6: Evaluate behavior, not just loss

Run a small smoke test before committing to the full dataset. Verify that training progresses, validation executes, and the saved adapter reloads successfully.

Then compare the tuned model with the untouched instruct checkpoint using:

  • Identical prompts and chat templates.
  • Identical generation settings.
  • The same held-out examples.
  • A rubric defined before inspecting results.

For support routing, measure JSON parse rate, schema validity, per-field accuracy, and critical misrouting. Report results separately for ambiguous requests, rare labels, and adversarial instructions.

Use deterministic decoding, such as do_sample=False, for stable classification comparisons. Validate parsed JSON rather than raw string equality, because whitespace and key order need not affect correctness.

Validation loss helps identify training problems, but it does not establish business value. A model can have lower loss while becoming overconfident or routing rare incidents incorrectly.

For free-text tasks, combine automated checks with blinded domain-expert review. Treat model-based graders as fallible evaluators, not ground truth.

Step 7: Package and deploy the adapter

The saved artifact is a LoRA adapter, not a standalone copy of LLaMA. Serving requires the compatible base checkpoint, the adapter, and the correct tokenizer.

Two deployment options are common:

  • Load the base plus adapter: easier adapter switching and smaller individual artifacts, but the serving engine must support the configuration.
  • Merge the adapter into a base model: simpler single-model packaging, but merging requires suitable precision and memory. Requantize afterward if needed, then reevaluate.

Transformers can support functional integration testing. vLLM or Hugging Face Text Generation Inference may suit higher-throughput serving; check current support for your model, adapter, and quantization combination.

Keep the production system prompt aligned with training. Benchmark realistic concurrency, prompt lengths, and output limits. Training memory measurements do not predict serving capacity, especially when the inference KV cache grows.

Roll out gradually with an immediate fallback to the previous model. For routing, validate outputs in application code and require human approval for high-impact actions.

Version the adapter alongside its base revision, dataset version, prompt contract, and evaluation report.

Common mistakes that undermine results

  • Treating documents as supervised examples: raw text is not equivalent to demonstrated task behavior.
  • Training on inconsistent labels: unresolved reviewer disagreements become model behavior.
  • Ignoring truncation: removed answers create weak or unusable training examples.
  • Changing templates between training and serving: mismatched role markers can degrade output.
  • Selecting solely by training loss: memorization is not held-out task performance.
  • Skipping baseline comparisons: a better prompt may solve the problem more cheaply.
  • Assuming constrained JSON means correct decisions: structural validity does not establish semantic accuracy.
  • Deploying without rollback: every release needs a recoverable artifact and observable failure criteria.

Frequently asked questions

How much data do I need to fine-tune LLaMA?

Start with a few hundred high-quality examples for a narrow feasibility test. Expand according to uncovered failure modes, not an arbitrary row count. Diverse inputs, consistent labels, and representative edge cases matter more than duplicated volume.

Can I fine-tune LLaMA on one GPU?

Often, yes. QLoRA makes an 8B checkpoint practical on some single-GPU setups, including approximately 24 GB configurations with conservative sequence lengths. Run a smoke test first; memory depends on context length, microbatch size, checkpointing, and implementation details.

Should I fine-tune LLaMA or use RAG?

Use fine-tuning primarily to change behavior and RAG primarily to supply retrievable knowledge. Combine them when the model must follow a specialized workflow while grounding answers in current documents. Evaluate the combined system end to end.

Does QLoRA produce a complete deployable model?

It usually produces adapter weights that depend on the original base model. You can serve both together or merge the adapter into a suitable base checkpoint. Test the actual serving artifact again after merging or quantization.

Make the first experiment a decision tool

Begin with one task, a defensible baseline, and a small curated dataset. Continue only if the tuned model improves held-out outcomes without unacceptable regression, latency, or operational complexity.

For related implementation walkthroughs, browse more Tutorials topics.

Have a question about this topic?

Ask the community and get answers from practitioners.

Start a discussion