Generative AI and LLM glossary
A practical glossary of generative AI and large language model terminology, connecting core definitions to architecture, procurement, security, and production decisions.
How to use this generative AI and LLM glossary
This generative ai and llm glossary explains the terms that shape AI architecture, vendor selection, and production delivery. For MyDiscussions readers, it connects definitions to practical decisions: when to use retrieval instead of fine-tuning, how to compare model costs, and why an agent needs different controls from a chatbot.
Use it as a shared vocabulary for engineering, product, procurement, and security teams. The central distinction is simple: a language model generates outputs; the surrounding application determines what information it sees, which actions it can take, and how failures are handled.
Foundational model terminology
Generative AI
Generative artificial intelligence refers to models that produce content, including text, images, audio, video, or code, based on patterns learned during training.
Not every AI system is generative. A fraud classifier assigns categories or scores; a generative model might draft an explanation of a flagged transaction. Systems often combine both approaches.
Large language model, or LLM
An LLM is a neural network trained on substantial amounts of language-related data to model and generate token sequences. Many widely deployed LLMs use transformer architectures.
OpenAI GPT models, Anthropic Claude, Google Gemini, and Meta Llama are examples of model families. They differ in capabilities, deployment options, licenses, and commercial terms. “Large” has no universal parameter threshold.
Foundation model
A foundation model is broadly trained and adaptable to multiple downstream tasks. An LLM can be a foundation model, but the category also includes models for images, audio, and other data types.
The business implication is reuse: one underlying model may support search, summarization, document processing, and coding assistance, with different application-level safeguards.
Transformer, attention, and parameters
A transformer is a neural network architecture built around attention mechanisms and related processing layers.
Attention lets the model weigh relationships between positions in its input or intermediate representations.
Parameters are learned numerical values, including model weights. Parameter count alone is a poor purchasing criterion: architecture, training data, inference settings, and task-specific performance also matter.
Tokens, tokenization, and context window
A token is a unit processed by a model, such as part of a word, punctuation, or a representation of non-text content. Tokenization converts inputs into these units.
The context window is the maximum sequence capacity available under a model’s documented limits. Input, output, and reasoning-token accounting vary by provider and model.
A large context window does not guarantee reliable use of every supplied detail. Test whether the model finds the right evidence in realistic documents, not merely whether the documents fit.
Multimodal model
A multimodal model processes or generates more than one modality, such as text and images or speech and text.
Check supported directions separately. A model that understands images may not generate images, and a speech-capable endpoint may have different latency and pricing from a text endpoint.
Training and customization terms
Pretraining and inference
Pretraining establishes a model’s broad capabilities by optimizing it on large datasets.
Inference is the use of a trained model to produce outputs. Most application teams consume inference through APIs or self-hosted services rather than pretraining models themselves.
This distinction matters financially: using an existing model avoids pretraining expense but still requires integration, evaluation, serving, and governance.
Fine-tuning and supervised fine-tuning
Fine-tuning updates model parameters using additional training data. Supervised fine-tuning, or SFT, commonly uses examples of desired input-output behavior.
Fine-tuning can improve format consistency, terminology, or recurring task behavior. It is usually not the best first choice for frequently changing company facts.
Before investing, establish a prompt-based baseline and measure the specific failures that additional training should address.
LoRA and parameter-efficient fine-tuning
Parameter-efficient fine-tuning, or PEFT, adapts a model while training only a subset of parameters or additional components.
Low-rank adaptation, or LoRA, learns compact weight updates rather than updating every original weight. Hugging Face PEFT provides implementations.
The benefit is reduced training resource demand relative to full fine-tuning. The trade-off is additional adapter management, compatibility testing, and deployment complexity.
Alignment, RLHF, and DPO
Alignment broadly describes efforts to make model behavior consistent with intended goals and constraints.
Reinforcement learning from human feedback, or RLHF, uses human preference information within a training process, often involving a learned reward model.
Direct preference optimization, or DPO, trains from preference pairs without the same explicit reinforcement-learning loop. Neither technique guarantees factual accuracy or safe behavior in every deployment.
Distillation and quantization
Distillation trains a model to reproduce aspects of another model’s behavior, often to obtain a smaller or cheaper system.
Quantization represents weights or other numerical values at lower precision to reduce memory use and potentially accelerate inference.
Both can improve deployment economics, but quality changes are task-dependent. Evaluate difficult examples, not just aggregate benchmark scores, and check applicable model terms before distilling outputs.
Prompting and generation controls
Prompt, system instructions, and few-shot learning
A prompt is the input supplied to guide a model’s response. System instructions typically define application-level behavior, subject to the provider’s message hierarchy.
Few-shot prompting includes a small set of examples in the input. Unlike fine-tuning, it does not update model weights.
Examples are useful for extraction schemas and classification boundaries, but they consume context space and may unintentionally teach patterns from unrepresentative samples.
Temperature and top-p
Temperature adjusts the distribution used when sampling outputs. Lower settings generally favor more probable tokens; higher settings generally increase variation.
Top-p, or nucleus sampling, restricts sampling to a set of tokens whose cumulative probability reaches a threshold.
Supported controls differ by model. A temperature of zero should not be treated as a universal reproducibility guarantee, and sampling settings do not establish factual correctness.
Structured outputs and tool calling
Structured outputs constrain responses to a specified format, often a JSON schema. Schema compliance does not establish that field values are correct.
Tool calling, also called function calling, lets a model request an application-defined operation with arguments. The application decides whether and how to execute it.
For example, a support assistant can request lookup_invoice, while server-side code enforces account ownership. See the OpenAI function-calling documentation for a concrete implementation pattern.
Retrieval and knowledge grounding
Embeddings and vector databases
An embedding is a numerical representation that captures useful relationships between items such as text passages.
A vector database indexes these representations for similarity search. Pinecone, Weaviate, and Milvus are examples; PostgreSQL can support vector search through pgvector.
Choose based on filtering, access controls, update frequency, operational complexity, and retrieval quality—not simply the number of vectors supported.
Retrieval-augmented generation, or RAG
RAG retrieves external information and supplies it to a generative model at request time.
A typical workflow searches an authorized document collection, selects relevant passages, and asks the model to answer from that evidence. Frameworks such as LlamaIndex and LangChain provide components for building these pipelines.
RAG improves access to current or private information, but does not eliminate hallucinations. Retrieval can miss evidence, and generation can misinterpret what it receives.
Chunking, hybrid search, and reranking
Chunking divides documents into retrievable units. Useful boundaries include headings, paragraphs, and complete procedures; arbitrary splits can separate rules from exceptions.
Hybrid search combines lexical matching, such as BM25, with vector similarity.
Reranking applies an additional relevance model to reorder retrieved candidates.
For technical documentation, hybrid search can preserve exact identifiers while semantic search captures paraphrases. Reranking may improve relevance at the cost of extra latency and computation.
Grounding and citations
Grounding ties generated output to supplied evidence or another defined source of truth.
Citations identify supporting sources. They are useful only when the referenced material actually supports the associated claim.
Evaluate evidence coverage and citation correctness separately. An answer containing links can still be unsupported.
Agents, tools, and orchestration
AI agent versus workflow
An AI agent uses model-driven decisions to select actions, invoke tools, or continue toward a goal.
A workflow follows more explicitly predefined control paths, although individual steps may use LLMs. Products use “agent” inconsistently, so examine actual autonomy rather than the label.
Use a constrained workflow when steps are predictable. Consider agentic behavior when paths vary substantially and the system can safely detect, contain, and recover from mistakes.
Orchestration, MCP, and guardrails
Orchestration coordinates model calls, retrieval, tools, state, retries, and human approvals. LangGraph and Microsoft Semantic Kernel are examples of relevant frameworks.
The Model Context Protocol, or MCP, standardizes aspects of connecting AI applications to tools and contextual resources. Protocol compatibility does not make a connected service trustworthy; consult the official MCP documentation.
Guardrails are controls around inputs, outputs, and actions. Strong implementations combine validation, permissions, monitoring, and approval gates rather than relying on a safety prompt alone.
Quality, security, and operational terms
Hallucination and evaluation
A hallucination is generated content that is false or unsupported but presented as if reliable.
Evaluation, often shortened to “evals,” systematically measures model or application behavior. Useful dimensions include task success, groundedness, retrieval relevance, refusal appropriateness, and tool-use correctness.
LLM-as-a-judge uses a model to assess outputs. It scales review but can introduce bias and inconsistent scoring, so calibrate it against human judgments.
Prompt injection and data leakage
Prompt injection occurs when untrusted content attempts to redirect a model’s behavior. Indirect injections can appear inside retrieved documents, web pages, or tool responses.
Data leakage includes unauthorized disclosure through generated answers, logs, retrieval results, or tool calls.
Treat retrieved text as data, enforce authorization outside the model, and restrict tool privileges. The OWASP Top 10 for LLM Applications provides a practical security reference.
Latency, throughput, and caching
Latency measures response delay. Time to first token measures how quickly output begins; total completion time captures the full wait.
Throughput measures work processed per unit of time.
Prompt caching can reuse computation for repeated input prefixes, subject to provider rules. Response caching reuses completed answers and requires careful handling of freshness and user permissions.
Open-weight, open-source, and self-hosted models
Open-weight means model weights are available under specified terms. It does not automatically imply unrestricted use or full training transparency.
Open-source AI carries broader expectations around permissions and available components; definitions and vendor terminology require careful reading.
Self-hosting means operating inference infrastructure yourself, potentially using vLLM or Hugging Face Text Generation Inference. It increases control but transfers responsibility for capacity, patching, reliability, and security to your team.
Choosing the right approach
These techniques solve different problems and can be combined.
| Approach | Best initial fit | Main trade-off | Concrete validation criterion |
|---|---|---|---|
| Prompting | Clear tasks with sufficient supplied context | Limited behavioral consistency | Success on representative examples |
| RAG | Current, private, or source-dependent knowledge | Retrieval and indexing complexity | Evidence recall and grounded answers |
| Fine-tuning | Repeated behavioral or formatting failures | Dataset and training maintenance | Improvement on held-out cases |
| Agentic execution | Tasks requiring variable action sequences | Greater security and reliability risk | Successful authorized actions without unsafe side effects |
| Self-hosting | Strong deployment-control requirements | Infrastructure and staffing burden | Quality and latency at expected concurrency |
A step-by-step process for applying the glossary
- Define the task and consequence of failure. Distinguish drafting assistance from actions such as issuing refunds or changing infrastructure.
- Build a representative test set. Include ambiguous requests, missing evidence, unauthorized access attempts, and domain-specific edge cases.
- Establish a simple baseline. Start with a capable hosted or self-hosted model, clear instructions, and validated outputs.
- Diagnose the failure category. Missing knowledge suggests retrieval; inconsistent behavior may justify examples or fine-tuning; unsafe actions require stronger application controls.
- Compare full-system economics. Measure cost per successful task, including retries, retrieval, tool execution, hosting, and human review.
- Pilot with limited permissions. Set acceptance thresholds, monitor quality and latency, and require approval for consequential actions.
- Version and reevaluate. Track model versions, prompts, retrieval settings, and datasets so updates can be tested and rolled back.
Common mistakes to avoid
- Equating context capacity with comprehension: test evidence use across realistic document lengths.
- Using fine-tuning as a live database: frequently changing facts usually belong in retrieval or transactional tools.
- Treating valid JSON as a correct answer: validate business rules and source values separately.
- Giving agents broad credentials: enforce least privilege and transaction limits outside the model.
- Comparing token prices alone: longer outputs, retries, and tool calls can change total cost substantially.
- Assuming vendor privacy terms are identical: verify retention, training use, residency, and contractual commitments for the specific service.
For adjacent terminology in software development, cloud computing, blockchain, and outsourcing, browse more Glossary topics.
Frequently asked questions
What is the difference between generative AI and an LLM?
Generative AI is the broader category of systems that produce content. An LLM focuses on language-related token sequences, although many modern models also support other modalities. Not every generative AI model is an LLM.
Should a business choose RAG or fine-tuning?
Use RAG primarily to supply relevant external knowledge. Consider fine-tuning to improve repeatable behavior, formatting, or task specialization. They can work together, but each should address a measured failure.
Does a larger model always produce better results?
No. A larger model may offer stronger general capabilities, but a smaller model can be sufficient for a narrow task with clear inputs. Compare quality, latency, deployment constraints, and cost using your own evaluation set.
Can an LLM safely execute business actions?
Yes, within a carefully controlled application—not because the model itself guarantees safety. Validate arguments, authorize actions server-side, constrain permissions, record execution, and require human approval where mistakes would have significant consequences.
Ask the community and get answers from practitioners.