AI and machine learning terms explained
Decode the AI and machine learning vocabulary that shapes architecture, procurement, and delivery decisions. Learn how core concepts connect to practical tools, measurable outcomes, and production trade-offs.
Why AI terminology matters in real projects
Having ai and machine learning terms explained in practical language helps teams distinguish a useful capability from a marketing label. For MyDiscussions readers evaluating platforms, building applications, or commissioning software, the vocabulary should clarify what a system learns, how it produces results, and what evidence makes those results trustworthy.
This glossary connects foundational concepts to engineering and procurement decisions. Rather than treating every term as an isolated definition, it shows how terminology affects data requirements, architecture, evaluation, operating costs, and risk.
AI, machine learning, and deep learning: the foundations
Artificial intelligence versus machine learning
Artificial intelligence (AI) is the broad field of building systems that perform tasks associated with intelligence, such as interpreting language, planning actions, or recognizing objects. AI includes both learned approaches and explicitly programmed methods, such as rule-based expert systems.
Machine learning (ML) is a subset of AI in which systems learn patterns from data rather than relying exclusively on manually written rules.
The practical distinction is important: a rules engine may be more predictable and easier to audit than a learned model when requirements are explicit. ML becomes attractive when useful patterns are too complex or variable to encode economically.
Deep learning and neural networks
A neural network is a parameterized model built from connected computational layers. Deep learning uses neural networks with multiple layers to learn increasingly complex representations.
Deep learning underpins much modern language, vision, and speech technology. However, it is not automatically the best choice for structured business data. Gradient-boosted tree frameworks such as XGBoost and LightGBM are strong candidates for tabular prediction problems.
PyTorch and TensorFlow support neural-network development. scikit-learn provides preprocessing, conventional ML algorithms, and evaluation utilities. Choose frameworks around the problem, deployment environment, and team expertise—not the perceived sophistication of the label.
Models, algorithms, parameters, and hyperparameters
These terms describe different parts of the learning process:
- Algorithm: A computational procedure, such as a learning method that fits a model.
- Model: The fitted artifact used to produce predictions or generated outputs.
- Parameter: A value learned during training, such as a neural-network weight.
- Hyperparameter: A configuration selected outside parameter fitting, such as learning rate or tree depth.
- Feature: An input variable or representation used by a model.
- Label: The target value supplied for supervised learning.
A customer-churn system might use account age and support history as features, cancellation as its label, and tree depth as a hyperparameter.
How machine learning systems learn
Supervised, unsupervised, and self-supervised learning
Supervised learning learns from examples paired with target outputs. Classification predicts categories, such as fraudulent versus legitimate transactions. Regression predicts numerical values, such as delivery time.
Unsupervised learning identifies structure without supplied target labels. Clustering groups similar records; dimensionality reduction compresses representations. Clusters are not automatically meaningful customer segments: someone must validate their stability and business usefulness.
Self-supervised learning creates training targets from the data itself. Predicting a missing word or the next token lets a system learn from text without humans labeling every example.
Self-supervised learning is central to many foundation models, but it does not mean human involvement disappears. Data selection, filtering, evaluation, and later adaptation still require deliberate choices.
Reinforcement learning and RLHF
Reinforcement learning (RL) trains an agent to choose actions based on rewards obtained through interaction with an environment. Relevant concepts include a policy, which determines behavior, and a reward function, which specifies the training objective.
Reinforcement learning from human feedback (RLHF) uses human preference information to guide model behavior, often through a learned reward model and reinforcement-learning optimization.
Not every preference-based training method is RLHF. For example, direct preference optimization adapts models using preference pairs without the same reinforcement-learning loop.
The key trade-off is objective design: optimizing an imperfect reward can encourage behavior that scores well without satisfying the underlying need.
Training, validation, and test data
Training adjusts model parameters. Validation supports choices such as model selection, hyperparameter tuning, and decision thresholds. Testing estimates performance on data kept separate from those choices.
Overfitting occurs when a model learns training-specific patterns that do not generalize. Underfitting occurs when it fails to capture important structure.
Data leakage happens when information unavailable at prediction time improperly enters training or evaluation. Examples include using post-cancellation account fields to predict cancellation or placing records from the same patient in both training and test sets.
Use time-based or group-based splits when the deployment setting demands them. Random splitting is not universally appropriate. The scikit-learn guide to common pitfalls explains how preprocessing and leakage can distort results.
Generative AI and large language model terms
Generative AI, foundation models, and LLMs
Generative AI produces content such as text, images, audio, or code. Predictive AI commonly refers to systems that estimate outcomes or assign categories, although generative models can also perform prediction tasks.
A foundation model is trained broadly enough to support adaptation across multiple tasks. A large language model (LLM) is a language-focused model capable of processing and generating token sequences.
OpenAI, Anthropic, and Google provide hosted model services. Hugging Face distributes models and tooling, including Transformers. Availability of model weights does not by itself establish unrestricted commercial use: review each model’s license.
Tokens, context windows, and inference
A token is a unit processed by a model. Depending on the tokenizer, it may represent a word fragment, punctuation mark, or another unit. Token counts are not equivalent to word counts and vary across languages and content types.
The context window is the sequence capacity available to a model invocation, subject to the provider’s input and output limits. A larger window allows more material, but does not guarantee reliable attention to every detail.
Inference is running a trained model to obtain an output. In production, distinguish:
- Latency: Time required to return a result.
- Time to first token: Delay before streaming output begins.
- Throughput: Work completed per unit of time.
- Concurrency: Requests being handled simultaneously.
These measures interact. Batching may improve throughput while increasing waiting time.
Prompts, temperature, and hallucinations
A prompt supplies instructions, examples, or context. Prompt engineering is the systematic design and testing of those inputs.
Temperature controls aspects of sampling randomness. Lower settings generally make outputs less variable, but they do not guarantee correctness or identical results across infrastructure and model changes.
A hallucination is generated content that is false, unsupported, or fabricated yet presented as plausible. It is not merely an unusual writing style.
For factual applications, evaluate claims against evidence. Requiring citations helps only if the cited sources exist and actually support the answer.
Embeddings and vector databases
An embedding represents an item—such as text or an image—as a numerical vector. Similarity measures can then identify related items.
A vector database stores and searches vectors, often alongside metadata filters. Examples include Pinecone, Weaviate, and Milvus. PostgreSQL’s pgvector extension adds vector capabilities to an existing relational database.
Vector similarity does not establish factual equivalence. For identifiers, product codes, and exact terminology, hybrid search combining lexical and vector retrieval may outperform either alone.
RAG, fine-tuning, and agents: choosing an architecture
Retrieval-augmented generation versus fine-tuning
Retrieval-augmented generation (RAG) retrieves relevant information and supplies it to a model as context. It is useful for answering questions over changing documentation or access-controlled organizational knowledge.
Fine-tuning updates a pretrained model’s parameters using additional training examples. It can improve task behavior, formatting, or domain-specific patterns, but is not a dependable substitute for a frequently updated knowledge store.
| Approach | Best suited to | Main trade-off | Evidence to collect |
|---|---|---|---|
| Prompting | Clear tasks with adequate existing model capability | Sensitive to wording and model changes | Task success across representative prompts |
| RAG | Answers grounded in external documents | Retrieval and permissions add complexity | Retrieval recall and source-supported answers |
| Fine-tuning | Repeated behavioral or formatting requirements | Training data and maintenance burden | Improvement over a strong prompt baseline |
| Tool-using agent | Workflows requiring actions across systems | More failure paths and security exposure | End-to-end success and safe failure handling |
These approaches can be combined. A fine-tuned model can use retrieval and call tools.
AI agents and tool calling
Tool calling lets a model request structured operations, such as querying inventory or creating a ticket. Application code must validate and execute those requests.
An AI agent typically combines a model with tools, state, and an execution loop to pursue a goal. The term is used inconsistently, so ask vendors what decisions the system makes autonomously.
Frameworks such as LangGraph and LlamaIndex can support orchestration. They do not remove the need for authorization checks, bounded retries, observability, or human approval.
Prefer a deterministic workflow when the sequence of actions is known. Add agentic decision-making only where flexible planning produces measurable value.
Evaluation, deployment, and governance vocabulary
Metrics that reflect the actual task
Accuracy is the proportion of correct classifications, but it can mislead when classes are imbalanced.
Precision measures how many predicted positives are truly positive. Recall measures how many actual positives are detected. F1 combines precision and recall, but does not encode your organization’s financial or safety costs.
Calibration describes whether predicted probabilities match observed frequencies. A well-ranked model can still produce poorly calibrated probabilities.
For generative applications, measure task-specific outcomes: source support, correct tool arguments, required-field completion, or successful issue resolution. Automated model-based grading can help, but validate graders against human-reviewed examples.
MLOps, drift, and monitoring
MLOps applies operational practices to the ML lifecycle: versioning, deployment, monitoring, and controlled updates. MLflow supports experiment tracking and model lifecycle management; cloud platforms such as Amazon SageMaker and Google Vertex AI provide managed capabilities.
Data drift is a change in input distributions. Concept drift is a change in the relationship between inputs and outcomes. Neither should be diagnosed solely from declining user satisfaction.
Record model versions, data lineage, prompts, retrieval configuration, and evaluation results. Monitor quality as well as uptime.
Explainability, privacy, and prompt injection
Explainability concerns understanding model behavior. Tools such as SHAP estimate feature contributions, but those explanations are not proof of causality.
Prompt injection occurs when untrusted content attempts to redirect an AI system’s behavior. Retrieved pages, documents, and tool outputs should be treated as data—not trusted instructions.
Evaluate data retention, access controls, training-use policies, and deletion support before sending sensitive information to a service. The NIST AI Risk Management Framework provides a structured foundation for identifying and managing AI risks.
A step-by-step process for applying these terms
- Step 1: Define the decision or task. Specify inputs, expected outputs, users, and unacceptable failures. “Adopt AI” is not a testable requirement.
- Step 2: Establish a baseline. Compare against manual work, rules, search, or a conventional ML model before introducing complex architectures.
- Step 3: Audit the data. Check label quality, permissions, freshness, representativeness, and leakage. Confirm that required features exist at inference time.
- Step 4: Select the simplest viable approach. Test prompting before fine-tuning; test a fixed workflow before an autonomous agent.
- Step 5: Build representative evaluations. Include routine cases, rare high-impact cases, ambiguous requests, and adversarial inputs. Set thresholds using actual error costs.
- Step 6: Measure operational economics. Track cost per successfully completed task, tail latency, human-review effort, and failure recovery—not just token price.
- Step 7: Pilot with controls. Limit permissions, require approval for consequential actions, and define rollback conditions before expanding access.
Common mistakes when comparing AI systems
- Equating benchmark scores with business performance. Public benchmarks may not resemble your documents, users, or error costs.
- Treating RAG as a hallucination cure. Retrieval can miss evidence, retrieve outdated content, or expose unauthorized documents.
- Fine-tuning before diagnosing failures. Missing context and poor retrieval usually require different fixes from inconsistent formatting.
- Assuming open weights mean low total cost. Hosting adds accelerator capacity, operations, security, and licensing responsibilities.
- Accepting vendor terminology without operational definitions. Ask what “reasoning,” “autonomous,” and “enterprise-ready” mean in demonstrable behavior.
Frequently asked questions
Is machine learning the same as AI?
No. AI is the broader field; machine learning is one approach within it. A system based entirely on explicit rules can qualify as AI without learning from data.
Do all AI applications need an LLM?
No. Forecasting, fraud scoring, optimization, and image inspection may use other model families—or no learned model at all. Start with the task and compare simpler alternatives.
Should we use RAG or fine-tuning for company knowledge?
RAG is usually the starting point for changing, attributable company information. Fine-tuning is more appropriate for learned behavior and output patterns. Neither replaces document permissions or factual evaluation.
Which AI terms matter most during vendor selection?
Prioritize inference, context limits, evaluation, latency, data retention, and deployment controls. Require vendors to demonstrate those concepts on representative workloads. For related technology definitions, browse more Glossary topics.
Ask the community and get answers from practitioners.