GUIDE PROBLEMS AND FIXES

Fixing slow CI/CD pipelines

Slow pipelines rarely have a single cause. Learn how to locate the critical path, prioritize fixes, and improve delivery speed without weakening release confidence.

Why CI/CD pipelines become slow

For engineering teams, fixing slow ci/cd pipelines means reducing the time from a code change to trustworthy feedback and a deployable release—not simply making individual jobs finish faster. A five-minute build provides little value if it waits behind other builds, repeats unnecessary work, or requires an unpredictable approval cycle.

Delays usually accumulate across several layers: runner queues, dependency downloads, compilation, tests, artifact handling, deployment orchestration, and human gates. Teams often optimize the most visible step while overlooking the actual bottleneck.

The right approach combines measurement with architectural changes. Decision-makers need to understand delivery impact and operating cost; practitioners need enough evidence to change the correct job, dependency, or execution environment.

Define what “slow” means before changing anything

A pipeline is slow when its delivery latency exceeds an agreed objective for that workflow. There is no universal acceptable duration: a documentation update and a regulated production release require different checks.

Separate these measurements:

  • Queue time: From a job becoming eligible to a runner starting it.
  • Execution time: From runner startup to job completion, including setup and cleanup.
  • Feedback latency: From submission to the first actionable failure or successful required checks.
  • Release lead time: From an eligible change to a verified deployment.
  • Compute consumption: Total runner time across jobs, including retries and duplicated runs.

Track median and tail latency, such as p50 and p95, by workflow and branch type. Do not mix canceled, failed, and successful runs without labeling them. A fast failure and a completed deployment represent different outcomes.

Set explicit objectives. For example, a team might require ordinary pull requests to return required checks within ten minutes. That is a team-selected target, not an industry benchmark.

Also define guardrails: acceptable infrastructure cost, test coverage expectations, security checks, and deployment failure rates. Faster pipelines that produce unreliable releases are not an improvement.

Diagnose the bottleneck from evidence

Inspect a representative set of recent runs, including peak development hours and unusually slow executions. One successful run can hide queue saturation, intermittent downloads, or repeated test failures.

Collect job timestamps, runner specifications, cache outcomes, test durations, artifact sizes, and retry counts. GitHub Actions, GitLab CI/CD, Jenkins, Azure Pipelines, and CircleCI expose useful execution data, although collecting queue and infrastructure metrics may require additional instrumentation.

Use a symptom-to-cause map

SymptomLikely causeEvidence to collectFirst intervention
Jobs wait before startingInsufficient runner capacity or concurrency limitsQueue duration, active jobs, quota usageAdjust capacity or scheduling
Setup dominates executionRepeated dependency installation or cold imagesDownload time, cache hits, image-pull durationImprove caching and runner images
Tests dominate every runSerial execution, imbalance, or expensive fixturesPer-test timing, shard duration, setup timeBalance shards and optimize fixtures
CPU remains busy throughout buildsCompilation bottleneck or ineffective incremental buildsCPU utilization, task execution logsTune build parallelism or reuse outputs
CPU is idle while jobs crawlNetwork, disk, or external service waitsI/O latency, network timings, timeout logsRemove or localize dependencies
Fast jobs wait for unrelated stagesOverly sequential workflow designDependency graph and timestampsReplace broad barriers with explicit dependencies
Deployment time varies widelyApprovals, rollout behavior, or environment contentionGate wait time, rollout events, lock durationFix the specific release bottleneck

Measure the critical path: the longest chain of dependent work that determines completion time. Reducing a job outside that path may save money without making the pipeline finish earlier.

A step-by-step process for fixing slow CI/CD pipelines

Step 1: Establish a reproducible baseline

Record workflow definitions, runner types, dependency lockfiles, and representative commits. Compare changes using equivalent workloads rather than unrelated pull requests.

Include both warm-cache and cold-cache runs. A cache-only benchmark can disguise poor performance after dependency updates, cache eviction, or runner replacement.

Store a small scorecard:

  • End-to-end p50 and p95 duration.
  • Queue time and critical-path execution time.
  • Runner minutes per successful workflow.
  • Retry frequency and known flaky-test failures.
  • Deployment success and rollback outcomes.

For decision-makers, connect delays to observable costs: developer interruptions, blocked release windows, or accumulated review queues. Avoid treating all waiting time as recoverable labor; developers may work on other tasks while checks run.

Step 2: Remove redundant executions

Before buying larger runners, identify work that should not run at all.

Common examples include duplicate branch and pull-request workflows, superseded commits still consuming capacity, and full application builds for documentation-only changes.

Use workflow concurrency controls to cancel obsolete pull-request runs. On GitHub Actions, concurrency groups can provide this behavior; GitLab offers auto-cancellation options.

Apply cancellation carefully. Do not automatically interrupt a production deployment midway unless its deployment mechanism supports safe cancellation and recovery.

Path filtering can skip irrelevant work, but shared configuration must trigger dependent checks. A change to a root lockfile, base container image, or shared library may affect many services. Ensure skipped workflows also interact correctly with required-check rules.

Step 3: Restructure the workflow around dependencies

A stage-based pipeline often forces unrelated jobs to wait. If packaging needs compilation but not documentation validation, represent those dependencies separately while retaining all release gates.

GitHub Actions and GitLab support dependency relationships through needs. Jenkins Pipeline supports parallel branches.

Run independent checks concurrently, then join them at the appropriate release boundary. Keep fast, high-signal checks early: formatting, type checking, configuration validation, and targeted unit tests often identify failures before expensive integration environments start.

Parallelism has limits. More jobs can increase queueing, duplicate setup, and overload databases or shared test services. Optimize wall-clock duration and total resource consumption together.

Step 4: Make caching correct before making it aggressive

Distinguish dependency caches, build caches, and artifacts:

  • Dependency caches reuse downloaded packages.
  • Build caches reuse outputs from previously executed build tasks.
  • Artifacts transfer specific outputs between jobs or retain release evidence.

Cache keys should reflect compatibility boundaries, including operating system, architecture, toolchain version, and relevant lockfiles. Build-output caches additionally need accurate input tracking.

GitHub’s dependency caching documentation explains key matching and access restrictions. Check platform-specific behavior rather than assuming caches are private, permanent, or available across every branch.

Measure restore and upload costs. A large cache can cost more time to transfer than it saves. Avoid caching temporary directories simply because they are large.

For remote build caches in Bazel, Gradle, or Nx, verify input completeness and access controls. Untrusted pull requests should not be able to publish outputs consumed by trusted release jobs.

Step 5: Optimize dependency installation and container builds

Pin dependencies and use reproducible installation modes, such as npm ci for Node.js projects. Prefer a package-manager download cache over reusing an installed dependency tree unless that tree’s compatibility is well understood.

Consider prebuilt runner images containing stable toolchains. This reduces repeated setup but introduces image patching and maintenance obligations.

For containers:

  • Keep the build context small with .dockerignore.
  • Copy dependency manifests before frequently changing source files.
  • Use multi-stage builds to separate build tooling from runtime contents.
  • Use BuildKit cache mounts or external caches where appropriate.

Docker’s build cache optimization guide describes these techniques. Remember that cache invalidation propagates through dependent layers: copying the entire repository too early can invalidate an otherwise reusable dependency-installation step.

Keep images and registries geographically close to runners when possible. Faster compilation cannot compensate for repeatedly pulling large images across a slow network.

Step 6: Rebalance and isolate the test suite

Start with test-duration reports rather than guessing which framework is slow. JUnit XML and framework-specific reporters can reveal expensive fixtures, long integration tests, and uneven shards.

Split suites using historical duration, not file count alone. Ten short files and ten long files do not create balanced workers.

Tools such as pytest-xdist, Jest workers, and Playwright sharding support concurrent testing, but each needs resource-aware configuration. Browser tests may exhaust memory before CPU capacity; integration tests may compete for database connections.

Make tests independent through isolated schemas, unique identifiers, and explicit cleanup. Shared mutable state turns parallel execution into intermittent failure.

Treat retries as a diagnostic signal, not a performance solution. Quarantining a flaky test should have an owner, an expiry condition, and compensating coverage where necessary. Otherwise, apparent speed improvements merely remove confidence.

Step 7: Right-size runners and capacity

Distinguish larger runners from more runners.

Larger runners can help CPU-bound compilation or memory-constrained tests. More runners primarily address queueing and independent jobs. Neither necessarily improves an external API bottleneck.

Compare configurations using cost per completed workflow, not just per-minute pricing. A more expensive runner may finish sufficiently faster to justify itself; additional cores may also sit unused.

Self-hosted runners provide control over hardware, local caches, and network placement. They also require patching, isolation, autoscaling, and cleanup. Persistent workspaces can hide undeclared dependencies or leak data between jobs.

For autoscaled pools, measure provisioning latency and consider limited warm capacity if startup dominates. Apply spending limits and monitor organization-level concurrency quotas.

Step 8: Shorten deployment without removing safeguards

Build an immutable artifact once and promote that artifact through environments. Rebuilding separately for staging and production wastes time and can produce different outputs.

Separate deployment waiting from deployment execution. Approval delays, environment locks, registry transfers, database migrations, and readiness checks require different fixes.

For Kubernetes, inspect scheduling events, image pulls, readiness probes, and rollout progress. The official Deployment documentation explains rollout behavior and progress reporting.

Do not shorten readiness probes merely to make dashboards look faster. Likewise, do not remove security scans or migration checks without an approved alternative. Optimize their scope, reuse valid results, or move suitable analysis earlier while preserving required release controls.

Choose changes by impact, risk, and maintenance

Prioritize fixes using four criteria:

  • Critical-path impact: Will this reduce end-to-end latency or only background work?
  • Frequency: Does it help most commits or only rare release scenarios?
  • Correctness risk: Could stale outputs, missed tests, or unsafe cancellation escape?
  • Ownership cost: Who maintains the runner image, cache policy, or test-selection logic?

Start with low-risk waste removal and instrumentation. Then address dependency structure, caching, and test balance. Reserve complex changes—remote execution, elaborate affected-project graphs, or custom runner fleets—for workloads that justify their operational cost.

Monorepo tools such as Nx and Bazel can avoid rebuilding unaffected targets, but they depend on accurate dependency graphs. Dynamic imports, generated code, and undeclared configuration dependencies can make selective execution unsafe.

Roll out changes incrementally, retain a straightforward way to restore the previous configuration, and compare representative runs after each intervention.

Common mistakes that make pipelines worse

Optimizing average duration alone. Tail latency may remain unacceptable because peak-hour queues and flaky tests dominate the worst experiences.

Adding unlimited parallelism. Excess workers create CPU contention, memory pressure, service throttling, and more setup overhead.

Using broad fallback caches for incompatible outputs. A fallback may be appropriate for reusable package downloads but unsafe for compiled outputs without compatibility guarantees.

Hiding retries inside scripts. Silent retry loops make unstable tests and unreliable services look like unexplained slowness.

Skipping checks without dependency awareness. File filters and affected-test selection need validation against shared inputs.

Ignoring artifact transfer. Uploading entire workspaces between jobs can erase gains from faster builds.

Maintaining no owner for performance. Pipeline latency regresses as suites and repositories grow. Assign responsibility for reviewing duration trends, cache effectiveness, and flaky-test backlogs.

Frequently asked questions

What should we fix first in a slow CI/CD pipeline?

Identify whether queueing or execution dominates, then inspect the critical path. If runners are unavailable, focus on capacity and redundant runs. If execution dominates, measure setup, build, test, and artifact-transfer time separately. Fix the largest recurring, avoidable delay before tuning small commands.

Does caching always make CI/CD faster?

No. Cache lookup, download, extraction, and upload all consume time. Caching helps when reuse savings exceed that overhead and hit rates are sufficient. Measure cold and warm runs, control cache size, and ensure keys prevent incompatible outputs from being reused.

Is it safe to run fewer tests on pull requests?

It can be, provided test selection uses a trustworthy dependency model and matches the project’s risk requirements. Keep essential checks mandatory, validate selection against periodic full runs, and run broader suites before appropriate release gates. Security-sensitive or heavily coupled changes may require full coverage.

Should we switch CI/CD providers to improve performance?

Only after separating provider constraints from workflow design. Migration may help with runner availability, hardware options, or network placement, but it will not automatically fix serial dependencies, bloated images, or flaky tests. Benchmark equivalent workloads and include migration effort, security controls, and operating costs.

Make pipeline performance a maintained capability

Sustainable improvements come from removing unnecessary work, shortening the critical path, and preserving trustworthy feedback. Define latency objectives, measure representative workloads, and evaluate every change against cost and correctness—not duration alone.

For related delivery and operations troubleshooting, browse more Problems and fixes topics.

Have a question about this topic?

Ask the community and get answers from practitioners.

Start a discussion