GUIDE PROBLEMS AND FIXES

Fixing memory leaks in Node.js apps

Diagnose Node.js memory growth, identify what keeps objects alive, and choose fixes that hold up under production traffic. This guide covers heap snapshots, buffers, caches, listeners, streams, and safe rollout criteria.

Why Node.js memory leaks require more than a restart

For engineering teams, fixing memory leaks in node.js apps means distinguishing legitimate memory demand from unintended retention, locating the references responsible, and proving that memory stabilizes under representative traffic. Restarting a process can restore service, but it does not explain whether the underlying problem is an unbounded cache, abandoned timer, native allocation, or overloaded request queue.

The operational consequences extend beyond an out-of-memory crash. Increasing garbage-collection work can degrade latency before the process fails. Container restarts can interrupt jobs, trigger retries, and multiply load on already strained dependencies.

For decision-makers, the useful deliverable is not simply “a smaller heap.” It is a documented cause, a bounded-memory design, and a regression test tied to realistic workloads. For practitioners, the starting point is knowing which memory measurement is growing.

Understand which Node.js memory is increasing

Node.js uses V8 for JavaScript execution, but JavaScript objects are only part of its memory footprint.

MeasurementWhat it representsHow to interpret growth
heapUsedV8 heap space occupied by JavaScript objectsA rising baseline after comparable collection cycles suggests retained objects
heapTotalMemory currently allocated for the V8 heapExpansion alone does not prove a leak
externalMemory associated with C++ objects bound to JavaScript objectsInvestigate buffers, native modules, and relevant library allocations
arrayBuffersMemory for ArrayBuffer and SharedArrayBuffer, including Node.js buffersInvestigate binary payloads and retained backing stores
rssResident memory for the entire processIncludes heap, code, stacks, and native allocations; may also reflect allocator behavior

Do not add these values together. They overlap: arrayBuffers is included in external, and RSS is a broader resident-memory measurement.

The official Node.js process memory documentation explains these fields and notes that allocator fragmentation can produce sustained RSS growth even when the V8 heap is stable.

Leak, capacity problem, or normal warm-up?

Use concrete criteria rather than a single alarming dashboard point:

  • Normal warm-up: Memory rises as modules load and caches populate, then settles into a repeatable range.
  • Likely JavaScript retention: Under comparable traffic, the low points of heapUsed keep rising across collection cycles.
  • Unbounded buffering: Memory follows queued requests, messages, or pending writes and falls when the backlog clears.
  • Possible native-memory issue: RSS grows while JavaScript heap remains broadly stable.
  • Capacity mismatch: Memory reaches a stable level, but that level exceeds the container’s safe operating budget.

Garbage collection creates a sawtooth pattern. Compare its lower envelope over time, not just its peaks. Likewise, process RSS need not immediately fall after objects become collectible.

Choose tools based on the evidence you need

Start with telemetry already available. Prometheus and Grafana can correlate process memory with throughput, latency, queue depth, and restarts. Datadog, New Relic, or another deployed observability platform can provide similar context, depending on its instrumentation.

For object-level investigation, use:

  • Chrome DevTools with the Node.js inspector: Inspect heap snapshots, allocation profiles, and retaining paths.
  • Node.js built-in heap snapshots: Capture retained object graphs without adding an application dependency.
  • Node.js heap profiling: Use --heap-prof for sampled allocation evidence with a different overhead profile from full snapshots.
  • Diagnostic reports: Capture runtime state around incidents; these complement snapshots but do not replace object-retention analysis.
  • Native profiling tools: Escalate to allocator or platform-specific tools when RSS growth is not explained by V8-visible memory.

An allocation profile shows where memory was allocated; a retaining path explains why it remains reachable. A function that allocates heavily may be working correctly. A small registry elsewhere may be responsible for keeping everything alive.

Never expose the inspector publicly. Restrict access to a trusted local connection or secured tunnel.

A step-by-step process for diagnosing Node.js memory leaks

1. Contain the incident without hiding the evidence

If production is unstable, prioritize continuity:

  • Reduce concurrency or pause memory-heavy jobs.
  • Apply admission control to prevent uncontrolled backlog growth.
  • Drain and restart affected instances in a controlled sequence.
  • Preserve memory charts, deployment identifiers, and workload characteristics.
  • Capture diagnostics only when sufficient headroom exists.

Increasing --max-old-space-size can buy time when the V8 heap limit is the immediate constraint. It is not a leak fix, and it does not cap total process memory. A larger heap may simply move failure from V8 to the container limit.

Keep capacity planning separate from root-cause remediation.

2. Establish a correlated memory baseline

Collect memory alongside workload measurements:

```js

const timer = setInterval(() => {

const memory = process.memoryUsage();

console.log({

time: new Date().toISOString(),

rss: memory.rss,

heapUsed: memory.heapUsed,

heapTotal: memory.heapTotal,

external: memory.external,

arrayBuffers: memory.arrayBuffers,

});

}, 15_000);

timer.unref();

```

This interval is illustrative, not a universal monitoring recommendation. Use your metrics pipeline in production and avoid excessive polling overhead.

Also record:

  • Requests and jobs in flight.
  • Queue length and consumer lag.
  • Active WebSocket connections.
  • Cache entries and, where possible, estimated bytes.
  • Payload sizes and endpoint mix.
  • Worker count and deployment version.

In clustered or worker-based applications, inspect the relevant processes and isolates. A single healthy process can obscure growth in another worker.

3. Reproduce with a representative, bounded workload

Use k6, Artillery, or an existing integration harness. Match the suspected trigger rather than merely increasing request volume.

For example:

  • Exercise unique cache keys if cardinality may be the cause.
  • Reconnect clients repeatedly if connection cleanup is suspect.
  • Include timeouts and cancellations if successful requests remain stable.
  • Repeat job failures if retry state accumulates.
  • Test large binary responses if external rises.

Warm the service first, run a fixed workload, allow outstanding work to drain, and repeat. Testing only one batch may miss cumulative retention.

Separate concurrency from total operations. A service can legitimately use more memory with more simultaneous work while still leaking a smaller amount on every completed operation.

4. Capture comparable heap snapshots safely

For a controlled environment, start Node.js with:

```bash

node --inspect=127.0.0.1:9229 app.js

```

Connect Chrome DevTools and use its Memory panel. Alternatively, supported Node.js versions can generate snapshots on a configured signal:

```bash

node --heapsnapshot-signal=SIGUSR2 app.js

kill -USR2 <pid>

```

The signal approach depends on the operating system and deployment setup.

Take snapshots after warm-up and after repeated equivalent workload cycles. Keep the application state and collection conditions as comparable as possible.

Heap snapshots pause execution and can require substantial additional memory. Prefer a replica removed from traffic or a staging environment with realistic data. The official Node.js heap snapshot guide documents the operational risks.

Snapshots can contain credentials, request bodies, and personal data. Restrict access, encrypt storage, and apply a short retention policy.

5. Follow retaining paths, not just large objects

Compare snapshots for growing object counts and retained size.

  • Shallow size measures the object itself.
  • Retained size estimates memory that would become collectible if that object and its exclusive retention disappeared.

A small array or map can retain a large graph. Follow references back toward a long-lived owner such as a module-level variable, event emitter, timer, cache, or connection registry.

Ask: Which owner should have released this reference, and at what lifecycle boundary?

Inspect objects surviving repeated workload cycles. A large response currently being processed may be legitimate; thousands of completed-request contexts reachable from a global collection are more suspicious.

6. Fix ownership, then rerun the same experiment

Prefer changes that enforce a lifecycle or a hard bound:

  • Remove entries when work completes.
  • Unsubscribe on every termination path.
  • Cancel timers and pending operations.
  • Bound caches and queues.
  • Stream large payloads with backpressure.

Repeat the original workload and snapshot comparison. Do not declare success because memory briefly falls after deployment: a fresh process naturally starts with less retained state.

Common leak patterns and practical fixes

Unbounded maps and application caches

A module-level Map survives for the process lifetime. If every customer, URL, or request produces a new key, it can grow indefinitely.

Choose a cache such as lru-cache with explicit capacity controls. Evaluate:

  • Maximum entries.
  • Maximum estimated bytes for variable-sized values.
  • Expiration behavior and cleanup timing.
  • Eviction effects on database load.
  • Key cardinality under hostile or accidental high-uniqueness traffic.

TTL alone does not guarantee an acceptable memory bound. Many entries can arrive before expiration, and some implementations remove expired entries lazily.

Moving the cache to Redis shifts storage out of Node.js but adds network latency, serialization, operational cost, and its own eviction requirements. It does not remove the need for limits.

Event listeners and timers that outlive requests

This pattern retains request-scoped state:

```js

app.get("/status", (req, res) => {

const onUpdate = update => {

// This closure references the response.

res.locals.latestUpdate = update;

};

sharedEmitter.on("update", onUpdate);

res.json({ ok: true });

});

```

The long-lived emitter retains every registered function and its captured references. For synchronous request-scoped work, remove the listener in finally; for asynchronous work, tie cleanup to completion, failure, and disconnection.

Use named handler references so they can be removed reliably. once() helps only if the event eventually fires; otherwise, the listener remains registered.

Apply the same reasoning to setInterval, delayed retries, subscriptions, and reconnect handlers. Calling unref() on a timer changes whether it keeps the process running; it does not release its callback.

Retained promises, jobs, and request contexts

An unresolved promise is not automatically a leak. The problem occurs when reachable registries, closures, or queues retain its associated state indefinitely.

A pending-operation map should delete entries on both success and failure:

```js

pending.set(jobId, context);

try {

return await runJob(context);

} finally {

pending.delete(jobId);

}

```

This still assumes runJob() settles. Define deadlines and cancellation behavior for stalled work. Where supported, pass an AbortSignal through the operation.

A timeout implemented with Promise.race() does not cancel the losing operation. Without real cancellation or another bound, timed-out work may continue consuming memory.

In Express or NestJS applications, avoid storing entire request objects in singleton services. Extract only the small, necessary fields.

Buffers, streams, and slow consumers

Binary workloads often increase external or arrayBuffers rather than predominantly increasing heapUsed.

Common causes include collecting all chunks before processing, ignoring writable backpressure, and retaining a small Buffer.subarray() view of a much larger buffer.

Use stream.pipeline() when appropriate to connect streams and propagate failures. The official Node.js stream documentation covers backpressure and pipeline behavior.

Copying a small retained slice can release the large backing store sooner, but introduces allocation and copying costs. Streaming lowers peak memory, but complicates retries and response handling after output has started.

For WebSockets and message consumers, bound per-client or per-consumer queues. Decide explicitly whether overload causes pausing, rejection, disconnection, or permissible message dropping.

Verify the fix and set release criteria

A useful acceptance test combines memory behavior with service quality.

CriterionEvidence to collect
Stable memory baselineHeap low points stop trending upward after warm-up
Bounded retained stateCache, listener, connection, and pending-work counts obey intended limits
Workload recoveryBacklogs drain and memory behavior returns to its expected range
Performance preservedLatency, throughput, and error rates meet service targets
Operational headroomPeak process and container usage remain safely below enforced limits

Run the test long enough to exercise relevant lifecycle events: cache expiration, credential refresh, reconnects, scheduled jobs, and repeated failures.

For deployment, use a canary and compare it with an unchanged instance under similar traffic. Confirm that the fix does not exchange memory growth for excessive cache misses, rejected work, or downstream overload.

Document the retaining owner, the missing cleanup or bound, and the regression test. That makes the investigation reusable rather than dependent on one engineer’s memory.

Common mistakes that prolong an investigation

  • Treating high RSS as proof of a JavaScript leak. Check heap, external memory, allocator behavior, and workload first.
  • Forcing garbage collection as a production fix. Reachable objects remain reachable.
  • Suppressing listener warnings. Raising the listener limit can hide evidence without correcting ownership.
  • Taking snapshots on an almost exhausted instance. The diagnostic operation itself may crash it.
  • Testing only successful requests. Cancellation, timeout, retry, and disconnect paths frequently miss cleanup.
  • Replacing every map with a WeakMap. Weak references suit specific ownership relationships, not general cache policy.
  • Using scheduled restarts as the final remedy. They provide containment, not proof of correctness.

For related operational troubleshooting, browse more Problems and fixes topics.

Frequently asked questions

How do I know whether Node.js has a memory leak?

Look for persistent growth under comparable workload conditions, especially rising heap low points across collection cycles. Confirm with retained-object growth or an expanding long-lived collection. One high reading, a larger heap reservation, or elevated RSS alone is insufficient.

Will increasing the Node.js heap limit fix the problem?

No. It may support a legitimately larger working set or temporarily delay failure. Retained objects still accumulate, and native allocations remain outside that limit. Account for total process and container memory before changing it.

Can Node.js leak memory even with garbage collection?

Yes. Garbage collection reclaims unreachable objects, not objects the application no longer needs semantically. Global maps, listener registries, timers, and queues can keep obsolete data reachable indefinitely.

Should I take heap snapshots directly in production?

Only with an explicit risk plan. Snapshots pause execution, consume additional memory, and may expose sensitive information. Prefer a drained replica with headroom; when that is unavailable, start with metrics and sampled profiling before attempting a snapshot.

Have a question about this topic?

Ask the community and get answers from practitioners.

Start a discussion