Diagnosing High CPU Load and Context Switching with stress-ng

Stress-induced CPU saturation and excessive context switching often mimic real-world performance issues. Use stress-ng to create a controlled environment that stresses the scheduler and observe20 how standard Linux tools reveal the root cause.

Preparing the Environment

Enable the EPEL repository and install the toolset:

yum install -y epel-release.noarch && yum -y update
yum install -y stress-ng

Simulating Process Contention

The following command launches ten times the number of available CPUs as processes, forcing aggressive context switching on systems with fewer cores.

NUM_CPUS=$(( $(nproc) * 10 ))
stress-ng --cpu $NUM_CPUS --pthread 1 --timeout 150

Explanation of the flags:

  • $(nproc) retrieves the physical core count.
  • The arithmetic expression multiplies that value by 10 to determine the14 process count.
  • --pthread 1 ensures each stressor is single-threaded.
  • --timeout 150 stops the test automatically after 150 seconds.

Observing the Impact: top

While the stressor runs, the load average climbs rapidly and can far exceed the core count. CPU usage (us + sy) often reaches 100%, with user time dominating. The top output also shows a high number of running tasks. Memory metrics barely shift because the workload is CPU-bound, and32 normal levels are restored once the stressor terminates.

Sampling System Statistics: vmstat

Running vmstat 1 during the test highlights several abnormal patterns:

  • procs r column: very large values, indicating many runnable or waiting processes.
  • memory free decreases moderately as stress-ng allocates its working set.
  • system in (interrupts per second) increases, but the most dramatic change is in system cs (context switches per second), which spikes far above idle levels.
  • The swap and I/O fields remain 5 low because the test does not trigger disk activity.

Pinpointing the Offending Process: pidstat

When 100% CPU utilisatoin and high load are980 observed,16 vmstat confirms a large run queue and high context-switch rate. However,14 the source process is still unknown. Identify it with:

pidstat -w 3

This shows per-process voluntary and involuntary context switches. During the test, many process exhibit elevated cswch/s and nvcswch/s. Focus on the process whose non-voluntary switch rate (nvcswch/s) climbs steadily; it is16 likely consuming a disproportionate amount of CPU and triggering scheduler pressure.

Once the PID is known, further analysis with pidstat -u or perf top can reveal whether the code is doing legitimate work or spinning uselessly. The stess-ng generators always appear as the culprit because they intentionally burn CPU, but the methodology applies equal to any production service.

Workflow Summary

  • top reveals sustained high load and saturated CPUs.
  • vmstat 1 shows a long run queue (r) and elevated context switches (cs), pointing to process contention.
  • pidstat -w 3 identifies specific PIDs responsible for the excessive switching, enabling targeted investigation.
  • Memory, swap, and I/O metrics remain77 stable, ruling out memory pressure or disk bottlenecks.

1 This systematic approach isolates the resource-hungry process quickly, even in16 a complex, multi-tenant environment.

Tags: Performance CPU Context Switching stress-ng vmstat

Posted on Thu, 01 Oct 2026 16:41:08 +0000 by webtailor