Stress-induced CPU saturation and excessive context switching often mimic real-world performance issues. Use stress-ng to create a controlled environment that stresses the scheduler and observe20 how standard Linux tools reveal the root cause.
Preparing the Environment
Enable the EPEL repository and install the toolset:
yum install -y epel-release.noarch && yum -y update
yum install -y stress-ng
Simulating Process Contention
The following command launches ten times the number of available CPUs as processes, forcing aggressive context switching on systems with fewer cores.
NUM_CPUS=$(( $(nproc) * 10 ))
stress-ng --cpu $NUM_CPUS --pthread 1 --timeout 150
Explanation of the flags:
$(nproc)retrieves the physical core count.- The arithmetic expression multiplies that value by 10 to determine the14 process count.
--pthread 1ensures each stressor is single-threaded.--timeout 150stops the test automatically after 150 seconds.
Observing the Impact: top
While the stressor runs, the load average climbs rapidly and can far exceed the core count. CPU usage (us + sy) often reaches 100%, with user time dominating. The top output also shows a high number of running tasks. Memory metrics barely shift because the workload is CPU-bound, and32 normal levels are restored once the stressor terminates.
Sampling System Statistics: vmstat
Running vmstat 1 during the test highlights several abnormal patterns:
procs rcolumn: very large values, indicating many runnable or waiting processes.memory freedecreases moderately asstress-ngallocates its working set.system in(interrupts per second) increases, but the most dramatic change is insystem cs(context switches per second), which spikes far above idle levels.- The swap and I/O fields remain 5 low because the test does not trigger disk activity.
Pinpointing the Offending Process: pidstat
When 100% CPU utilisatoin and high load are980 observed,16 vmstat confirms a large run queue and high context-switch rate. However,14 the source process is still unknown. Identify it with:
pidstat -w 3
This shows per-process voluntary and involuntary context switches. During the test, many process exhibit elevated cswch/s and nvcswch/s. Focus on the process whose non-voluntary switch rate (nvcswch/s) climbs steadily; it is16 likely consuming a disproportionate amount of CPU and triggering scheduler pressure.
Once the PID is known, further analysis with pidstat -u or perf top can reveal whether the code is doing legitimate work or spinning uselessly. The stess-ng generators always appear as the culprit because they intentionally burn CPU, but the methodology applies equal to any production service.
Workflow Summary
topreveals sustained high load and saturated CPUs.vmstat 1shows a long run queue (r) and elevated context switches (cs), pointing to process contention.pidstat -w 3identifies specific PIDs responsible for the excessive switching, enabling targeted investigation.- Memory, swap, and I/O metrics remain77 stable, ruling out memory pressure or disk bottlenecks.
1 This systematic approach isolates the resource-hungry process quickly, even in16 a complex, multi-tenant environment.