Skip to content

Diagnose Linux Memory Pressure and OOM Kills

When a Linux host runs low on memory, applications may slow down, allocations can fail, and the kernel may terminate a process through the out-of-memory (OOM) killer. A high memory reading alone does not prove a problem: Linux uses otherwise idle RAM for cache. Correlate symptoms with pressure, swap activity, kernel messages, and the affected process before changing limits or restarting services.

The commands in this guide are read-only. Run them during or soon after an incident when possible; process and pressure snapshots show only the current state, while kernel logs may preserve evidence of an earlier OOM event.


Step 1: Check Memory, Swap, and Pressure

01

Take a Read-Only Memory Snapshot

Triage

Check available memory and swap, then inspect the kernel’s pressure-stall information (PSI). PSI measures time tasks spend stalled for memory; sustained pressure is more useful than a single percentage showing RAM in use. vmstat reports memory and swap activity over one-second intervals after its initial average line.

Terminal window
free -h
vmstat 1 5
cat /proc/pressure/memory
❯ View Expected Console Output
total used free shared buff/cache available
Mem: 15Gi 13Gi 420Mi 1.1Gi 2.2Gi 1.4Gi
Swap: 4.0Gi 2.8Gi 1.2Gi
procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
r b swpd free buff cache si so bi bo in cs us sy id wa st
2 0 2936012 430080 92160 2211840 12 64 120 210 800 1400 18 7 70 5 0
some avg10=4.21 avg60=2.87 avg300=1.30 total=123456789
full avg10=1.14 avg60=0.82 avg300=0.35 total=34567890

Step 2: Find Kernel OOM Evidence

02

Review Kernel Messages Around the Incident

Confirm OOM

Search the current boot’s kernel journal for OOM reports. If you know when the incident occurred, bound the query to a time window. The log may identify the killed process, its PID, memory usage, and the affected memory cgroup. Lack of a match does not rule out an OOM event if logs rotated, the host rebooted, or the kernel ring buffer was lost.

Terminal window
sudo journalctl -k -b --no-pager | grep -Ei 'out of memory|oom-kill|killed process'
# Narrow the search to an incident window (adjust the timestamps).
sudo journalctl -k -b --since '2026-10-06 09:00:00' --until '2026-10-06 09:30:00' --no-pager
❯ View Expected Console Output
kernel: memory: usage 2048MB, limit 2048MB, failcnt 37
kernel: Memory cgroup out of memory: Killed process 2481 (worker) total-vm:3120040kB, anon-rss:1842200kB

Step 3: Identify Current Memory Consumers

03

Compare Process Memory with System Totals

Locate Usage

Sort processes by resident memory to find large current consumers. RSS is a useful lead, but it can count shared pages in more than one process and will not explain all kernel or cgroup memory. Capture the owning service and workload before deciding whether a process is expected to use that amount.

Terminal window
ps -eo pid,ppid,user,comm,rss,%mem --sort=-rss | head -n 15
❯ View Expected Console Output
PID PPID USER COMMAND RSS %MEM
2481 2100 app worker 1842200 11.4
902 1 root java 1324000 8.2

Step 4: Check the Service’s Cgroup Limit

04

Compare Service Usage with Its Limit

Check Scope

A host can have available RAM while a service is OOM-killed because its cgroup has a lower memory limit. For a systemd service, inspect its control group and resource properties. On cgroup v2 systems, read memory.current, memory.max, and memory.events from that group’s directory. Replace the example unit with the service involved; permissions may require sudo.

Terminal window
systemctl show example.service -p ControlGroup -p MemoryCurrent -p MemoryMax
# Use the ControlGroup path printed above beneath /sys/fs/cgroup.
sudo cat /sys/fs/cgroup/system.slice/example.service/memory.current
sudo cat /sys/fs/cgroup/system.slice/example.service/memory.max
sudo cat /sys/fs/cgroup/system.slice/example.service/memory.events
❯ View Expected Console Output
ControlGroup=/system.slice/example.service
MemoryCurrent=2147483648
MemoryMax=2147483648
low 0
high 0
max 42
oom 3
oom_kill 2

Step 5: Preserve Evidence and Choose a Response

Record the incident time, affected service, kernel OOM lines, free and vmstat output, PSI values, and cgroup counters. Compare with the service’s normal workload and recent changes. If the cgroup counters show OOM kills while host memory remained available, investigate whether the configured limit matches the workload. If host-wide pressure is sustained, identify the workload and memory trend before considering capacity, application tuning, or workload scheduling changes.

Avoid killing large processes or raising memory limits as an automatic response: either action can interrupt work or shift pressure to the whole host. Validate any proposed limit change against host capacity, competing services, and the application’s documented requirements, then monitor PSI, swap activity, and OOM counters after the change.

Comments