The Silent Autopsy: Forensics in a Post-Crash Server

📅 Oct 03, 2026 ★★★★★ 📚 Operating Systems, DevOps, Incident Response
#Linux Forensics #OOM Killer #File Descriptors #dmesg #strace

Scenario: The Missing Gigabytes

You place the interviewee in a terminal simulation. A critical Linux production server went completely offline 15 minutes ago. It has since rebooted and the application is running normally again. The CTO demands a post-mortem to understand exactly why the server died. The candidate has root access but no APM (Application Performance Monitoring) dashboards—only the raw Linux command line.

Q:The candidate runs uptime and sees the load average is low. They run cat /var/log/syslog, but the logs from the time of the crash are missing. What specific command must they use to view the kernel’s persistent hardware and memory logs from the PREVIOUS boot cycle to find the cause of death? Reveal â–¾

They must use journalctl -k -b -1.

The -k flag filters for strictly kernel-level messages (dmesg), and the -b -1 flag instructs systemd to load the logs from the previous boot cycle before the crash occurred. This is critical because user-space logging daemons (like rsyslog) often freeze or are killed when a system is heavily thrashing, leaving the kernel ring buffer as the only surviving witness to the crash.

Q:The kernel logs reveal: Out of memory: Killed process 4012 (java). The candidate checks the application configuration and sees the JVM is strictly limited to a 2GB maximum heap size (-Xmx2g). The server has 8GB of physical RAM. The candidate is confused how a 2GB JVM could exhaust an 8GB server. Explain the memory black hole. Reveal â–¾

The candidate is confusing the JVM Heap with the total memory footprint of a process.

While the Java garbage-collected heap is capped at 2GB, the JVM process allocates vast amounts of Native Memory (Off-Heap). This includes Thread Stacks (every thread takes 1MB by default), Metaspace (class definitions), the JVM’s internal C++ structures, and most dangerously, Direct Byte Buffers (used extensively by NIO libraries like Netty or Kafka clients for zero-copy network I/O). A memory leak in native C/C++ libraries called via JNI completely bypasses the JVM’s memory limits and will silently consume all server RAM until the Linux OOM-killer intervenes.

Q:The application is restarted. Five minutes later, an alert fires: Disk Space is at 100%. The candidate finds a massive 50GB error.log file and immediately runs rm error.log. They run df -h, but the disk is STILL at 100% capacity. Why didn’t deleting the 50GB file free up any space? Reveal â–¾

Because deleting a file in Linux only unlinks the filename from the directory tree; it does not delete the underlying data blocks if a process still holds an open File Descriptor.

The Java application is still running and actively holding a file handle to error.log. The kernel will not release the 50GB of disk blocks until that specific file descriptor is closed. The candidate must run lsof | grep deleted to identify the ghost file. To actually fix the disk space without restarting the application, they should have truncated the file in place using > error.log or truncate -s 0 error.log, which clears the blocks while preserving the file descriptor.

Variations & Real-World Impact

  • eBPF Observability: Modern incident response is moving away from reactive post-mortems toward proactive tracing. Companies now expect candidates to understand tools like bcc or bpftrace, which use eBPF to safely hook into kernel events (like memory allocations or TCP retransmits) in real-time without crashing the production server.
  • Zombie Processes: Candidates are often tested on process states. If they see a process in state Z (Zombie), they cannot simply kill -9 it. A zombie is already dead; it is merely an entry in the process table waiting for its parent to read its exit status. You must kill the parent process to reap the zombies.

Further Exploration

Discussion & Comments

SDB Watermark