Skip to content

Python System Automation Deep Dive

Topics: subprocess, os, pathlib, shutil, psutil, socket, platform, signal, pwd/grp, crontab, systemd, watchdog, tempfile, glob, fnmatch, stat, struct, fcntl Strategy: Bash equivalent → Python replacement, building real tools Level: L1–L2 (Foundations → Operations) Time: 120–150 minutes Prerequisites: Basic Python (variables, functions, loops, dicts). If you came from Python for Ops — The Bash Expert's Bridge, you're set.


The Mission

You maintain 40 Linux servers. Some bare-metal, some VMs, a few LXC containers. You have a ~/bin/ full of bash scripts: disk cleanup, log rotation, user provisioning, service restarts, port checks, cron management. They work. They've worked for years.

But they're brittle. One script parses /etc/passwd with awk -F:, another does ps aux | grep | grep -v grep, a third does df -h | tail -n +2 | awk '{print $5}' and breaks when a filesystem name wraps to the next line. You've been burned by spaces in filenames, locale-dependent output, and the classic "it worked on my machine" because one server has GNU coreutils and another has busybox.

Today you rewrite all of it in Python. Not because Python is always better — but because for system automation that needs to be reliable, testable, and maintainable, Python's standard library gives you structured access to everything bash was scraping with text parsing.


Part 1 — The Standard Library Tour: Your New Toolbox

Before we write anything, meet the modules. These are all standard library — no pip install required.

The Big Table: Bash → Python

Bash Python module What it does
ls, find pathlib, glob, os.walk File discovery
cp, mv, rm, mkdir -p shutil, pathlib File operations
cat, head, tail open(), itertools.islice File reading
chmod, chown, stat os.chmod, os.chown, os.stat, stat module Permissions
df, du shutil.disk_usage, os.statvfs Disk usage
ps, kill, top psutil (or /proc parsing) Process management
free, uptime, uname psutil, platform, os.uname System info
hostname, ip addr socket, netifaces Networking
useradd, usermod pwd, grp, subprocess User management
systemctl subprocess, dbus Service management
crontab -l, crontab -e python-crontab, sched Scheduling
inotifywait watchdog, inotify File watching
mktemp tempfile Temp files
trap signal, atexit Signal handling
flock fcntl.flock, filelock File locking
date, sleep datetime, time Time operations
sed, awk, grep re, str methods Text processing
tar, gzip, zip tarfile, gzip, zipfile Archives
env, export os.environ Environment variables

Rule of thumb: If you're shelling out to a coreutils command just to parse its text output, there's almost certainly a Python module that gives you the data as a native object.


Part 2 — File and Directory Operations

pathlib: The One Module to Rule Them All

pathlib replaced os.path as the Pythonic way to handle filesystem paths. If you learn one module from this lesson, make it this one.

from pathlib import Path

# --- Creating paths ---
home = Path.home()                    # /home/yourusername
config = home / ".config" / "myapp"   # operator / joins paths
log_dir = Path("/var/log/myapp")

# --- Checking existence ---
if config.exists():
    print(f"Config dir exists: {config}")
if not log_dir.is_dir():
    log_dir.mkdir(parents=True, exist_ok=True)  # mkdir -p

# --- Listing files ---
# Like: ls /var/log/*.log
for f in Path("/var/log").glob("*.log"):
    print(f.name, f.stat().st_size)

# Like: find /var/log -name "*.log" -type f (recursive)
for f in Path("/var/log").rglob("*.log"):
    print(f)

# --- Reading and writing ---
# Like: cat /etc/hostname
hostname = Path("/etc/hostname").read_text().strip()

# Like: echo "data" > /tmp/output.txt
Path("/tmp/output.txt").write_text("data\n")

# Like: cat >> /tmp/output.txt (append)
with open("/tmp/output.txt", "a") as fh:
    fh.write("more data\n")

# --- File metadata ---
p = Path("/etc/passwd")
print(f"Size: {p.stat().st_size}")
print(f"Modified: {p.stat().st_mtime}")
print(f"Owner UID: {p.stat().st_uid}")
print(f"Permissions: {oct(p.stat().st_mode)}")

# --- Path manipulation ---
p = Path("/var/log/syslog.1.gz")
print(p.name)        # syslog.1.gz
print(p.stem)        # syslog.1
print(p.suffix)      # .gz
print(p.suffixes)    # ['.1', '.gz']
print(p.parent)      # /var/log
print(p.parts)       # ('/', 'var', 'log', 'syslog.1.gz')
print(p.resolve())   # resolves symlinks to absolute path

Why not os.path? You can still use it, but compare:

# os.path (old school)
import os
full = os.path.join(os.path.expanduser("~"), ".config", "myapp", "settings.json")
if os.path.isfile(full):
    with open(full) as f:
        data = f.read()

# pathlib (modern)
from pathlib import Path
p = Path.home() / ".config" / "myapp" / "settings.json"
if p.is_file():
    data = p.read_text()

shutil: Bulk File Operations

import shutil

# cp -r /src /dst
shutil.copytree("/opt/myapp", "/opt/myapp.bak")

# cp file1 file2
shutil.copy2("/etc/nginx/nginx.conf", "/tmp/nginx.conf.bak")  # preserves metadata

# mv /src /dst
shutil.move("/tmp/report.csv", "/archive/report-2026.csv")

# rm -rf (CAREFUL)
shutil.rmtree("/tmp/build-artifacts")

# df -h (disk usage)
usage = shutil.disk_usage("/")
print(f"Total: {usage.total // (1024**3)} GB")
print(f"Used:  {usage.used // (1024**3)} GB")
print(f"Free:  {usage.free // (1024**3)} GB")
pct = usage.used / usage.total * 100
print(f"Usage: {pct:.1f}%")

tempfile: Safe Temporary Files

import tempfile
from pathlib import Path

# Like: mktemp
with tempfile.NamedTemporaryFile(mode="w", suffix=".conf", delete=False) as tmp:
    tmp.write("worker_processes auto;\n")
    print(f"Wrote to {tmp.name}")
    # File persists after the block because delete=False

# Like: mktemp -d
with tempfile.TemporaryDirectory(prefix="build-") as tmpdir:
    build_dir = Path(tmpdir)
    (build_dir / "output.txt").write_text("build artifact")
    # Directory and contents auto-deleted when block exits

File Locking: No More Race Conditions

In bash you might use flock. Python equivalent:

import fcntl
from pathlib import Path

lockfile = Path("/var/run/myapp.lock")

def acquire_lock():
    """Prevent multiple instances from running."""
    fh = open(lockfile, "w")
    try:
        fcntl.flock(fh, fcntl.LOCK_EX | fcntl.LOCK_NB)
        fh.write(str(os.getpid()))
        fh.flush()
        return fh
    except BlockingIOError:
        print("Another instance is already running")
        raise SystemExit(1)

# Or use the filelock library for cross-platform support:
# pip install filelock
from filelock import FileLock
lock = FileLock("/var/run/myapp.lock", timeout=10)
with lock:
    # critical section — only one process at a time
    do_work()

Part 3 — Process Management

subprocess: Running Commands Properly

You already know subprocess.run() from the bridge lesson, but let's go deeper.

import subprocess

# --- Basic execution ---
# Like: ls -la /var/log
result = subprocess.run(
    ["ls", "-la", "/var/log"],
    capture_output=True,
    text=True,
    check=True,  # raises CalledProcessError on non-zero exit
)
print(result.stdout)

# --- NEVER do this ---
# subprocess.run(f"ls -la {user_input}", shell=True)  # COMMAND INJECTION
# Always use a list of args, never shell=True with untrusted input

# --- Piping (the safe way) ---
# Like: cat /var/log/syslog | grep ERROR | wc -l
# Don't chain shell pipes. Do the filtering in Python:
result = subprocess.run(
    ["cat", "/var/log/syslog"],
    capture_output=True, text=True,
)
error_count = sum(1 for line in result.stdout.splitlines() if "ERROR" in line)

# --- Timeout ---
try:
    result = subprocess.run(
        ["ping", "-c", "4", "8.8.8.8"],
        capture_output=True, text=True,
        timeout=10,
    )
except subprocess.TimeoutExpired:
    print("Command timed out")

# --- Environment variables ---
import os
env = os.environ.copy()
env["MY_VAR"] = "custom_value"
result = subprocess.run(["printenv", "MY_VAR"], capture_output=True, text=True, env=env)

# --- Working directory ---
result = subprocess.run(
    ["git", "status", "--short"],
    capture_output=True, text=True,
    cwd="/opt/myapp",
)

# --- Streaming output (for long-running commands) ---
proc = subprocess.Popen(
    ["tail", "-f", "/var/log/syslog"],
    stdout=subprocess.PIPE,
    text=True,
)
for line in proc.stdout:
    if "error" in line.lower():
        print(f"ALERT: {line.strip()}")
    # Break after 100 lines for demo purposes

psutil: Process and System Info Without Parsing Text

This is where Python absolutely destroys bash. Instead of ps aux | grep | awk, you get structured data.

pip install psutil
import psutil

# --- System info ---
# Like: uptime
boot = psutil.boot_time()
print(f"Up since: {datetime.fromtimestamp(boot)}")

# Like: free -h
mem = psutil.virtual_memory()
print(f"Total: {mem.total // (1024**3)} GB")
print(f"Available: {mem.available // (1024**3)} GB")
print(f"Used: {mem.percent}%")

swap = psutil.swap_memory()
print(f"Swap used: {swap.percent}%")

# Like: nproc / lscpu
print(f"CPUs (logical): {psutil.cpu_count()}")
print(f"CPUs (physical): {psutil.cpu_count(logical=False)}")
print(f"CPU usage: {psutil.cpu_percent(interval=1)}%")
# Per-CPU usage:
for i, pct in enumerate(psutil.cpu_percent(interval=1, percpu=True)):
    print(f"  CPU {i}: {pct}%")

# Like: df -h
for part in psutil.disk_partitions():
    try:
        usage = psutil.disk_usage(part.mountpoint)
        print(f"{part.device} → {part.mountpoint}: "
              f"{usage.percent}% used ({usage.free // (1024**3)} GB free)")
    except PermissionError:
        pass

# --- Process management ---
# Like: ps aux | grep nginx
for proc in psutil.process_iter(["pid", "name", "username", "cpu_percent", "memory_percent"]):
    if "nginx" in proc.info["name"]:
        print(f"PID {proc.info['pid']}: {proc.info['name']} "
              f"(user={proc.info['username']}, "
              f"cpu={proc.info['cpu_percent']}%, "
              f"mem={proc.info['memory_percent']:.1f}%)")

# Like: kill -9 <pid>
proc = psutil.Process(12345)
proc.terminate()  # SIGTERM (graceful)
proc.wait(timeout=5)
# proc.kill()     # SIGKILL (force)

# Like: pgrep -f myapp
def find_procs(name):
    """Find processes by name (no grep | grep -v grep nonsense)."""
    return [p for p in psutil.process_iter(["name", "cmdline"])
            if name in (p.info["name"] or "")]

# --- Network connections ---
# Like: ss -tlnp / netstat -tlnp
for conn in psutil.net_connections(kind="tcp"):
    if conn.status == "LISTEN":
        print(f"PID {conn.pid} listening on {conn.laddr.ip}:{conn.laddr.port}")

# Like: top (one-shot)
def top_procs(n=10):
    """Top N processes by memory usage."""
    procs = []
    for p in psutil.process_iter(["pid", "name", "memory_percent", "cpu_percent"]):
        procs.append(p.info)
    procs.sort(key=lambda x: x["memory_percent"] or 0, reverse=True)
    for p in procs[:n]:
        print(f"PID {p['pid']:>6} | {p['name']:<20} | "
              f"MEM {p['memory_percent']:>5.1f}% | CPU {p['cpu_percent']:>5.1f}%")

Signal Handling: Graceful Shutdowns

import signal
import sys

# Like: trap 'cleanup' SIGTERM SIGINT
def handle_shutdown(signum, frame):
    sig_name = signal.Signals(signum).name
    print(f"\nReceived {sig_name}, cleaning up...")
    # Close connections, flush buffers, remove pid files, etc.
    cleanup()
    sys.exit(0)

signal.signal(signal.SIGTERM, handle_shutdown)
signal.signal(signal.SIGINT, handle_shutdown)

# Like: trap 'cleanup' EXIT
import atexit
def cleanup():
    Path("/var/run/myapp.pid").unlink(missing_ok=True)
    print("Cleaned up")

atexit.register(cleanup)

Part 4 — System Information

platform: What Am I Running On?

import platform

print(platform.system())         # Linux
print(platform.release())        # 5.15.0-91-generic
print(platform.version())        # #101-Ubuntu SMP...
print(platform.machine())        # x86_64
print(platform.node())           # hostname
print(platform.python_version()) # 3.11.6

# Like: uname -a (all at once)
print(platform.uname())

# Like: lsb_release -a (on Linux)
try:
    import distro  # pip install distro
    print(f"{distro.name()} {distro.version()} ({distro.codename()})")
except ImportError:
    # fallback
    print(platform.freedesktop_os_release().get("PRETTY_NAME", "Unknown"))

Reading /proc Directly

Sometimes the standard library isn't enough and you need to go straight to the source. Every Linux sysadmin should know that /proc is a goldmine.

from pathlib import Path

# Like: cat /proc/loadavg
load = Path("/proc/loadavg").read_text().split()
print(f"Load: {load[0]} {load[1]} {load[2]}")
print(f"Running/Total threads: {load[3]}")

# Like: cat /proc/meminfo | grep MemAvailable
meminfo = {}
for line in Path("/proc/meminfo").read_text().splitlines():
    key, _, value = line.partition(":")
    meminfo[key.strip()] = value.strip()
print(f"Available: {meminfo['MemAvailable']}")

# Like: cat /proc/<pid>/status
def proc_info(pid):
    """Read structured process info from /proc."""
    status = Path(f"/proc/{pid}/status").read_text()
    info = {}
    for line in status.splitlines():
        key, _, value = line.partition(":")
        info[key.strip()] = value.strip()
    return info

# Like: ls /proc/*/fd | wc -l (open file descriptors)
def open_fds(pid):
    """Count open file descriptors for a process."""
    fd_dir = Path(f"/proc/{pid}/fd")
    try:
        return len(list(fd_dir.iterdir()))
    except PermissionError:
        return -1

Part 5 — User and Group Management

pwd and grp: Reading User/Group Databases

import pwd
import grp

# Like: getent passwd username
user = pwd.getpwnam("www-data")
print(f"UID: {user.pw_uid}")
print(f"GID: {user.pw_gid}")
print(f"Home: {user.pw_dir}")
print(f"Shell: {user.pw_shell}")

# Like: getent passwd (all users)
for u in pwd.getpwall():
    if u.pw_uid >= 1000 and u.pw_uid < 65534:  # human users
        print(f"{u.pw_name} (UID {u.pw_uid}): {u.pw_dir}")

# Like: getent group
for g in grp.getgrall():
    if g.gr_mem:  # groups with members
        print(f"{g.gr_name} (GID {g.gr_gid}): {', '.join(g.gr_mem)}")

# Like: id username
def user_info(username):
    """Get full user info like the `id` command."""
    u = pwd.getpwnam(username)
    groups = [g.gr_name for g in grp.getgrall() if username in g.gr_mem]
    primary = grp.getgrgid(u.pw_gid).gr_name
    return {
        "uid": u.pw_uid,
        "gid": u.pw_gid,
        "primary_group": primary,
        "groups": groups,
        "home": u.pw_dir,
        "shell": u.pw_shell,
    }

# Like: who / w (logged-in users)
import utmp  # pip install utmp — or parse /var/run/utmp manually
# Or the quick way:
import subprocess
result = subprocess.run(["who"], capture_output=True, text=True)
print(result.stdout)

User Provisioning Script

In bash you'd write useradd -m -s /bin/bash -G sudo,docker $user. Here's the Python equivalent that's actually maintainable:

import subprocess
import pwd
import secrets
import string
from pathlib import Path

def create_user(username, groups=None, shell="/bin/bash", ssh_key=None):
    """Create a user account with optional group membership and SSH key."""
    # Check if user exists
    try:
        pwd.getpwnam(username)
        print(f"User {username} already exists")
        return False
    except KeyError:
        pass

    # Build useradd command
    cmd = ["useradd", "-m", "-s", shell]
    if groups:
        cmd.extend(["-G", ",".join(groups)])
    cmd.append(username)

    subprocess.run(cmd, check=True)

    # Generate temporary password
    alphabet = string.ascii_letters + string.digits + string.punctuation
    temp_pass = "".join(secrets.choice(alphabet) for _ in range(16))
    subprocess.run(
        ["chpasswd"],
        input=f"{username}:{temp_pass}",
        text=True,
        check=True,
    )

    # Force password change on first login
    subprocess.run(["passwd", "-e", username], check=True)

    # Set up SSH key if provided
    if ssh_key:
        user = pwd.getpwnam(username)
        ssh_dir = Path(user.pw_dir) / ".ssh"
        ssh_dir.mkdir(mode=0o700, exist_ok=True)
        auth_keys = ssh_dir / "authorized_keys"
        auth_keys.write_text(ssh_key + "\n")
        auth_keys.chmod(0o600)
        # chown to the new user
        import os
        os.chown(ssh_dir, user.pw_uid, user.pw_gid)
        os.chown(auth_keys, user.pw_uid, user.pw_gid)

    print(f"Created user {username} (temp password: {temp_pass})")
    return True

Part 6 — Networking

socket: DNS, Ports, and Connections

import socket

# Like: hostname
print(socket.gethostname())

# Like: hostname -f
print(socket.getfqdn())

# Like: dig / nslookup
ip = socket.gethostbyname("google.com")
print(f"google.com → {ip}")

# Reverse DNS: like dig -x
try:
    hostname = socket.gethostbyaddr("8.8.8.8")
    print(f"8.8.8.8 → {hostname[0]}")
except socket.herror:
    print("No reverse DNS")

# Like: getent hosts
info = socket.getaddrinfo("google.com", 443, proto=socket.IPPROTO_TCP)
for family, socktype, proto, canonname, sockaddr in info:
    print(f"  {sockaddr[0]}:{sockaddr[1]}")

# --- Port checking ---
# Like: nc -zv host port / nmap -p port host
def check_port(host, port, timeout=3):
    """Check if a TCP port is open."""
    try:
        with socket.create_connection((host, port), timeout=timeout):
            return True
    except (ConnectionRefusedError, TimeoutError, OSError):
        return False

# Scan common ports
services = {22: "SSH", 80: "HTTP", 443: "HTTPS", 3306: "MySQL", 5432: "PostgreSQL",
            6379: "Redis", 8080: "Alt-HTTP", 9090: "Prometheus"}
for port, name in services.items():
    status = "OPEN" if check_port("localhost", port) else "closed"
    print(f"  {port:>5} ({name:<12}): {status}")

Getting Local IP Addresses

import socket
import fcntl
import struct

# Quick way: what IP would we use to reach the internet?
def get_primary_ip():
    """Get the primary outbound IP address."""
    with socket.socket(socket.AF_INET, socket.SOCK_DGRAM) as s:
        s.connect(("8.8.8.8", 80))  # doesn't actually send anything
        return s.getsockname()[0]

print(f"Primary IP: {get_primary_ip()}")

# All interfaces (using psutil — much cleaner than parsing ip addr)
import psutil

for iface, addrs in psutil.net_if_addrs().items():
    for addr in addrs:
        if addr.family == socket.AF_INET:
            print(f"{iface}: {addr.address}/{addr.netmask}")

Part 7 — Service Management

systemd via subprocess

There's no pure-Python systemd library in the standard library, but subprocess gives you clean access.

import subprocess
import json

def systemctl(action, service):
    """Run a systemctl command and return success/failure."""
    result = subprocess.run(
        ["systemctl", action, service],
        capture_output=True, text=True,
    )
    return result.returncode == 0

def service_status(service):
    """Get structured service status."""
    result = subprocess.run(
        ["systemctl", "show", service,
         "--property=ActiveState,SubState,MainPID,ExecMainStartTimestamp,"
         "MemoryCurrent,CPUUsageNSec,NRestarts"],
        capture_output=True, text=True,
    )
    if result.returncode != 0:
        return None

    info = {}
    for line in result.stdout.strip().splitlines():
        key, _, value = line.partition("=")
        info[key] = value
    return info

# --- Usage ---
# Like: systemctl status nginx
status = service_status("nginx")
if status:
    print(f"State: {status['ActiveState']} ({status['SubState']})")
    print(f"PID: {status['MainPID']}")
    print(f"Restarts: {status['NRestarts']}")

# Like: systemctl restart nginx
if not systemctl("restart", "nginx"):
    print("Failed to restart nginx!")

# Like: systemctl is-active nginx
if systemctl("is-active", "nginx"):
    print("nginx is running")

# --- List failed units ---
# Like: systemctl --failed
result = subprocess.run(
    ["systemctl", "list-units", "--failed", "--no-legend", "--plain"],
    capture_output=True, text=True,
)
failed = [line.split()[0] for line in result.stdout.strip().splitlines() if line.strip()]
if failed:
    print(f"FAILED UNITS: {', '.join(failed)}")

Journal Querying

import subprocess
import json
from datetime import datetime, timedelta

def journal_errors(service, since_minutes=60):
    """Get recent error-level journal entries for a service."""
    since = (datetime.now() - timedelta(minutes=since_minutes)).strftime("%Y-%m-%d %H:%M:%S")
    result = subprocess.run(
        ["journalctl", "-u", service, "--since", since,
         "-p", "err", "-o", "json", "--no-pager"],
        capture_output=True, text=True,
    )
    entries = []
    for line in result.stdout.strip().splitlines():
        try:
            entry = json.loads(line)
            entries.append({
                "timestamp": entry.get("__REALTIME_TIMESTAMP"),
                "message": entry.get("MESSAGE"),
                "priority": entry.get("PRIORITY"),
            })
        except json.JSONDecodeError:
            pass
    return entries

errors = journal_errors("nginx", since_minutes=30)
for e in errors:
    print(f"  {e['message']}")

Part 8 — File Watching

watchdog: React to Filesystem Changes

In bash you'd use inotifywait -m -r /path. Python's watchdog library is the standard answer.

pip install watchdog
import time
from watchdog.observers import Observer
from watchdog.events import FileSystemEventHandler

class LogHandler(FileSystemEventHandler):
    """React to file changes in a directory."""

    def on_modified(self, event):
        if not event.is_directory and event.src_path.endswith(".log"):
            print(f"Modified: {event.src_path}")

    def on_created(self, event):
        if not event.is_directory:
            print(f"New file: {event.src_path}")

    def on_deleted(self, event):
        if not event.is_directory:
            print(f"Deleted: {event.src_path}")

# Watch /var/log for changes
observer = Observer()
observer.schedule(LogHandler(), "/var/log", recursive=True)
observer.start()

try:
    while True:
        time.sleep(1)
except KeyboardInterrupt:
    observer.stop()
observer.join()

Part 9 — Scheduling and Cron

python-crontab: Manage Cron Jobs Programmatically

pip install python-crontab
from crontab import CronTab

# Like: crontab -l
cron = CronTab(user="root")
for job in cron:
    print(f"  {job}")

# Like: (crontab -l; echo "0 2 * * * /usr/local/bin/backup.sh") | crontab -
job = cron.new(command="/usr/local/bin/backup.sh", comment="nightly backup")
job.setall("0 2 * * *")  # 2:00 AM daily
cron.write()

# Like: crontab -l | grep -v backup | crontab -
cron.remove_all(comment="nightly backup")
cron.write()

# Validate a schedule
job = cron.new(command="/bin/true")
job.setall("*/5 * * * *")  # every 5 minutes
print(f"Valid: {job.is_valid()}")
print(f"Next run: {job.schedule().get_next()}")

sched: In-Process Scheduling

For scripts that need to run tasks on a schedule without cron:

import sched
import time

scheduler = sched.scheduler(time.time, time.sleep)

def check_disk():
    """Check disk usage and alert if > 90%."""
    import shutil
    usage = shutil.disk_usage("/")
    pct = usage.used / usage.total * 100
    if pct > 90:
        print(f"ALERT: Disk usage at {pct:.1f}%")
    # Re-schedule self (every 300 seconds)
    scheduler.enter(300, 1, check_disk)

# Start the loop
scheduler.enter(0, 1, check_disk)
scheduler.run()

Part 10 — Archives and Compression

import tarfile
import zipfile
import gzip
import shutil
from pathlib import Path

# --- tar.gz ---
# Like: tar czf backup.tar.gz /opt/myapp/
with tarfile.open("/tmp/backup.tar.gz", "w:gz") as tar:
    tar.add("/opt/myapp", arcname="myapp")

# Like: tar xzf backup.tar.gz -C /tmp/restore/
with tarfile.open("/tmp/backup.tar.gz", "r:gz") as tar:
    tar.extractall("/tmp/restore", filter="data")  # filter= for safety (Python 3.12+)

# Like: tar tzf backup.tar.gz
with tarfile.open("/tmp/backup.tar.gz", "r:gz") as tar:
    for member in tar.getmembers():
        print(f"  {member.name} ({member.size} bytes)")

# --- zip ---
# Like: zip -r backup.zip /opt/myapp/
with zipfile.ZipFile("/tmp/backup.zip", "w", zipfile.ZIP_DEFLATED) as zf:
    for f in Path("/opt/myapp").rglob("*"):
        if f.is_file():
            zf.write(f, f.relative_to("/opt"))

# --- gzip a single file ---
# Like: gzip -k access.log
with open("/var/log/access.log", "rb") as f_in:
    with gzip.open("/var/log/access.log.gz", "wb") as f_out:
        shutil.copyfileobj(f_in, f_out)

Part 11 — Environment and Configuration

os.environ: Environment Variables Done Right

import os

# Like: echo $HOME
home = os.environ["HOME"]

# Like: echo ${DB_HOST:-localhost} (with default)
db_host = os.environ.get("DB_HOST", "localhost")

# Like: export MY_VAR=value
os.environ["MY_VAR"] = "value"

# Like: env | grep DB_
db_vars = {k: v for k, v in os.environ.items() if k.startswith("DB_")}

# Like: source .env (parse a .env file)
def load_dotenv(path=".env"):
    """Minimal .env loader (for when you don't want python-dotenv)."""
    env_file = Path(path)
    if not env_file.exists():
        return
    for line in env_file.read_text().splitlines():
        line = line.strip()
        if not line or line.startswith("#"):
            continue
        key, _, value = line.partition("=")
        # Strip optional quotes
        value = value.strip().strip("'\"")
        os.environ[key.strip()] = value

configparser: INI Files

import configparser

# Like: parsing my.cnf, php.ini, systemd unit files
config = configparser.ConfigParser()
config.read("/etc/myapp/config.ini")

db_host = config.get("database", "host", fallback="localhost")
db_port = config.getint("database", "port", fallback=5432)
debug = config.getboolean("app", "debug", fallback=False)

Part 12 — Putting It All Together: A System Health Check Script

Here's the kind of script that replaces a dozen bash one-liners:

#!/usr/bin/env python3
"""System health check — replaces 200 lines of bash."""

import shutil
import socket
from datetime import datetime
from pathlib import Path

import psutil

def check_disk(threshold=85):
    """Check all mounted filesystems."""
    alerts = []
    for part in psutil.disk_partitions():
        try:
            usage = psutil.disk_usage(part.mountpoint)
            if usage.percent >= threshold:
                alerts.append(
                    f"DISK {part.mountpoint}: {usage.percent}% "
                    f"({usage.free // (1024**3)} GB free)"
                )
        except PermissionError:
            pass
    return alerts

def check_memory(threshold=90):
    """Check RAM and swap usage."""
    alerts = []
    mem = psutil.virtual_memory()
    if mem.percent >= threshold:
        alerts.append(f"RAM: {mem.percent}% used ({mem.available // (1024**2)} MB available)")
    swap = psutil.swap_memory()
    if swap.percent >= threshold:
        alerts.append(f"SWAP: {swap.percent}% used")
    return alerts

def check_load():
    """Check system load average."""
    load1, load5, load15 = psutil.getloadavg()
    cpus = psutil.cpu_count()
    alerts = []
    if load5 > cpus * 0.8:
        alerts.append(f"LOAD: {load5:.2f} (5min) on {cpus} CPUs")
    return alerts

def check_services(services):
    """Check if critical services are running."""
    import subprocess
    alerts = []
    for svc in services:
        result = subprocess.run(
            ["systemctl", "is-active", svc],
            capture_output=True, text=True,
        )
        if result.stdout.strip() != "active":
            alerts.append(f"SERVICE {svc}: {result.stdout.strip()}")
    return alerts

def check_ports(ports):
    """Check if expected ports are listening."""
    alerts = []
    for port, name in ports.items():
        try:
            with socket.create_connection(("localhost", port), timeout=2):
                pass
        except (ConnectionRefusedError, TimeoutError, OSError):
            alerts.append(f"PORT {port} ({name}): not listening")
    return alerts

def check_zombies():
    """Find zombie processes."""
    zombies = []
    for proc in psutil.process_iter(["pid", "name", "status"]):
        if proc.info["status"] == psutil.STATUS_ZOMBIE:
            zombies.append(f"ZOMBIE: PID {proc.info['pid']} ({proc.info['name']})")
    return zombies

def main():
    hostname = socket.gethostname()
    now = datetime.now().strftime("%Y-%m-%d %H:%M:%S")
    print(f"=== Health Check: {hostname} @ {now} ===\n")

    all_alerts = []

    checks = [
        ("Disk", check_disk),
        ("Memory", check_memory),
        ("Load", check_load),
        ("Zombies", check_zombies),
        ("Services", lambda: check_services(["nginx", "postgresql", "redis"])),
        ("Ports", lambda: check_ports({
            80: "HTTP", 443: "HTTPS", 5432: "PostgreSQL", 6379: "Redis",
        })),
    ]

    for name, check_fn in checks:
        alerts = check_fn()
        if alerts:
            all_alerts.extend(alerts)
            for a in alerts:
                print(f"  WARN  {a}")
        else:
            print(f"  OK    {name}")

    print()
    if all_alerts:
        print(f"RESULT: {len(all_alerts)} warning(s)")
        return 1
    else:
        print("RESULT: All checks passed")
        return 0

if __name__ == "__main__":
    raise SystemExit(main())

Part 13 — Common Patterns and Idioms

Pattern: Retry with Backoff

import time
import random

def retry(fn, max_attempts=3, base_delay=1, max_delay=30):
    """Retry a function with exponential backoff and jitter."""
    for attempt in range(1, max_attempts + 1):
        try:
            return fn()
        except Exception as e:
            if attempt == max_attempts:
                raise
            delay = min(base_delay * (2 ** (attempt - 1)), max_delay)
            delay *= (0.5 + random.random())  # jitter
            print(f"Attempt {attempt} failed ({e}), retrying in {delay:.1f}s...")
            time.sleep(delay)

Pattern: PID File

import os
import sys
from pathlib import Path

def write_pidfile(path="/var/run/myapp.pid"):
    """Write PID file, exit if already running."""
    pidfile = Path(path)
    if pidfile.exists():
        old_pid = int(pidfile.read_text().strip())
        try:
            os.kill(old_pid, 0)  # signal 0 = check if alive
            print(f"Already running (PID {old_pid})")
            sys.exit(1)
        except ProcessLookupError:
            pass  # stale pidfile, proceed
    pidfile.write_text(str(os.getpid()))

def remove_pidfile(path="/var/run/myapp.pid"):
    Path(path).unlink(missing_ok=True)

Pattern: Atomic File Write

import tempfile
import os
from pathlib import Path

def atomic_write(path, content, mode=0o644):
    """Write a file atomically (write to temp, then rename).
    Prevents partial writes if the process crashes mid-write.
    """
    target = Path(path)
    fd, tmp_path = tempfile.mkstemp(dir=target.parent, prefix=f".{target.name}.")
    try:
        with os.fdopen(fd, "w") as f:
            f.write(content)
        os.chmod(tmp_path, mode)
        os.rename(tmp_path, target)  # atomic on same filesystem
    except BaseException:
        os.unlink(tmp_path)
        raise

Pattern: Log Tailer

import time
from pathlib import Path

def tail_follow(path, callback):
    """Like tail -f, but in Python. Handles log rotation."""
    p = Path(path)
    pos = p.stat().st_size  # start at end
    inode = p.stat().st_ino

    while True:
        stat = p.stat()
        # Check for log rotation (inode changed)
        if stat.st_ino != inode:
            pos = 0
            inode = stat.st_ino

        if stat.st_size > pos:
            with open(p) as f:
                f.seek(pos)
                for line in f:
                    callback(line.rstrip())
                pos = f.tell()

        time.sleep(0.5)

# Usage:
# tail_follow("/var/log/syslog", lambda line: print(line) if "ERROR" in line else None)

Module Quick Reference

Standard Library (no install needed)

Module Use for
os Environment, UID/GID, file descriptors, os.walk
os.path Legacy path manipulation (prefer pathlib)
pathlib Modern path objects, file I/O, globbing
shutil Copy/move/delete trees, disk usage, which()
subprocess Running external commands
signal Signal handlers (SIGTERM, SIGINT, etc.)
atexit Cleanup on exit
tempfile Temp files and directories
glob Shell-style wildcards (prefer pathlib.glob)
fnmatch Filename pattern matching
stat File permission constants (S_IRUSR, etc.)
fcntl File locking (flock)
socket DNS, port checks, hostname
platform OS/arch/Python version info
pwd User database (/etc/passwd)
grp Group database (/etc/group)
sched In-process event scheduler
configparser INI file parsing
tarfile tar/tar.gz/tar.bz2 creation and extraction
zipfile ZIP creation and extraction
gzip gzip compression
struct Binary data packing/unpacking
datetime Date/time manipulation
time Sleep, monotonic clock, epoch time
re Regex (replaces grep, sed, awk patterns)
json JSON parsing/generation
csv CSV parsing/generation
logging Structured logging
argparse CLI argument parsing
textwrap Text wrapping/dedenting
itertools Efficient iteration (islice for head/tail)
collections Counter, defaultdict, deque
concurrent.futures Thread/process pools

Third-Party (pip install)

Package Use for Replaces
psutil Process/system monitoring ps, top, free, df, netstat
python-crontab Cron management crontab -l/-e
watchdog Filesystem event monitoring inotifywait
distro Linux distribution info lsb_release
filelock Cross-platform file locking flock
paramiko SSH connections ssh, scp
fabric Remote command execution ssh in loops
pexpect Interactive process control expect
click CLI framework argparse (more ergonomic)
rich Terminal formatting/tables printf, column -t
python-dotenv .env file loading source .env
schedule Human-friendly scheduling cron expressions
sh Shell command wrapper subprocess (terser syntax)
plumbum Shell command piping Shell pipes

When to Stay in Bash

Python isn't always the right tool. Stay in bash when:

  • One-liners: grep ERROR /var/log/syslog | tail -20 — don't write 15 lines of Python for this
  • Glue between commands: A 5-line script that runs terraform plan, checks the exit code, and runs terraform apply — bash is fine
  • Interactive/ad-hoc: You're poking around a server debugging something. Use bash.
  • Boot scripts: /etc/init.d/, early boot — Python might not be available yet
  • Performance-critical text processing: awk processing a 10GB log file will beat Python. Use the right tool.

Switch to Python when: - The script is over ~50 lines of bash - You need error handling beyond set -e - You're parsing structured data (JSON, YAML, CSV) - You need retry logic, backoff, or complex control flow - You need to maintain state between runs - Multiple people will maintain the script - You need tests - You're tired of quoting bugs and word-splitting surprises


Exercises

  1. Disk Alert Script: Write a script that checks all mounted filesystems, alerts if any exceed 85% usage, and writes results to a JSON file. Use only shutil and psutil.

  2. Process Inventory: Write a tool that lists all processes, groups them by user, and shows total CPU/memory per user. Sort by memory descending. (Hint: psutil.process_iter

  3. collections.defaultdict)

  4. Port Scanner: Write a function that takes a hostname and a range of ports, checks which are open using socket, and returns results. Add a --timeout and --threads CLI flag using argparse. Use concurrent.futures.ThreadPoolExecutor for parallelism.

  5. Log Rotator: Write a script that rotates log files: app.log → app.log.1 → app.log.2 → ... → delete app.log.5. Handle the case where the process has the file open (use copytruncate strategy).

  6. Service Monitor: Write a daemon that checks a list of systemd services every 60 seconds, logs state changes, and writes a status file. Use signal for graceful shutdown and atexit for cleanup.

  7. User Audit: Write a script that compares /etc/passwd users against an expected list (from a YAML file), reports additions/removals, and checks that all human users (UID >= 1000) have a valid shell and home directory that exists.


Further Reading

  • psutil docs: Process and system monitoring — the single most useful third-party library for system automation
  • pathlib docs: PEP 428 — your new best friend for file operations
  • subprocess docs: Especially the security considerations section
  • signal docs: Signal handling and the caveats around threads
  • The training/library/lessons/python-for-ops-the-bash-experts-bridge.md lesson in this repo covers subprocess, pathlib, and requests in more depth
  • The training/library/lessons/python-automating-everything-apis-and-infrastructure.md lesson covers the API/cloud side of automation