Python System Automation Deep Dive¶
Topics: subprocess, os, pathlib, shutil, psutil, socket, platform, signal, pwd/grp, crontab, systemd, watchdog, tempfile, glob, fnmatch, stat, struct, fcntl Strategy: Bash equivalent → Python replacement, building real tools Level: L1–L2 (Foundations → Operations) Time: 120–150 minutes Prerequisites: Basic Python (variables, functions, loops, dicts). If you came from Python for Ops — The Bash Expert's Bridge, you're set.
The Mission¶
You maintain 40 Linux servers. Some bare-metal, some VMs, a few LXC containers. You have
a ~/bin/ full of bash scripts: disk cleanup, log rotation, user provisioning, service
restarts, port checks, cron management. They work. They've worked for years.
But they're brittle. One script parses /etc/passwd with awk -F:, another does ps aux |
grep | grep -v grep, a third does df -h | tail -n +2 | awk '{print $5}' and breaks
when a filesystem name wraps to the next line. You've been burned by spaces in filenames,
locale-dependent output, and the classic "it worked on my machine" because one server has
GNU coreutils and another has busybox.
Today you rewrite all of it in Python. Not because Python is always better — but because for system automation that needs to be reliable, testable, and maintainable, Python's standard library gives you structured access to everything bash was scraping with text parsing.
Part 1 — The Standard Library Tour: Your New Toolbox¶
Before we write anything, meet the modules. These are all standard library — no
pip install required.
The Big Table: Bash → Python¶
| Bash | Python module | What it does |
|---|---|---|
ls, find |
pathlib, glob, os.walk |
File discovery |
cp, mv, rm, mkdir -p |
shutil, pathlib |
File operations |
cat, head, tail |
open(), itertools.islice |
File reading |
chmod, chown, stat |
os.chmod, os.chown, os.stat, stat module |
Permissions |
df, du |
shutil.disk_usage, os.statvfs |
Disk usage |
ps, kill, top |
psutil (or /proc parsing) |
Process management |
free, uptime, uname |
psutil, platform, os.uname |
System info |
hostname, ip addr |
socket, netifaces |
Networking |
useradd, usermod |
pwd, grp, subprocess |
User management |
systemctl |
subprocess, dbus |
Service management |
crontab -l, crontab -e |
python-crontab, sched |
Scheduling |
inotifywait |
watchdog, inotify |
File watching |
mktemp |
tempfile |
Temp files |
trap |
signal, atexit |
Signal handling |
flock |
fcntl.flock, filelock |
File locking |
date, sleep |
datetime, time |
Time operations |
sed, awk, grep |
re, str methods |
Text processing |
tar, gzip, zip |
tarfile, gzip, zipfile |
Archives |
env, export |
os.environ |
Environment variables |
Rule of thumb: If you're shelling out to a coreutils command just to parse its text output, there's almost certainly a Python module that gives you the data as a native object.
Part 2 — File and Directory Operations¶
pathlib: The One Module to Rule Them All¶
pathlib replaced os.path as the Pythonic way to handle filesystem paths. If you learn
one module from this lesson, make it this one.
from pathlib import Path
# --- Creating paths ---
home = Path.home() # /home/yourusername
config = home / ".config" / "myapp" # operator / joins paths
log_dir = Path("/var/log/myapp")
# --- Checking existence ---
if config.exists():
print(f"Config dir exists: {config}")
if not log_dir.is_dir():
log_dir.mkdir(parents=True, exist_ok=True) # mkdir -p
# --- Listing files ---
# Like: ls /var/log/*.log
for f in Path("/var/log").glob("*.log"):
print(f.name, f.stat().st_size)
# Like: find /var/log -name "*.log" -type f (recursive)
for f in Path("/var/log").rglob("*.log"):
print(f)
# --- Reading and writing ---
# Like: cat /etc/hostname
hostname = Path("/etc/hostname").read_text().strip()
# Like: echo "data" > /tmp/output.txt
Path("/tmp/output.txt").write_text("data\n")
# Like: cat >> /tmp/output.txt (append)
with open("/tmp/output.txt", "a") as fh:
fh.write("more data\n")
# --- File metadata ---
p = Path("/etc/passwd")
print(f"Size: {p.stat().st_size}")
print(f"Modified: {p.stat().st_mtime}")
print(f"Owner UID: {p.stat().st_uid}")
print(f"Permissions: {oct(p.stat().st_mode)}")
# --- Path manipulation ---
p = Path("/var/log/syslog.1.gz")
print(p.name) # syslog.1.gz
print(p.stem) # syslog.1
print(p.suffix) # .gz
print(p.suffixes) # ['.1', '.gz']
print(p.parent) # /var/log
print(p.parts) # ('/', 'var', 'log', 'syslog.1.gz')
print(p.resolve()) # resolves symlinks to absolute path
Why not os.path? You can still use it, but compare:
# os.path (old school)
import os
full = os.path.join(os.path.expanduser("~"), ".config", "myapp", "settings.json")
if os.path.isfile(full):
with open(full) as f:
data = f.read()
# pathlib (modern)
from pathlib import Path
p = Path.home() / ".config" / "myapp" / "settings.json"
if p.is_file():
data = p.read_text()
shutil: Bulk File Operations¶
import shutil
# cp -r /src /dst
shutil.copytree("/opt/myapp", "/opt/myapp.bak")
# cp file1 file2
shutil.copy2("/etc/nginx/nginx.conf", "/tmp/nginx.conf.bak") # preserves metadata
# mv /src /dst
shutil.move("/tmp/report.csv", "/archive/report-2026.csv")
# rm -rf (CAREFUL)
shutil.rmtree("/tmp/build-artifacts")
# df -h (disk usage)
usage = shutil.disk_usage("/")
print(f"Total: {usage.total // (1024**3)} GB")
print(f"Used: {usage.used // (1024**3)} GB")
print(f"Free: {usage.free // (1024**3)} GB")
pct = usage.used / usage.total * 100
print(f"Usage: {pct:.1f}%")
tempfile: Safe Temporary Files¶
import tempfile
from pathlib import Path
# Like: mktemp
with tempfile.NamedTemporaryFile(mode="w", suffix=".conf", delete=False) as tmp:
tmp.write("worker_processes auto;\n")
print(f"Wrote to {tmp.name}")
# File persists after the block because delete=False
# Like: mktemp -d
with tempfile.TemporaryDirectory(prefix="build-") as tmpdir:
build_dir = Path(tmpdir)
(build_dir / "output.txt").write_text("build artifact")
# Directory and contents auto-deleted when block exits
File Locking: No More Race Conditions¶
In bash you might use flock. Python equivalent:
import fcntl
from pathlib import Path
lockfile = Path("/var/run/myapp.lock")
def acquire_lock():
"""Prevent multiple instances from running."""
fh = open(lockfile, "w")
try:
fcntl.flock(fh, fcntl.LOCK_EX | fcntl.LOCK_NB)
fh.write(str(os.getpid()))
fh.flush()
return fh
except BlockingIOError:
print("Another instance is already running")
raise SystemExit(1)
# Or use the filelock library for cross-platform support:
# pip install filelock
from filelock import FileLock
lock = FileLock("/var/run/myapp.lock", timeout=10)
with lock:
# critical section — only one process at a time
do_work()
Part 3 — Process Management¶
subprocess: Running Commands Properly¶
You already know subprocess.run() from the bridge lesson, but let's go deeper.
import subprocess
# --- Basic execution ---
# Like: ls -la /var/log
result = subprocess.run(
["ls", "-la", "/var/log"],
capture_output=True,
text=True,
check=True, # raises CalledProcessError on non-zero exit
)
print(result.stdout)
# --- NEVER do this ---
# subprocess.run(f"ls -la {user_input}", shell=True) # COMMAND INJECTION
# Always use a list of args, never shell=True with untrusted input
# --- Piping (the safe way) ---
# Like: cat /var/log/syslog | grep ERROR | wc -l
# Don't chain shell pipes. Do the filtering in Python:
result = subprocess.run(
["cat", "/var/log/syslog"],
capture_output=True, text=True,
)
error_count = sum(1 for line in result.stdout.splitlines() if "ERROR" in line)
# --- Timeout ---
try:
result = subprocess.run(
["ping", "-c", "4", "8.8.8.8"],
capture_output=True, text=True,
timeout=10,
)
except subprocess.TimeoutExpired:
print("Command timed out")
# --- Environment variables ---
import os
env = os.environ.copy()
env["MY_VAR"] = "custom_value"
result = subprocess.run(["printenv", "MY_VAR"], capture_output=True, text=True, env=env)
# --- Working directory ---
result = subprocess.run(
["git", "status", "--short"],
capture_output=True, text=True,
cwd="/opt/myapp",
)
# --- Streaming output (for long-running commands) ---
proc = subprocess.Popen(
["tail", "-f", "/var/log/syslog"],
stdout=subprocess.PIPE,
text=True,
)
for line in proc.stdout:
if "error" in line.lower():
print(f"ALERT: {line.strip()}")
# Break after 100 lines for demo purposes
psutil: Process and System Info Without Parsing Text¶
This is where Python absolutely destroys bash. Instead of ps aux | grep | awk, you get
structured data.
import psutil
# --- System info ---
# Like: uptime
boot = psutil.boot_time()
print(f"Up since: {datetime.fromtimestamp(boot)}")
# Like: free -h
mem = psutil.virtual_memory()
print(f"Total: {mem.total // (1024**3)} GB")
print(f"Available: {mem.available // (1024**3)} GB")
print(f"Used: {mem.percent}%")
swap = psutil.swap_memory()
print(f"Swap used: {swap.percent}%")
# Like: nproc / lscpu
print(f"CPUs (logical): {psutil.cpu_count()}")
print(f"CPUs (physical): {psutil.cpu_count(logical=False)}")
print(f"CPU usage: {psutil.cpu_percent(interval=1)}%")
# Per-CPU usage:
for i, pct in enumerate(psutil.cpu_percent(interval=1, percpu=True)):
print(f" CPU {i}: {pct}%")
# Like: df -h
for part in psutil.disk_partitions():
try:
usage = psutil.disk_usage(part.mountpoint)
print(f"{part.device} → {part.mountpoint}: "
f"{usage.percent}% used ({usage.free // (1024**3)} GB free)")
except PermissionError:
pass
# --- Process management ---
# Like: ps aux | grep nginx
for proc in psutil.process_iter(["pid", "name", "username", "cpu_percent", "memory_percent"]):
if "nginx" in proc.info["name"]:
print(f"PID {proc.info['pid']}: {proc.info['name']} "
f"(user={proc.info['username']}, "
f"cpu={proc.info['cpu_percent']}%, "
f"mem={proc.info['memory_percent']:.1f}%)")
# Like: kill -9 <pid>
proc = psutil.Process(12345)
proc.terminate() # SIGTERM (graceful)
proc.wait(timeout=5)
# proc.kill() # SIGKILL (force)
# Like: pgrep -f myapp
def find_procs(name):
"""Find processes by name (no grep | grep -v grep nonsense)."""
return [p for p in psutil.process_iter(["name", "cmdline"])
if name in (p.info["name"] or "")]
# --- Network connections ---
# Like: ss -tlnp / netstat -tlnp
for conn in psutil.net_connections(kind="tcp"):
if conn.status == "LISTEN":
print(f"PID {conn.pid} listening on {conn.laddr.ip}:{conn.laddr.port}")
# Like: top (one-shot)
def top_procs(n=10):
"""Top N processes by memory usage."""
procs = []
for p in psutil.process_iter(["pid", "name", "memory_percent", "cpu_percent"]):
procs.append(p.info)
procs.sort(key=lambda x: x["memory_percent"] or 0, reverse=True)
for p in procs[:n]:
print(f"PID {p['pid']:>6} | {p['name']:<20} | "
f"MEM {p['memory_percent']:>5.1f}% | CPU {p['cpu_percent']:>5.1f}%")
Signal Handling: Graceful Shutdowns¶
import signal
import sys
# Like: trap 'cleanup' SIGTERM SIGINT
def handle_shutdown(signum, frame):
sig_name = signal.Signals(signum).name
print(f"\nReceived {sig_name}, cleaning up...")
# Close connections, flush buffers, remove pid files, etc.
cleanup()
sys.exit(0)
signal.signal(signal.SIGTERM, handle_shutdown)
signal.signal(signal.SIGINT, handle_shutdown)
# Like: trap 'cleanup' EXIT
import atexit
def cleanup():
Path("/var/run/myapp.pid").unlink(missing_ok=True)
print("Cleaned up")
atexit.register(cleanup)
Part 4 — System Information¶
platform: What Am I Running On?¶
import platform
print(platform.system()) # Linux
print(platform.release()) # 5.15.0-91-generic
print(platform.version()) # #101-Ubuntu SMP...
print(platform.machine()) # x86_64
print(platform.node()) # hostname
print(platform.python_version()) # 3.11.6
# Like: uname -a (all at once)
print(platform.uname())
# Like: lsb_release -a (on Linux)
try:
import distro # pip install distro
print(f"{distro.name()} {distro.version()} ({distro.codename()})")
except ImportError:
# fallback
print(platform.freedesktop_os_release().get("PRETTY_NAME", "Unknown"))
Reading /proc Directly¶
Sometimes the standard library isn't enough and you need to go straight to the source.
Every Linux sysadmin should know that /proc is a goldmine.
from pathlib import Path
# Like: cat /proc/loadavg
load = Path("/proc/loadavg").read_text().split()
print(f"Load: {load[0]} {load[1]} {load[2]}")
print(f"Running/Total threads: {load[3]}")
# Like: cat /proc/meminfo | grep MemAvailable
meminfo = {}
for line in Path("/proc/meminfo").read_text().splitlines():
key, _, value = line.partition(":")
meminfo[key.strip()] = value.strip()
print(f"Available: {meminfo['MemAvailable']}")
# Like: cat /proc/<pid>/status
def proc_info(pid):
"""Read structured process info from /proc."""
status = Path(f"/proc/{pid}/status").read_text()
info = {}
for line in status.splitlines():
key, _, value = line.partition(":")
info[key.strip()] = value.strip()
return info
# Like: ls /proc/*/fd | wc -l (open file descriptors)
def open_fds(pid):
"""Count open file descriptors for a process."""
fd_dir = Path(f"/proc/{pid}/fd")
try:
return len(list(fd_dir.iterdir()))
except PermissionError:
return -1
Part 5 — User and Group Management¶
pwd and grp: Reading User/Group Databases¶
import pwd
import grp
# Like: getent passwd username
user = pwd.getpwnam("www-data")
print(f"UID: {user.pw_uid}")
print(f"GID: {user.pw_gid}")
print(f"Home: {user.pw_dir}")
print(f"Shell: {user.pw_shell}")
# Like: getent passwd (all users)
for u in pwd.getpwall():
if u.pw_uid >= 1000 and u.pw_uid < 65534: # human users
print(f"{u.pw_name} (UID {u.pw_uid}): {u.pw_dir}")
# Like: getent group
for g in grp.getgrall():
if g.gr_mem: # groups with members
print(f"{g.gr_name} (GID {g.gr_gid}): {', '.join(g.gr_mem)}")
# Like: id username
def user_info(username):
"""Get full user info like the `id` command."""
u = pwd.getpwnam(username)
groups = [g.gr_name for g in grp.getgrall() if username in g.gr_mem]
primary = grp.getgrgid(u.pw_gid).gr_name
return {
"uid": u.pw_uid,
"gid": u.pw_gid,
"primary_group": primary,
"groups": groups,
"home": u.pw_dir,
"shell": u.pw_shell,
}
# Like: who / w (logged-in users)
import utmp # pip install utmp — or parse /var/run/utmp manually
# Or the quick way:
import subprocess
result = subprocess.run(["who"], capture_output=True, text=True)
print(result.stdout)
User Provisioning Script¶
In bash you'd write useradd -m -s /bin/bash -G sudo,docker $user. Here's the Python
equivalent that's actually maintainable:
import subprocess
import pwd
import secrets
import string
from pathlib import Path
def create_user(username, groups=None, shell="/bin/bash", ssh_key=None):
"""Create a user account with optional group membership and SSH key."""
# Check if user exists
try:
pwd.getpwnam(username)
print(f"User {username} already exists")
return False
except KeyError:
pass
# Build useradd command
cmd = ["useradd", "-m", "-s", shell]
if groups:
cmd.extend(["-G", ",".join(groups)])
cmd.append(username)
subprocess.run(cmd, check=True)
# Generate temporary password
alphabet = string.ascii_letters + string.digits + string.punctuation
temp_pass = "".join(secrets.choice(alphabet) for _ in range(16))
subprocess.run(
["chpasswd"],
input=f"{username}:{temp_pass}",
text=True,
check=True,
)
# Force password change on first login
subprocess.run(["passwd", "-e", username], check=True)
# Set up SSH key if provided
if ssh_key:
user = pwd.getpwnam(username)
ssh_dir = Path(user.pw_dir) / ".ssh"
ssh_dir.mkdir(mode=0o700, exist_ok=True)
auth_keys = ssh_dir / "authorized_keys"
auth_keys.write_text(ssh_key + "\n")
auth_keys.chmod(0o600)
# chown to the new user
import os
os.chown(ssh_dir, user.pw_uid, user.pw_gid)
os.chown(auth_keys, user.pw_uid, user.pw_gid)
print(f"Created user {username} (temp password: {temp_pass})")
return True
Part 6 — Networking¶
socket: DNS, Ports, and Connections¶
import socket
# Like: hostname
print(socket.gethostname())
# Like: hostname -f
print(socket.getfqdn())
# Like: dig / nslookup
ip = socket.gethostbyname("google.com")
print(f"google.com → {ip}")
# Reverse DNS: like dig -x
try:
hostname = socket.gethostbyaddr("8.8.8.8")
print(f"8.8.8.8 → {hostname[0]}")
except socket.herror:
print("No reverse DNS")
# Like: getent hosts
info = socket.getaddrinfo("google.com", 443, proto=socket.IPPROTO_TCP)
for family, socktype, proto, canonname, sockaddr in info:
print(f" {sockaddr[0]}:{sockaddr[1]}")
# --- Port checking ---
# Like: nc -zv host port / nmap -p port host
def check_port(host, port, timeout=3):
"""Check if a TCP port is open."""
try:
with socket.create_connection((host, port), timeout=timeout):
return True
except (ConnectionRefusedError, TimeoutError, OSError):
return False
# Scan common ports
services = {22: "SSH", 80: "HTTP", 443: "HTTPS", 3306: "MySQL", 5432: "PostgreSQL",
6379: "Redis", 8080: "Alt-HTTP", 9090: "Prometheus"}
for port, name in services.items():
status = "OPEN" if check_port("localhost", port) else "closed"
print(f" {port:>5} ({name:<12}): {status}")
Getting Local IP Addresses¶
import socket
import fcntl
import struct
# Quick way: what IP would we use to reach the internet?
def get_primary_ip():
"""Get the primary outbound IP address."""
with socket.socket(socket.AF_INET, socket.SOCK_DGRAM) as s:
s.connect(("8.8.8.8", 80)) # doesn't actually send anything
return s.getsockname()[0]
print(f"Primary IP: {get_primary_ip()}")
# All interfaces (using psutil — much cleaner than parsing ip addr)
import psutil
for iface, addrs in psutil.net_if_addrs().items():
for addr in addrs:
if addr.family == socket.AF_INET:
print(f"{iface}: {addr.address}/{addr.netmask}")
Part 7 — Service Management¶
systemd via subprocess¶
There's no pure-Python systemd library in the standard library, but subprocess gives you clean access.
import subprocess
import json
def systemctl(action, service):
"""Run a systemctl command and return success/failure."""
result = subprocess.run(
["systemctl", action, service],
capture_output=True, text=True,
)
return result.returncode == 0
def service_status(service):
"""Get structured service status."""
result = subprocess.run(
["systemctl", "show", service,
"--property=ActiveState,SubState,MainPID,ExecMainStartTimestamp,"
"MemoryCurrent,CPUUsageNSec,NRestarts"],
capture_output=True, text=True,
)
if result.returncode != 0:
return None
info = {}
for line in result.stdout.strip().splitlines():
key, _, value = line.partition("=")
info[key] = value
return info
# --- Usage ---
# Like: systemctl status nginx
status = service_status("nginx")
if status:
print(f"State: {status['ActiveState']} ({status['SubState']})")
print(f"PID: {status['MainPID']}")
print(f"Restarts: {status['NRestarts']}")
# Like: systemctl restart nginx
if not systemctl("restart", "nginx"):
print("Failed to restart nginx!")
# Like: systemctl is-active nginx
if systemctl("is-active", "nginx"):
print("nginx is running")
# --- List failed units ---
# Like: systemctl --failed
result = subprocess.run(
["systemctl", "list-units", "--failed", "--no-legend", "--plain"],
capture_output=True, text=True,
)
failed = [line.split()[0] for line in result.stdout.strip().splitlines() if line.strip()]
if failed:
print(f"FAILED UNITS: {', '.join(failed)}")
Journal Querying¶
import subprocess
import json
from datetime import datetime, timedelta
def journal_errors(service, since_minutes=60):
"""Get recent error-level journal entries for a service."""
since = (datetime.now() - timedelta(minutes=since_minutes)).strftime("%Y-%m-%d %H:%M:%S")
result = subprocess.run(
["journalctl", "-u", service, "--since", since,
"-p", "err", "-o", "json", "--no-pager"],
capture_output=True, text=True,
)
entries = []
for line in result.stdout.strip().splitlines():
try:
entry = json.loads(line)
entries.append({
"timestamp": entry.get("__REALTIME_TIMESTAMP"),
"message": entry.get("MESSAGE"),
"priority": entry.get("PRIORITY"),
})
except json.JSONDecodeError:
pass
return entries
errors = journal_errors("nginx", since_minutes=30)
for e in errors:
print(f" {e['message']}")
Part 8 — File Watching¶
watchdog: React to Filesystem Changes¶
In bash you'd use inotifywait -m -r /path. Python's watchdog library is the standard
answer.
import time
from watchdog.observers import Observer
from watchdog.events import FileSystemEventHandler
class LogHandler(FileSystemEventHandler):
"""React to file changes in a directory."""
def on_modified(self, event):
if not event.is_directory and event.src_path.endswith(".log"):
print(f"Modified: {event.src_path}")
def on_created(self, event):
if not event.is_directory:
print(f"New file: {event.src_path}")
def on_deleted(self, event):
if not event.is_directory:
print(f"Deleted: {event.src_path}")
# Watch /var/log for changes
observer = Observer()
observer.schedule(LogHandler(), "/var/log", recursive=True)
observer.start()
try:
while True:
time.sleep(1)
except KeyboardInterrupt:
observer.stop()
observer.join()
Part 9 — Scheduling and Cron¶
python-crontab: Manage Cron Jobs Programmatically¶
from crontab import CronTab
# Like: crontab -l
cron = CronTab(user="root")
for job in cron:
print(f" {job}")
# Like: (crontab -l; echo "0 2 * * * /usr/local/bin/backup.sh") | crontab -
job = cron.new(command="/usr/local/bin/backup.sh", comment="nightly backup")
job.setall("0 2 * * *") # 2:00 AM daily
cron.write()
# Like: crontab -l | grep -v backup | crontab -
cron.remove_all(comment="nightly backup")
cron.write()
# Validate a schedule
job = cron.new(command="/bin/true")
job.setall("*/5 * * * *") # every 5 minutes
print(f"Valid: {job.is_valid()}")
print(f"Next run: {job.schedule().get_next()}")
sched: In-Process Scheduling¶
For scripts that need to run tasks on a schedule without cron:
import sched
import time
scheduler = sched.scheduler(time.time, time.sleep)
def check_disk():
"""Check disk usage and alert if > 90%."""
import shutil
usage = shutil.disk_usage("/")
pct = usage.used / usage.total * 100
if pct > 90:
print(f"ALERT: Disk usage at {pct:.1f}%")
# Re-schedule self (every 300 seconds)
scheduler.enter(300, 1, check_disk)
# Start the loop
scheduler.enter(0, 1, check_disk)
scheduler.run()
Part 10 — Archives and Compression¶
import tarfile
import zipfile
import gzip
import shutil
from pathlib import Path
# --- tar.gz ---
# Like: tar czf backup.tar.gz /opt/myapp/
with tarfile.open("/tmp/backup.tar.gz", "w:gz") as tar:
tar.add("/opt/myapp", arcname="myapp")
# Like: tar xzf backup.tar.gz -C /tmp/restore/
with tarfile.open("/tmp/backup.tar.gz", "r:gz") as tar:
tar.extractall("/tmp/restore", filter="data") # filter= for safety (Python 3.12+)
# Like: tar tzf backup.tar.gz
with tarfile.open("/tmp/backup.tar.gz", "r:gz") as tar:
for member in tar.getmembers():
print(f" {member.name} ({member.size} bytes)")
# --- zip ---
# Like: zip -r backup.zip /opt/myapp/
with zipfile.ZipFile("/tmp/backup.zip", "w", zipfile.ZIP_DEFLATED) as zf:
for f in Path("/opt/myapp").rglob("*"):
if f.is_file():
zf.write(f, f.relative_to("/opt"))
# --- gzip a single file ---
# Like: gzip -k access.log
with open("/var/log/access.log", "rb") as f_in:
with gzip.open("/var/log/access.log.gz", "wb") as f_out:
shutil.copyfileobj(f_in, f_out)
Part 11 — Environment and Configuration¶
os.environ: Environment Variables Done Right¶
import os
# Like: echo $HOME
home = os.environ["HOME"]
# Like: echo ${DB_HOST:-localhost} (with default)
db_host = os.environ.get("DB_HOST", "localhost")
# Like: export MY_VAR=value
os.environ["MY_VAR"] = "value"
# Like: env | grep DB_
db_vars = {k: v for k, v in os.environ.items() if k.startswith("DB_")}
# Like: source .env (parse a .env file)
def load_dotenv(path=".env"):
"""Minimal .env loader (for when you don't want python-dotenv)."""
env_file = Path(path)
if not env_file.exists():
return
for line in env_file.read_text().splitlines():
line = line.strip()
if not line or line.startswith("#"):
continue
key, _, value = line.partition("=")
# Strip optional quotes
value = value.strip().strip("'\"")
os.environ[key.strip()] = value
configparser: INI Files¶
import configparser
# Like: parsing my.cnf, php.ini, systemd unit files
config = configparser.ConfigParser()
config.read("/etc/myapp/config.ini")
db_host = config.get("database", "host", fallback="localhost")
db_port = config.getint("database", "port", fallback=5432)
debug = config.getboolean("app", "debug", fallback=False)
Part 12 — Putting It All Together: A System Health Check Script¶
Here's the kind of script that replaces a dozen bash one-liners:
#!/usr/bin/env python3
"""System health check — replaces 200 lines of bash."""
import shutil
import socket
from datetime import datetime
from pathlib import Path
import psutil
def check_disk(threshold=85):
"""Check all mounted filesystems."""
alerts = []
for part in psutil.disk_partitions():
try:
usage = psutil.disk_usage(part.mountpoint)
if usage.percent >= threshold:
alerts.append(
f"DISK {part.mountpoint}: {usage.percent}% "
f"({usage.free // (1024**3)} GB free)"
)
except PermissionError:
pass
return alerts
def check_memory(threshold=90):
"""Check RAM and swap usage."""
alerts = []
mem = psutil.virtual_memory()
if mem.percent >= threshold:
alerts.append(f"RAM: {mem.percent}% used ({mem.available // (1024**2)} MB available)")
swap = psutil.swap_memory()
if swap.percent >= threshold:
alerts.append(f"SWAP: {swap.percent}% used")
return alerts
def check_load():
"""Check system load average."""
load1, load5, load15 = psutil.getloadavg()
cpus = psutil.cpu_count()
alerts = []
if load5 > cpus * 0.8:
alerts.append(f"LOAD: {load5:.2f} (5min) on {cpus} CPUs")
return alerts
def check_services(services):
"""Check if critical services are running."""
import subprocess
alerts = []
for svc in services:
result = subprocess.run(
["systemctl", "is-active", svc],
capture_output=True, text=True,
)
if result.stdout.strip() != "active":
alerts.append(f"SERVICE {svc}: {result.stdout.strip()}")
return alerts
def check_ports(ports):
"""Check if expected ports are listening."""
alerts = []
for port, name in ports.items():
try:
with socket.create_connection(("localhost", port), timeout=2):
pass
except (ConnectionRefusedError, TimeoutError, OSError):
alerts.append(f"PORT {port} ({name}): not listening")
return alerts
def check_zombies():
"""Find zombie processes."""
zombies = []
for proc in psutil.process_iter(["pid", "name", "status"]):
if proc.info["status"] == psutil.STATUS_ZOMBIE:
zombies.append(f"ZOMBIE: PID {proc.info['pid']} ({proc.info['name']})")
return zombies
def main():
hostname = socket.gethostname()
now = datetime.now().strftime("%Y-%m-%d %H:%M:%S")
print(f"=== Health Check: {hostname} @ {now} ===\n")
all_alerts = []
checks = [
("Disk", check_disk),
("Memory", check_memory),
("Load", check_load),
("Zombies", check_zombies),
("Services", lambda: check_services(["nginx", "postgresql", "redis"])),
("Ports", lambda: check_ports({
80: "HTTP", 443: "HTTPS", 5432: "PostgreSQL", 6379: "Redis",
})),
]
for name, check_fn in checks:
alerts = check_fn()
if alerts:
all_alerts.extend(alerts)
for a in alerts:
print(f" WARN {a}")
else:
print(f" OK {name}")
print()
if all_alerts:
print(f"RESULT: {len(all_alerts)} warning(s)")
return 1
else:
print("RESULT: All checks passed")
return 0
if __name__ == "__main__":
raise SystemExit(main())
Part 13 — Common Patterns and Idioms¶
Pattern: Retry with Backoff¶
import time
import random
def retry(fn, max_attempts=3, base_delay=1, max_delay=30):
"""Retry a function with exponential backoff and jitter."""
for attempt in range(1, max_attempts + 1):
try:
return fn()
except Exception as e:
if attempt == max_attempts:
raise
delay = min(base_delay * (2 ** (attempt - 1)), max_delay)
delay *= (0.5 + random.random()) # jitter
print(f"Attempt {attempt} failed ({e}), retrying in {delay:.1f}s...")
time.sleep(delay)
Pattern: PID File¶
import os
import sys
from pathlib import Path
def write_pidfile(path="/var/run/myapp.pid"):
"""Write PID file, exit if already running."""
pidfile = Path(path)
if pidfile.exists():
old_pid = int(pidfile.read_text().strip())
try:
os.kill(old_pid, 0) # signal 0 = check if alive
print(f"Already running (PID {old_pid})")
sys.exit(1)
except ProcessLookupError:
pass # stale pidfile, proceed
pidfile.write_text(str(os.getpid()))
def remove_pidfile(path="/var/run/myapp.pid"):
Path(path).unlink(missing_ok=True)
Pattern: Atomic File Write¶
import tempfile
import os
from pathlib import Path
def atomic_write(path, content, mode=0o644):
"""Write a file atomically (write to temp, then rename).
Prevents partial writes if the process crashes mid-write.
"""
target = Path(path)
fd, tmp_path = tempfile.mkstemp(dir=target.parent, prefix=f".{target.name}.")
try:
with os.fdopen(fd, "w") as f:
f.write(content)
os.chmod(tmp_path, mode)
os.rename(tmp_path, target) # atomic on same filesystem
except BaseException:
os.unlink(tmp_path)
raise
Pattern: Log Tailer¶
import time
from pathlib import Path
def tail_follow(path, callback):
"""Like tail -f, but in Python. Handles log rotation."""
p = Path(path)
pos = p.stat().st_size # start at end
inode = p.stat().st_ino
while True:
stat = p.stat()
# Check for log rotation (inode changed)
if stat.st_ino != inode:
pos = 0
inode = stat.st_ino
if stat.st_size > pos:
with open(p) as f:
f.seek(pos)
for line in f:
callback(line.rstrip())
pos = f.tell()
time.sleep(0.5)
# Usage:
# tail_follow("/var/log/syslog", lambda line: print(line) if "ERROR" in line else None)
Module Quick Reference¶
Standard Library (no install needed)¶
| Module | Use for |
|---|---|
os |
Environment, UID/GID, file descriptors, os.walk |
os.path |
Legacy path manipulation (prefer pathlib) |
pathlib |
Modern path objects, file I/O, globbing |
shutil |
Copy/move/delete trees, disk usage, which() |
subprocess |
Running external commands |
signal |
Signal handlers (SIGTERM, SIGINT, etc.) |
atexit |
Cleanup on exit |
tempfile |
Temp files and directories |
glob |
Shell-style wildcards (prefer pathlib.glob) |
fnmatch |
Filename pattern matching |
stat |
File permission constants (S_IRUSR, etc.) |
fcntl |
File locking (flock) |
socket |
DNS, port checks, hostname |
platform |
OS/arch/Python version info |
pwd |
User database (/etc/passwd) |
grp |
Group database (/etc/group) |
sched |
In-process event scheduler |
configparser |
INI file parsing |
tarfile |
tar/tar.gz/tar.bz2 creation and extraction |
zipfile |
ZIP creation and extraction |
gzip |
gzip compression |
struct |
Binary data packing/unpacking |
datetime |
Date/time manipulation |
time |
Sleep, monotonic clock, epoch time |
re |
Regex (replaces grep, sed, awk patterns) |
json |
JSON parsing/generation |
csv |
CSV parsing/generation |
logging |
Structured logging |
argparse |
CLI argument parsing |
textwrap |
Text wrapping/dedenting |
itertools |
Efficient iteration (islice for head/tail) |
collections |
Counter, defaultdict, deque |
concurrent.futures |
Thread/process pools |
Third-Party (pip install)¶
| Package | Use for | Replaces |
|---|---|---|
psutil |
Process/system monitoring | ps, top, free, df, netstat |
python-crontab |
Cron management | crontab -l/-e |
watchdog |
Filesystem event monitoring | inotifywait |
distro |
Linux distribution info | lsb_release |
filelock |
Cross-platform file locking | flock |
paramiko |
SSH connections | ssh, scp |
fabric |
Remote command execution | ssh in loops |
pexpect |
Interactive process control | expect |
click |
CLI framework | argparse (more ergonomic) |
rich |
Terminal formatting/tables | printf, column -t |
python-dotenv |
.env file loading |
source .env |
schedule |
Human-friendly scheduling | cron expressions |
sh |
Shell command wrapper | subprocess (terser syntax) |
plumbum |
Shell command piping | Shell pipes |
When to Stay in Bash¶
Python isn't always the right tool. Stay in bash when:
- One-liners:
grep ERROR /var/log/syslog | tail -20— don't write 15 lines of Python for this - Glue between commands: A 5-line script that runs
terraform plan, checks the exit code, and runsterraform apply— bash is fine - Interactive/ad-hoc: You're poking around a server debugging something. Use bash.
- Boot scripts:
/etc/init.d/, early boot — Python might not be available yet - Performance-critical text processing:
awkprocessing a 10GB log file will beat Python. Use the right tool.
Switch to Python when:
- The script is over ~50 lines of bash
- You need error handling beyond set -e
- You're parsing structured data (JSON, YAML, CSV)
- You need retry logic, backoff, or complex control flow
- You need to maintain state between runs
- Multiple people will maintain the script
- You need tests
- You're tired of quoting bugs and word-splitting surprises
Exercises¶
-
Disk Alert Script: Write a script that checks all mounted filesystems, alerts if any exceed 85% usage, and writes results to a JSON file. Use only
shutilandpsutil. -
Process Inventory: Write a tool that lists all processes, groups them by user, and shows total CPU/memory per user. Sort by memory descending. (Hint:
psutil.process_iter -
collections.defaultdict) -
Port Scanner: Write a function that takes a hostname and a range of ports, checks which are open using
socket, and returns results. Add a--timeoutand--threadsCLI flag usingargparse. Useconcurrent.futures.ThreadPoolExecutorfor parallelism. -
Log Rotator: Write a script that rotates log files:
app.log→app.log.1→app.log.2→ ... → deleteapp.log.5. Handle the case where the process has the file open (usecopytruncatestrategy). -
Service Monitor: Write a daemon that checks a list of systemd services every 60 seconds, logs state changes, and writes a status file. Use
signalfor graceful shutdown andatexitfor cleanup. -
User Audit: Write a script that compares
/etc/passwdusers against an expected list (from a YAML file), reports additions/removals, and checks that all human users (UID >= 1000) have a valid shell and home directory that exists.
Further Reading¶
- psutil docs: Process and system monitoring — the single most useful third-party library for system automation
- pathlib docs: PEP 428 — your new best friend for file operations
- subprocess docs: Especially the security considerations section
- signal docs: Signal handling and the caveats around threads
- The
training/library/lessons/python-for-ops-the-bash-experts-bridge.mdlesson in this repo covers subprocess, pathlib, and requests in more depth - The
training/library/lessons/python-automating-everything-apis-and-infrastructure.mdlesson covers the API/cloud side of automation