---
tags:
- devops
- l1
- flashcard-deck
- database-internals
---
<!-- wiki:breadcrumb:start -->
[Portal](../../../../library/portal/index.md) | **Level:** [L1: Foundations](../../../../library/portal/levels.md) | **Topics:** [Database Internals](../../../../library/portal/topics.md) | **Domain:** DevOps & Tooling
<!-- wiki:breadcrumb:end -->

id	category	difficulty	tags	question	answer	source_path
db-internals/a1b2c3d4f5e6	db-internals	easy	database, indexing, btree	What is the default index type in PostgreSQL and what queries does it support?	B-tree. It supports equality (=) and range queries (<, >, BETWEEN) on sortable data. It is the most commonly used index type.\n\nRemember: an index is a data structure (usually B-tree) that speeds up lookups at the cost of slower writes and extra storage. Like a book's index — you don't read every page to find a topic.	training/library/topics/database-internals/primer.md
db-internals/b2c3d4e5a6f7	db-internals	easy	database, indexing, explain	What command shows the actual execution plan with real timings for a query in PostgreSQL?	EXPLAIN (ANALYZE, BUFFERS, FORMAT TEXT) followed by the query. ANALYZE is required to actually execute the query and show real timing and row counts instead of estimates.\n\nRemember: EXPLAIN ANALYZE actually executes the query — be careful with UPDATE/DELETE! Wrap in a transaction: BEGIN; EXPLAIN ANALYZE ...; ROLLBACK;	training/library/topics/database-internals/primer.md
db-internals/c3d4e5f6b7a8	db-internals	easy	database, indexing, statistics	What PostgreSQL command updates table statistics used by the query planner?	ANALYZE table_name; (for a specific table) or just ANALYZE; (for the entire database). Stale statistics cause the planner to choose suboptimal execution plans.\n\nRemember: EXPLAIN ANALYZE actually runs the query and shows real execution times. Look for: Seq Scan (missing index), Nested Loop on large tables (consider Hash Join), high actual vs estimated rows (stale statistics).	training/library/topics/database-internals/primer.md
db-internals/d4e5f6a7c8b9	db-internals	medium	database, indexing, missing	How can you identify tables that may be missing indexes in PostgreSQL?	Query pg_stat_user_tables for tables with high seq_scan counts and high seq_tup_read values relative to idx_scan. A large table with many sequential scans and few or no index scans likely needs an index on commonly filtered columns.\n\nExample: SELECT schemaname, relname, seq_scan, idx_scan FROM pg_stat_user_tables WHERE seq_scan > 1000 AND idx_scan < 50 ORDER BY seq_scan DESC; — high seq_scan + low idx_scan = missing index candidate.	training/library/topics/database-internals/primer.md
db-internals/e5f6a7b8d9c0	db-internals	medium	database, indexing, explain	In EXPLAIN ANALYZE output, what does a Seq Scan on a large table indicate and how do you fix it?	A sequential scan on a large table typically means there is no suitable index for the query's WHERE clause or join condition. Fix by creating an index on the filtered columns. Verify improvement by running EXPLAIN ANALYZE again and confirming an Index Scan or Index Only Scan.\n\nRemember: EXPLAIN ANALYZE actually runs the query and shows real execution times. Look for: Seq Scan (missing index), Nested Loop on large tables (consider Hash Join), high actual vs estimated rows (stale statistics).	training/library/topics/database-internals/primer.md
db-internals/f6a7b8c9e0d1	db-internals	medium	database, indexing, pg_stat_statements	How do you find the most expensive queries in PostgreSQL using pg_stat_statements?	SELECT query, calls, total_exec_time, mean_exec_time, rows FROM pg_stat_statements ORDER BY total_exec_time DESC LIMIT 10; This requires the pg_stat_statements extension to be loaded.\n\nRemember: PostgreSQL statistics views (pg_stat_user_tables, pg_stat_statements) are your best friends for performance tuning. Enable pg_stat_statements in shared_preload_libraries.	training/library/topics/database-internals/primer.md
db-internals/a7b8c9d0f1e2	db-internals	medium	database, indexing, bloat	What is index bloat and how do you fix it online in PostgreSQL 12+?	Index bloat occurs when dead tuples accumulate in an index, wasting disk space and slowing queries. Fix with REINDEX INDEX CONCURRENTLY idx_name, which rebuilds the index without locking the table for writes.\n\nRemember: an index is a data structure (usually B-tree) that speeds up lookups at the cost of slower writes and extra storage. Like a book's index — you don't read every page to find a topic.	training/library/topics/database-internals/primer.md
db-internals/b8c9d0e1a2f3	db-internals	hard	database, indexing, covering	What is a covering index and how do you create one in PostgreSQL?	A covering index includes all columns needed by a query so it can be satisfied entirely from the index (index-only scan) without accessing the table. Create with: CREATE INDEX idx_name ON table (filter_col) INCLUDE (col1, col2). The INCLUDE columns are stored in the index but not used for searching.\n\nRemember: an index is a data structure (usually B-tree) that speeds up lookups at the cost of slower writes and extra storage. Like a book's index — you don't read every page to find a topic.	training/library/topics/database-internals/primer.md
db-internals/c9d0e1f2b3a4	db-internals	hard	database, indexing, composite	Why does column order matter in a composite index?	A composite index on (A, B) efficiently supports queries filtering on A alone or on A and B together, but not on B alone. The index follows a leftmost-prefix rule — it can only skip to a specific value of the first column, then scan within it. Ordering columns by selectivity and query patterns is critical.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/d0e1f2a3c4b5	db-internals	hard	database, indexing, mysql	How does covering index behavior differ between PostgreSQL and MySQL InnoDB?	"In PostgreSQL, you explicitly create covering indexes with INCLUDE columns. In MySQL InnoDB, every secondary index implicitly includes the primary key columns, and the clustered index (primary key) stores the full row. Look for ""Using index"" in MySQL EXPLAIN Extra column to confirm an index-only scan.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain."	training/library/topics/database-internals/primer.md
db-internals/e1f2a3b4c5d6	db-internals	medium	database, indexing, composite, ordering	How should you decide column order in a composite index when multiple columns are filtered?	Place the most selective column (fewest matching rows) first, then the next most selective, unless one column is used in range scans — put equality columns before range columns. For example, (status, created_at) is better than (created_at, status) if status is tested with = and created_at with BETWEEN.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	
db-internals/f2a3b4c5d6e7	db-internals	medium	database, indexing, index-only-scan	What is an index-only scan and what condition must be met for PostgreSQL to use one?	"An index-only scan satisfies the query entirely from the index without visiting the heap (table). PostgreSQL can use it when all columns in SELECT, WHERE, and ORDER BY are in the index AND the visibility map shows the pages are all-visible (recently vacuumed). Check EXPLAIN for ""Index Only Scan"" and watch the ""Heap Fetches"" count.\n\nRemember: an index is a data structure (usually B-tree) that speeds up lookups at the cost of slower writes and extra storage. Like a book's index — you don't read every page to find a topic."	
db-internals/a3b4c5d6e7f8	db-internals	easy	database, indexing, anti-pattern	When should you NOT add an index to a table?	Avoid indexing when the table is very small (seq scan is faster), the column has very low selectivity (e.g., boolean with 50/50 distribution), the table is write-heavy with few reads (indexes slow down INSERT/UPDATE/DELETE), or you already have too many indexes causing write amplification and bloat.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	
db-internals/b4c5d6e7f8a9	db-internals	medium	database, indexing, partial	What is a partial index and when is it useful?	A partial index includes only rows matching a WHERE predicate: CREATE INDEX idx ON orders (created_at) WHERE status = 'pending'. It is smaller and faster than a full index because it skips rows that don't match the predicate. Useful when queries consistently filter on a subset of rows.\n\nRemember: an index is a data structure (usually B-tree) that speeds up lookups at the cost of slower writes and extra storage. Like a book's index — you don't read every page to find a topic.	
db-internals/c5d6e7f8a9b0	db-internals	hard	database, indexing, bloat, maintenance	What is the difference between REINDEX and ANALYZE, and when do you run each?	REINDEX rebuilds an index from scratch to eliminate bloat (dead space from updates/deletes). ANALYZE updates the planner's statistics about data distribution. Run REINDEX when index bloat causes slow scans or excessive disk usage. Run ANALYZE after bulk loads or major data changes so the planner picks optimal plans.\n\nRemember: an index is a data structure (usually B-tree) that speeds up lookups at the cost of slower writes and extra storage. Like a book's index — you don't read every page to find a topic.	
db-internals/d6e7f8a9b0c1	db-internals	hard	database, indexing, btree, hash, gin, gist	Compare B-tree, hash, GIN, and GiST index types in PostgreSQL.	B-tree: default, supports equality and range queries on sortable data. Hash: equality only, smaller than B-tree for that case but not WAL-logged before PG 10.\nGIN (Generalized Inverted Index): best for multi-valued columns like arrays, JSONB, and full-text search.\nGiST (Generalized Search Tree): best for geometric, range, and proximity queries (e.g., PostGIS, tsquery).\n\nRemember: B-tree = balanced tree where each node can have many children. Keeps data sorted, supports O(log n) lookups. Used by PostgreSQL, MySQL, SQLite for indexes.	
db-internals/e7f8a9b0c1d2	db-internals	medium	database, indexing, unused	How do you find and safely remove unused indexes in PostgreSQL?	Query pg_stat_user_indexes for indexes with idx_scan = 0 (or very low) over a representative time period. Before dropping, verify the index is not used for unique constraints or foreign key lookups. Use DROP INDEX CONCURRENTLY to avoid locking the table during removal.\n\nExample: SELECT schemaname, relname, seq_scan, idx_scan FROM pg_stat_user_tables WHERE seq_scan > 1000 AND idx_scan < 50 ORDER BY seq_scan DESC; — high seq_scan + low idx_scan = missing index candidate.	
db-internals/f8a9b0c1d2e3	db-internals	hard	database, indexing, expression	What is an expression index and when would you use one?	An expression index indexes the result of a function or expression: CREATE INDEX idx ON users (lower(email)). Use it when queries filter on transformed values (e.g., case-insensitive search, date truncation). Without it, PostgreSQL cannot use a regular index on the column because the expression changes the lookup value.\n\nRemember: an index is a data structure (usually B-tree) that speeds up lookups at the cost of slower writes and extra storage. Like a book's index — you don't read every page to find a topic.	
db-internals/a1b2c3e4d5f6	db-internals	easy	database, locking, types	What PostgreSQL lock level do INSERT, UPDATE, and DELETE operations acquire?	RowExclusiveLock. This lock conflicts with ShareLock and AccessExclusiveLock but allows concurrent row modifications on different rows.\n\nRemember: lock granularity: row > page > table > database. Finer locks = more concurrency but more overhead. Most OLTP workloads use row-level locks.	training/library/topics/database-internals/primer.md
db-internals/b2c3d4f5e6a7	db-internals	easy	database, locking, exclusive	What operations acquire an AccessExclusiveLock in PostgreSQL and why is it dangerous?	ALTER TABLE, DROP TABLE, and TRUNCATE acquire AccessExclusiveLock, which conflicts with every other lock type. This blocks all concurrent reads and writes on the table, potentially causing connection pile-ups if the operation is slow.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/c3d4e5a6f7b8	db-internals	easy	database, locking, timeout	How do you prevent a PostgreSQL query from waiting indefinitely for a lock?	SET lock_timeout = '5s'; This causes the query to fail immediately with an error if it cannot acquire the requested lock within the timeout period, instead of blocking indefinitely.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/d4e5f6b7a8c9	db-internals	medium	database, locking, viewing	How do you find which queries are blocked by locks in PostgreSQL?	Join pg_locks with pg_stat_activity: query pg_locks for rows WHERE NOT granted to find blocked locks, then join to pg_stat_activity on pid to see the blocked and blocking queries, users, and durations.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/e5f6a7c8b9d0	db-internals	medium	database, locking, deadlock	How does PostgreSQL detect and resolve deadlocks?	PostgreSQL runs a deadlock detector every deadlock_timeout interval (default 1 second). When it detects a cycle of lock dependencies, it terminates the youngest transaction in the cycle and returns a deadlock_detected error to that client.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/f6a7b8d9c0e1	db-internals	medium	database, locking, optimistic, pessimistic	What is the difference between optimistic and pessimistic locking?	Pessimistic locking acquires locks before modifying data (SELECT FOR UPDATE), preventing conflicts upfront — best for high-contention scenarios. Optimistic locking checks a version or timestamp at commit time and retries on conflict — best for low-contention scenarios with longer transactions.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/a7b8c9e0d1f2	db-internals	medium	database, locking, for-update	What does SELECT FOR UPDATE do and when would you use it?	It acquires a row-level lock (RowShareLock) on the selected rows, preventing other transactions from modifying or locking them until the current transaction completes. Use it for pessimistic locking patterns like inventory decrement where you need to read-then-write atomically.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/b8c9d0f1e2a3	db-internals	hard	database, locking, migration	Why should you set lock_timeout before running ALTER TABLE in production, and what happens if it times out?	ALTER TABLE acquires AccessExclusiveLock, which blocks all reads and writes. If the table has active queries, the ALTER waits for them, and new queries queue behind it — creating a cascading pile-up. Setting lock_timeout = '3s' causes the ALTER to fail fast if it cannot get the lock, avoiding the pile-up. On timeout, the ALTER is aborted and no schema change is made.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/c9d0e1a2f3b4	db-internals	hard	database, locking, skip-locked	What do NOWAIT and SKIP LOCKED do with SELECT FOR UPDATE, and what are their use cases?	NOWAIT causes the query to fail immediately if any selected row is already locked (instead of waiting). SKIP LOCKED silently skips rows that are locked by other transactions. SKIP LOCKED is ideal for work-queue patterns where multiple workers pull jobs from the same table without blocking each other.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/d0e1f2b3a4c5	db-internals	hard	database, locking, mysql	How does MySQL InnoDB lock wait behavior differ from PostgreSQL?	MySQL InnoDB uses innodb_lock_wait_timeout (default 50 seconds) compared to PostgreSQL lock_timeout (not set by default). MySQL InnoDB also uses gap locks at REPEATABLE READ isolation to prevent phantom reads, which can cause unexpected lock conflicts on ranges of rows that PostgreSQL avoids through its MVCC implementation.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/54eb883e	db-internals	easy	database, locking, row-level, table-level	What is the difference between row-level and table-level locks in PostgreSQL?	Row-level locks (e.g., from SELECT FOR UPDATE, UPDATE, DELETE) lock individual rows, allowing concurrent access to other rows in the same table. Table-level locks (e.g., AccessExclusiveLock from ALTER TABLE) lock the entire table, blocking all concurrent operations. Row-level locks scale better for concurrent workloads.\n\nRemember: lock granularity: row > page > table > database. Finer locks = more concurrency but more overhead. Most OLTP workloads use row-level locks.	
db-internals/ec6c7bc7	db-internals	medium	database, locking, deadlock, resolution	What are practical strategies to prevent deadlocks in application code?	Always acquire locks in a consistent order across all transactions (e.g., sort rows by primary key before locking). Keep transactions short to reduce the window for conflicts. Use lock_timeout to fail fast instead of waiting indefinitely. Retry deadlocked transactions with exponential backoff in application code.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	
db-internals/cc93b06d	db-internals	medium	database, locking, escalation	What is lock escalation and does PostgreSQL do it?	Lock escalation is when a database engine automatically converts many fine-grained locks (row-level) into a coarser lock (table-level) to reduce memory overhead. PostgreSQL does NOT escalate locks — it maintains row-level locks regardless of count. SQL Server and some other engines do escalate, which can cause unexpected blocking.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	
db-internals/7707fa91	db-internals	medium	database, locking, advisory	What are advisory locks in PostgreSQL and when would you use them?	Advisory locks are application-defined locks that PostgreSQL manages but does not enforce on any table or row. Acquire with pg_advisory_lock(key) or pg_try_advisory_lock(key). Use them for application-level coordination like ensuring only one worker processes a job, rate limiting, or distributed mutex patterns. They must be explicitly released.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	
db-internals/4fc034e8	db-internals	hard	database, locking, long-running	How do long-running transactions cause lock pile-ups in production?	A long-running transaction holds its locks for the entire duration. If an ALTER TABLE then arrives, it requests AccessExclusiveLock and queues behind the long transaction. All subsequent queries on that table queue behind the ALTER, creating a cascading pile-up. Even read queries are blocked because they cannot acquire AccessShareLock while AccessExclusiveLock is pending.\n\nRemember: transaction isolation levels (low to high): Read Uncommitted, Read Committed, Repeatable Read, Serializable. Higher isolation = fewer anomalies but more locking/overhead.	
db-internals/09487d09	db-internals	medium	database, locking, monitoring, pg_stat_activity	How do you use pg_stat_activity and pg_locks together to diagnose lock waits?	Join pg_stat_activity (for query text and duration) with pg_locks (for lock details): SELECT blocked.pid, blocked.query, blocking.pid AS blocker_pid, blocking.query AS blocker_query FROM pg_stat_activity blocked JOIN pg_locks bl ON bl.pid = blocked.pid JOIN pg_locks bk ON bk.relation = bl.relation AND bk.granted AND bl.pid != bk.pid JOIN pg_stat_activity blocking ON blocking.pid = bk.pid WHERE NOT bl.granted;\n\nRemember: PostgreSQL statistics views (pg_stat_user_tables, pg_stat_statements) are your best friends for performance tuning. Enable pg_stat_statements in shared_preload_libraries.	
db-internals/d63e76c9	db-internals	hard	database, locking, optimistic, implementation	How do you implement optimistic locking with a version column?	Add a version integer column to the table. On read, fetch the current version. On update: UPDATE items SET data = $1, version = version + 1 WHERE id = $2 AND version = $3. If zero rows are affected, another transaction modified the row — retry the entire read-modify-write cycle. This avoids holding locks during user think time.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	
db-internals/73680640	db-internals	easy	database, locking, share	What is a ShareLock and when does PostgreSQL acquire one?	ShareLock is acquired by CREATE INDEX (non-concurrent). It allows concurrent reads but blocks writes (INSERT, UPDATE, DELETE). This is why CREATE INDEX on a large table can cause write outages. Use CREATE INDEX CONCURRENTLY instead, which takes a weaker lock and allows writes to continue.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	
db-internals/a1b2d3c4e5f6	db-internals	easy	database, read-replicas, purpose	What is the primary purpose of read replicas in a database architecture?	To scale read traffic by offloading read queries from the primary to one or more replica nodes, reducing load on the primary and improving overall throughput for read-heavy workloads.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/b2c3e4d5f6a7	db-internals	easy	database, read-replicas, monitoring	How do you check if a PostgreSQL instance is currently running as a replica?	SELECT pg_is_in_recovery(); — returns true if the instance is in recovery mode (i.e., running as a replica), false if it is the primary.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/c3d4f5e6a7b8	db-internals	easy	database, read-replicas, lag	How do you measure replication delay in seconds on a PostgreSQL replica?	SELECT now() - pg_last_xact_replay_timestamp() AS replication_delay; This shows the time since the last replayed transaction from the primary.\n\nRemember: streaming replication sends WAL records to replicas in real-time. Synchronous = guaranteed consistency, higher latency. Asynchronous = faster, risk of data loss on failover.	training/library/topics/database-internals/primer.md
db-internals/d4e5a6f7b8c9	db-internals	medium	database, read-replicas, routing	Name three tools used for connection routing between a PostgreSQL primary and read replicas.	PgBouncer (connection pooler, does not route by query type), Pgpool-II (query-aware load balancer that routes SELECTs to replicas), and HAProxy (TCP-level load balancing with health checks).\n\nRemember: each database connection consumes memory (~5-10 MB in PostgreSQL). Use connection pooling (PgBouncer, ProxySQL) to serve thousands of app connections with dozens of DB connections.	training/library/topics/database-internals/primer.md
db-internals/e5f6b7a8c9d0	db-internals	medium	database, read-replicas, read-after-write	What is the read-after-write consistency problem with read replicas?	When a user writes data to the primary and then immediately reads from a replica, the replica may not have replayed the write yet due to replication lag. The user sees stale data or gets a 404 for a record they just created. Fix by routing read-your-own-writes to the primary or a synchronous replica.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/f6a7c8b9d0e1	db-internals	medium	database, read-replicas, consistency	When should reads go to the primary vs an async replica?	Reads requiring strong consistency (authentication, payments, inventory checks) should go to the primary or a synchronous replica. Reads tolerating eventual consistency (dashboards, search, analytics, reporting) can safely go to async replicas.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/a7b8d9c0e1f2	db-internals	medium	database, read-replicas, alerting	What are typical alerting thresholds for replication lag on read replicas?	Warning at replication lag > 5 seconds (sustained for 2 minutes), Critical at lag > 30 seconds or replica disconnected (for 1 minute), Page when replica count drops below the minimum healthy count (e.g., fewer than 2 of 3 replicas connected).\n\nRemember: good alerting follows the RED method for services (Rate, Errors, Duration) and USE method for resources (Utilization, Saturation, Errors).\n\nRemember: streaming replication sends WAL records to replicas in real-time. Synchronous = guaranteed consistency, higher latency. Asynchronous = faster, risk of data loss on failover.	training/library/topics/database-internals/primer.md
db-internals/b8c9e0d1f2a3	db-internals	hard	database, read-replicas, synchronous-commit	How does PostgreSQL's synchronous_commit setting enable per-transaction consistency control with replicas?	synchronous_commit can be set per transaction: 'remote_apply' ensures the write is replayed on the sync replica before commit returns (safe for immediate reads from that replica). 'local' only waits for local WAL flush (faster but replica may lag). This lets you choose strong vs eventual consistency on a per-query basis.\n\nRemember: transaction isolation levels (low to high): Read Uncommitted, Read Committed, Repeatable Read, Serializable. Higher isolation = fewer anomalies but more locking/overhead.	training/library/topics/database-internals/primer.md
db-internals/c9d0f1e2a3b4	db-internals	hard	database, read-replicas, architecture	In a read replica architecture, what happens if all replicas become unreachable?	All read traffic either fails (if the application requires replicas) or falls back to the primary (if routing allows it), potentially overloading the primary. A well-designed system has circuit breakers to detect failed replicas, routes reads to the primary as a fallback with rate limiting, and alerts immediately so ops can restore replicas or add capacity.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/d0e1a2f3b4c5	db-internals	hard	database, read-replicas, primary-check	How do you verify from the primary that all expected replicas are connected and healthy?	SELECT application_name, client_addr, state, pg_wal_lsn_diff(pg_current_wal_lsn(), replay_lsn) AS replay_lag_bytes, replay_lag FROM pg_stat_replication; Check that the count matches expected replicas, all show state = 'streaming', and replay_lag_bytes is within acceptable thresholds.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/74ad501c	db-internals	medium	database, read-replicas, lag, measurement	What are two ways to measure replication lag and what are their limitations?	Time-based: SELECT now() - pg_last_xact_replay_timestamp(). Limitation: shows false lag when the primary is idle (no new transactions to replay).\nByte-based: compare pg_current_wal_lsn() on primary with pg_last_wal_replay_lsn() on replica. Limitation: bytes don't directly translate to time — a large transaction may inflate byte lag without meaning the replica is far behind.\n\nRemember: streaming replication sends WAL records to replicas in real-time. Synchronous = guaranteed consistency, higher latency. Asynchronous = faster, risk of data loss on failover.	
db-internals/129a500e	db-internals	hard	database, read-replicas, read-after-write, strategies	Name three strategies for achieving read-after-write consistency with async replicas.	1. Route read-your-own-writes to the primary for a short window after each write (e.g., 5 seconds).\n2. Include the primary's LSN in the write response and only read from a replica once its replay LSN has caught up.\n3. Use synchronous_commit = 'remote_apply' for critical writes so the replica is guaranteed to have replayed them before the commit returns.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	
db-internals/22cd2169	db-internals	medium	database, read-replicas, routing, proxysql, pgpool	How does Pgpool-II route queries differently from PgBouncer?	Pgpool-II parses SQL and routes SELECT statements to replicas and writes to the primary automatically. PgBouncer is a connection pooler only — it does not inspect queries or route by type. For query-aware routing with PgBouncer, the application must use separate connection strings for read and write pools.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	
db-internals/b5fa976a	db-internals	hard	database, read-replicas, promotion, failover	What steps are involved in promoting a replica to primary during failover?	1. Verify the replica is the most up-to-date (least lag).\n2. Fence or stop the old primary to prevent split-brain.\n3. Promote the replica: SELECT pg_promote() or pg_ctl promote.\n4. Update DNS or connection routing to point at the new primary.\n5. Reconfigure remaining replicas to follow the new primary.\n6. Verify replication resumes and applications reconnect.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	
db-internals/bb179e65	db-internals	medium	database, read-replicas, cascading	What is cascading replication and when is it useful?	Cascading replication allows a replica to stream WAL to other replicas instead of all replicas connecting directly to the primary. Configure with primary_conninfo pointing to an upstream replica. Useful when you have many replicas — it reduces network and CPU load on the primary by fanning out through intermediate replicas.\n\nRemember: streaming replication sends WAL records to replicas in real-time. Synchronous = guaranteed consistency, higher latency. Asynchronous = faster, risk of data loss on failover.	
db-internals/3084fbd0	db-internals	medium	database, read-replicas, health, monitoring	What metrics should you monitor on read replicas to detect problems early?	Replication lag (seconds and bytes), replay rate (WAL replayed per second), connection count vs max_connections, CPU and I/O utilization, streaming state in pg_stat_replication (should be 'streaming'), and query cancellation rate on the replica due to recovery conflicts (pg_stat_database_conflicts).\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	
db-internals/9727c0e6	db-internals	hard	database, read-replicas, anti-pattern	When can adding read replicas make performance worse instead of better?	When the workload is write-heavy (replicas must replay all writes but serve few reads), when replication lag causes frequent cache invalidation or application retries, when queries on replicas conflict with WAL replay causing cancellations (max_standby_streaming_delay), or when the connection routing layer adds more latency than the replica offloads.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	
db-internals/cf4d22fe	db-internals	easy	database, read-replicas, conflicts	What are recovery conflicts on a PostgreSQL replica and how do you handle them?	Recovery conflicts occur when a query running on a replica blocks WAL replay (e.g., the primary vacuums a row the replica query is reading). PostgreSQL cancels the query after max_standby_streaming_delay (default 30s). Handle by increasing the delay, using hot_standby_feedback = on (tells primary not to vacuum rows replicas need), or accepting occasional query cancellations.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	
db-internals/a1b2c3d4e5f0	db-internals	easy	database, replication, wal	How does PostgreSQL stream changes from a primary to replicas?	Via WAL (Write-Ahead Log) shipping. Every change is written to WAL first, then streamed to replicas for replay.\n\nRemember: database internals knowledge is what separates a developer who uses databases from an engineer who operates them. Understanding the storage engine prevents mysterious performance issues.\n\nGotcha: always test database configuration changes on a non-production replica first. A bad setting can crash the database or corrupt data.	training/library/topics/database-internals/primer.md
db-internals/b2c3d4e5f0a1	db-internals	easy	database, replication, async	What is the key difference between synchronous and asynchronous replication?	In synchronous replication, the primary waits for at least one replica to confirm the write before committing. In asynchronous, the primary commits immediately without waiting, which is faster but risks data loss on failover.\n\nRemember: streaming replication sends WAL records to replicas in real-time. Synchronous = guaranteed consistency, higher latency. Asynchronous = faster, risk of data loss on failover.	training/library/topics/database-internals/primer.md
db-internals/c3d4e5f0a1b2	db-internals	easy	database, replication, monitoring	What PostgreSQL system view shows the current replication status and lag?	pg_stat_replication. It shows client_addr, state, sent_lsn, write_lsn, flush_lsn, replay_lsn, and can be used to calculate replication lag in bytes.\n\nRemember: streaming replication sends WAL records to replicas in real-time. Synchronous = guaranteed consistency, higher latency. Asynchronous = faster, risk of data loss on failover.	training/library/topics/database-internals/primer.md
db-internals/d4e5f0a1b2c3	db-internals	medium	database, replication, lag	Name three common causes of replication lag in PostgreSQL.	Under-provisioned replica (CPU/IO), long-running queries on the replica blocking WAL replay, network saturation between primary and replica, and large transactions generating excessive WAL.\n\nRemember: streaming replication sends WAL records to replicas in real-time. Synchronous = guaranteed consistency, higher latency. Asynchronous = faster, risk of data loss on failover.	training/library/topics/database-internals/primer.md
db-internals/e5f0a1b2c3d4	db-internals	medium	database, replication, failover	How do you promote a PostgreSQL standby to primary (two methods)?	Using pg_ctl promote -D /path/to/data on the command line, or via SQL with SELECT pg_promote(). Both convert the standby from recovery mode to a read-write primary.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/f0a1b2c3d4e5	db-internals	medium	database, replication, split-brain	What is split-brain in database replication and why is it dangerous?	Split-brain occurs when two nodes both believe they are the primary and accept writes simultaneously. This causes data divergence that is extremely difficult to reconcile, potentially leading to data loss or corruption.\n\nRemember: streaming replication sends WAL records to replicas in real-time. Synchronous = guaranteed consistency, higher latency. Asynchronous = faster, risk of data loss on failover.	training/library/topics/database-internals/primer.md
db-internals/a1c2e3f4b5d6	db-internals	medium	database, replication, consensus	Name three strategies for preventing split-brain in database clusters.	Use a consensus mechanism (e.g., Patroni with etcd/ZooKeeper/Consul), implement fencing/STONITH to power off the old primary, and require quorum-based decisions where a majority must agree before promotion.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/b2d3f4a5c6e7	db-internals	hard	database, replication, mysql	How does MySQL semi-synchronous replication differ from PostgreSQL synchronous replication?	MySQL semi-sync (rpl_semi_sync_master_enabled) waits for at least one replica to acknowledge receipt of the binary log event, but not necessarily its replay. PostgreSQL synchronous replication can be configured to wait for write, flush, or replay (remote_apply) on the replica, offering more granular control.\n\nRemember: streaming replication sends WAL records to replicas in real-time. Synchronous = guaranteed consistency, higher latency. Asynchronous = faster, risk of data loss on failover.	training/library/topics/database-internals/primer.md
db-internals/c3e4a5b6d7f8	db-internals	hard	database, replication, failover, tools	What is Patroni and how does it prevent split-brain during PostgreSQL failover?	Patroni is an automated failover manager for PostgreSQL that uses a distributed consensus store (etcd, ZooKeeper, or Consul) to elect a leader. Only the node holding the leader key in the consensus store can operate as primary, preventing split-brain because consensus stores guarantee only one leader at a time.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/d4f5b6c7e8a9	db-internals	hard	database, replication, wal, lag	How would you calculate replication lag in seconds on a PostgreSQL replica, and what is the limitation of this approach?	Use: SELECT EXTRACT(EPOCH FROM now() - pg_last_xact_replay_timestamp()) AS lag_seconds. The limitation is that if no writes are happening on the primary, the replay timestamp stays stale, making lag appear high even though the replica is fully caught up. Check pg_last_wal_receive_lsn() = pg_last_wal_replay_lsn() first to handle this edge case.\n\nRemember: streaming replication sends WAL records to replicas in real-time. Synchronous = guaranteed consistency, higher latency. Asynchronous = faster, risk of data loss on failover.	training/library/topics/database-internals/primer.md
db-internals/e5a6b7c8d9f0	db-internals	medium	database, replication, sync, async, tradeoffs	What are the practical trade-offs between synchronous and asynchronous replication?	Synchronous: zero data loss on failover, but higher write latency (every commit waits for replica ACK) and reduced availability (if the sync replica goes down, writes block unless you configure synchronous_standby_names with multiple candidates).\nAsynchronous: lower write latency and no availability impact from replica failures, but potential data loss on failover equal to the replication lag at crash time.\n\nRemember: streaming replication sends WAL records to replicas in real-time. Synchronous = guaranteed consistency, higher latency. Asynchronous = faster, risk of data loss on failover.	
db-internals/f6b7c8d9e0a1	db-internals	medium	database, replication, wal-shipping, logical	What is the difference between WAL shipping (physical) and logical replication?	WAL shipping sends raw WAL bytes to replicas that replay them identically — the replica is a byte-for-byte copy. Logical replication decodes WAL into logical change events (INSERT, UPDATE, DELETE) and applies them, allowing selective table replication, cross-version replication, and different indexes or schemas on the subscriber. Logical replication has higher overhead but more flexibility.\n\nRemember: WAL = Write-Ahead Log. All changes are written to the log BEFORE being applied to data files. Ensures crash recovery: replay the WAL to recover uncommitted changes.	
db-internals/a7c8d9e0f1b2	db-internals	hard	database, replication, split-brain, prevention	How do fencing and STONITH prevent split-brain after a failover?	Fencing isolates the old primary so it cannot accept writes. STONITH (Shoot The Other Node In The Head) powers off or reboots the old primary at the hardware level. Without fencing, a network partition could leave the old primary still accepting writes while the new primary also accepts writes, causing data divergence that is nearly impossible to reconcile automatically.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	
db-internals/b8d9e0f1a2c3	db-internals	medium	database, replication, slots	What are replication slots and why do they matter for WAL retention?	A replication slot tells the primary to retain WAL segments until the connected replica has consumed them, preventing the primary from recycling WAL that the replica still needs. Without slots, a slow or disconnected replica may fall too far behind and require a full base backup to resync. The risk: a dead replica with an active slot causes unbounded WAL accumulation on the primary, filling the disk.\n\nRemember: WAL = Write-Ahead Log. All changes are written to the log BEFORE being applied to data files. Ensures crash recovery: replay the WAL to recover uncommitted changes.	
db-internals/c9e0f1a2b3d4	db-internals	hard	database, replication, multi-master, conflicts	What are the main challenges of multi-master (multi-primary) replication?	Write conflicts: two nodes may update the same row simultaneously, requiring conflict resolution rules (last-write-wins, application-defined merge, or manual resolution). Increased complexity in schema changes (DDL must be coordinated). Higher replication overhead. Most PostgreSQL deployments avoid multi-master; tools like BDR (Bi-Directional Replication) handle it but add operational complexity.\n\nRemember: streaming replication sends WAL records to replicas in real-time. Synchronous = guaranteed consistency, higher latency. Asynchronous = faster, risk of data loss on failover.	
db-internals/d0f1a2b3c4e5	db-internals	medium	database, replication, streaming, setup	What are the key steps to set up streaming replication in PostgreSQL?	1. On the primary: set wal_level = replica, max_wal_senders >= number of replicas, create a replication user.\n2. Take a base backup: pg_basebackup -h primary -D /data -U replicator -Fp -Xs -P.\n3. On the replica: configure primary_conninfo in postgresql.conf (or recovery.conf for PG < 12) and create standby.signal.\n4. Start the replica — it connects and streams WAL continuously.\n\nRemember: streaming replication sends WAL records to replicas in real-time. Synchronous = guaranteed consistency, higher latency. Asynchronous = faster, risk of data loss on failover.	
db-internals/e1a2b3c4d5f6	db-internals	easy	database, replication, monitoring, lag	What is the simplest way to monitor replication lag from the primary side?	Query pg_stat_replication on the primary: SELECT client_addr, state, replay_lag FROM pg_stat_replication; The replay_lag column (PG 10+) shows the time since the last WAL replayed on each replica. Alert if replay_lag exceeds your SLA threshold (e.g., > 5 seconds for warning, > 30 seconds for critical).\n\nRemember: streaming replication sends WAL records to replicas in real-time. Synchronous = guaranteed consistency, higher latency. Asynchronous = faster, risk of data loss on failover.	
db-internals/f2b3c4d5e6a7	db-internals	hard	database, replication, logical, use-cases	Name three use cases where logical replication is preferred over physical streaming replication.	1. Replicating a subset of tables (e.g., sharing only the orders table with an analytics database).\n2. Replicating between different PostgreSQL major versions during a rolling upgrade.\n3. Replicating to a database with different indexes, triggers, or additional columns (e.g., a search-optimized replica with extra GIN indexes).\n\nRemember: streaming replication sends WAL records to replicas in real-time. Synchronous = guaranteed consistency, higher latency. Asynchronous = faster, risk of data loss on failover.	
db-internals/a1f2b3c4d5e6	db-internals	easy	database, transactions, acid	What do the four letters in ACID stand for and what does each mean?	Atomicity (all or nothing), Consistency (data moves between valid states), Isolation (concurrent transactions don't interfere), Durability (committed data survives crashes).\n\nRemember: ACID = Atomicity (all or nothing), Consistency (valid state), Isolation (concurrent transactions don't interfere), Durability (committed data survives crashes). Mnemonic: 'All Changes Isolated Durably.'	training/library/topics/database-internals/primer.md
db-internals/b2a3c4d5e6f7	db-internals	easy	database, transactions, isolation	What is the default isolation level in PostgreSQL?	Read Committed. This means a transaction can only see data committed before each statement (no dirty reads), but may see different results for the same query if another transaction commits between statements (non-repeatable reads).\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/c3b4d5e6f7a8	db-internals	easy	database, transactions, timeout	How do you set a statement timeout in PostgreSQL to prevent runaway queries?	SET statement_timeout = '30s'; for per-session, or configure statement_timeout in postgresql.conf (in milliseconds) for a global default.\n\nRemember: database internals knowledge is what separates a developer who uses databases from an engineer who operates them. Understanding the storage engine prevents mysterious performance issues.\n\nGotcha: always test database configuration changes on a non-production replica first. A bad setting can crash the database or corrupt data.	training/library/topics/database-internals/primer.md
db-internals/d4c5e6f7a8b9	db-internals	medium	database, transactions, isolation	List the four SQL isolation levels from weakest to strongest.	Read Uncommitted (allows dirty reads), Read Committed (no dirty reads), Repeatable Read (no non-repeatable reads), Serializable (no phantom reads — full isolation). PostgreSQL treats Read Uncommitted as Read Committed.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/e5d6f7a8b9c0	db-internals	medium	database, transactions, deadlock	What happens when PostgreSQL detects a deadlock?	PostgreSQL automatically detects deadlocks (checking every deadlock_timeout interval, default 1s) and terminates the youngest transaction involved, allowing the other to proceed. The terminated transaction receives an error that the application should handle with a retry.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/f6e7a8b9c0d1	db-internals	medium	database, transactions, long-running	How do you find transactions running longer than 5 minutes in PostgreSQL?	SELECT pid, now() - xact_start AS duration, query, state FROM pg_stat_activity WHERE state != 'idle' AND xact_start < now() - interval '5 minutes' ORDER BY duration DESC;\n\nRemember: transaction isolation levels (low to high): Read Uncommitted, Read Committed, Repeatable Read, Serializable. Higher isolation = fewer anomalies but more locking/overhead.	training/library/topics/database-internals/primer.md
db-internals/a7f8b9c0d1e2	db-internals	medium	database, transactions, long-running	Why are long-running transactions dangerous in PostgreSQL?	They hold locks that block other transactions (including DDL like ALTER TABLE), prevent VACUUM from reclaiming dead tuples (causing table and index bloat), increase MVCC overhead, and can exhaust connection pool resources.\n\nRemember: transaction isolation levels (low to high): Read Uncommitted, Read Committed, Repeatable Read, Serializable. Higher isolation = fewer anomalies but more locking/overhead.	training/library/topics/database-internals/primer.md
db-internals/b8a9c0d1e2f3	db-internals	hard	database, transactions, mysql	How does MySQL's default isolation level differ from PostgreSQL's, and what is the practical impact?	MySQL defaults to REPEATABLE READ, while PostgreSQL defaults to READ COMMITTED. MySQL uses gap locks at Repeatable Read to prevent phantom reads, which can increase lock contention. PostgreSQL's MVCC-based Repeatable Read avoids extra locking but may produce serialization failures that require application retries.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/c9b0d1e2f3a4	db-internals	hard	database, transactions, serializable	When would you use SERIALIZABLE isolation and what is the operational cost?	Use Serializable for transactions that must appear to execute sequentially (e.g., financial transfers, inventory reservations). The cost is higher abort rates due to serialization conflicts, increased CPU overhead from dependency tracking, and the need for application-level retry logic. PostgreSQL implements it via Serializable Snapshot Isolation (SSI).\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/d0c1e2f3a4b5	db-internals	hard	database, transactions, terminate	How do you safely terminate a specific backend in PostgreSQL, and what is the difference between pg_cancel_backend and pg_terminate_backend?	pg_cancel_backend(pid) cancels the current query but keeps the connection alive (equivalent to sending SIGINT). pg_terminate_backend(pid) kills the entire backend process and closes the connection (equivalent to SIGTERM). Use cancel first for a gentle stop; use terminate when the backend is stuck or unresponsive.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	training/library/topics/database-internals/primer.md
db-internals/e1d2f3a4b5c6	db-internals	easy	database, transactions, acid, durability	What does Durability mean in practice and how does PostgreSQL guarantee it?	Durability means once a transaction is committed, the data survives even if the server crashes immediately after. PostgreSQL guarantees this by writing all changes to the WAL (Write-Ahead Log) and fsyncing the WAL to disk before returning the commit acknowledgment. On recovery, any committed but unwritten data pages are replayed from WAL.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	
db-internals/f2e3a4b5c6d7	db-internals	medium	database, transactions, isolation, read-committed, serializable	What is the practical difference between Read Committed and Serializable isolation?	Read Committed sees only data committed before each individual statement — two identical SELECTs in the same transaction can return different results if another transaction commits between them. Serializable guarantees the transaction sees a frozen snapshot and behaves as if all transactions ran one at a time. Serializable detects conflicts and aborts one transaction, requiring retry logic.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	
db-internals/a3f4b5c6d7e8	db-internals	medium	database, transactions, phantom-reads, dirty-reads	What are dirty reads, non-repeatable reads, and phantom reads?	Dirty read: reading uncommitted data from another transaction (not possible in PostgreSQL). Non-repeatable read: re-reading a row and getting different values because another transaction committed an update. Phantom read: re-running a query and getting different rows because another transaction committed an insert or delete matching the WHERE clause.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	
db-internals/b4a5c6d7e8f9	db-internals	medium	database, transactions, savepoints	What are savepoints and how do they help inside a transaction?	A savepoint marks a point within a transaction you can roll back to without aborting the entire transaction: SAVEPOINT sp1; ... ROLLBACK TO sp1; This is useful for retrying a portion of work (e.g., an insert that might violate a unique constraint) while keeping earlier work in the same transaction intact.\n\nRemember: transaction isolation levels (low to high): Read Uncommitted, Read Committed, Repeatable Read, Serializable. Higher isolation = fewer anomalies but more locking/overhead.	
db-internals/c5b6d7e8f9a0	db-internals	hard	database, transactions, two-phase-commit	What is two-phase commit (2PC) and when is it needed?	2PC coordinates a transaction across multiple databases or services. Phase 1 (prepare): all participants confirm they can commit. Phase 2 (commit): the coordinator tells all to commit. If any participant fails to prepare, all abort. In PostgreSQL, use PREPARE TRANSACTION 'id' and COMMIT PREPARED 'id'. Needed for distributed transactions, but adds latency and complexity — prefer saga patterns where possible.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	
db-internals/d6c7e8f9a0b1	db-internals	hard	database, transactions, xid-wraparound	What is transaction ID wraparound in PostgreSQL and why is it dangerous?	PostgreSQL uses 32-bit transaction IDs (XIDs), which wrap around after ~4 billion transactions. If VACUUM does not freeze old row versions in time, the database must shut down to prevent data loss (it would interpret old committed rows as being in the future). Monitor with: SELECT datname, age(datfrozenxid) FROM pg_database; Alert when age exceeds 500 million.\n\nRemember: transaction isolation levels (low to high): Read Uncommitted, Read Committed, Repeatable Read, Serializable. Higher isolation = fewer anomalies but more locking/overhead.	
db-internals/e7d8f9a0b1c2	db-internals	medium	database, transactions, connection-pool	How should connection pools handle transactions?	Connection pools (PgBouncer, HikariCP) must ensure a transaction runs entirely on a single connection — mid-transaction connection switching corrupts state. In PgBouncer, use transaction pooling mode (not statement mode) and avoid SET commands that persist beyond the transaction. Always close transactions promptly to return connections to the pool; idle-in-transaction connections waste pool capacity.\n\nRemember: transaction isolation levels (low to high): Read Uncommitted, Read Committed, Repeatable Read, Serializable. Higher isolation = fewer anomalies but more locking/overhead.\n\nRemember: each database connection consumes memory (~5-10 MB in PostgreSQL). Use connection pooling (PgBouncer, ProxySQL) to serve thousands of app connections with dozens of DB connections.	
db-internals/f8e9a0b1c2d3	db-internals	easy	database, transactions, begin-commit	What happens if you run statements without BEGIN in PostgreSQL?	PostgreSQL runs each statement in an implicit auto-commit transaction — the statement is automatically committed on success or rolled back on failure. Wrapping multiple statements in BEGIN...COMMIT groups them into a single atomic transaction where either all succeed or all are rolled back together.\n\nUnder the hood: understanding database internals (page layout, WAL, MVCC, buffer pool) helps you diagnose performance issues that high-level monitoring alone cannot explain.	

<!-- wiki:related:start -->
---

## Wiki Navigation

### Related Content

- [Database Internals](../../../../library/topics/database-internals/index.md) (Topic Pack, L1) — Database Internals

<!-- wiki:related:end -->
