---
tags:
- devops
- l1
- flashcard-deck
- sql
---
<!-- wiki:breadcrumb:start -->
[Portal](../../../../library/portal/index.md) | **Level:** [L1: Foundations](../../../../library/portal/levels.md) | **Topics:** [SQL](../../../../library/portal/topics.md) | **Domain:** DevOps & Tooling
<!-- wiki:breadcrumb:end -->

id	category	difficulty	tags	question	answer	source_path
databases/04e588929901	databases	hard	databases, performance	You find out your database became a bottleneck and users experience issues accessing data. How can you deal with such situation?	Not much information provided as to why it became a bottleneck and what is current architecture, so one general approach could be \nto reduce the load on your database by moving frequently-accessed data to in-memory structure.\n\nRemember: database fundamentals (ACID, indexing, normalization, query optimization) transfer across all database systems. Master the concepts, then learn vendor-specific syntax.	projects/knowledge/interview/databases/005-you-find-out-your-database-became-a-bottleneck-and.txt
databases/0d309df63755	databases	medium	databases, sql	What is ORM? What benefits it provides in regards to relational databases usage?	"[Wikipedia](https://en.wikipedia.org/wiki/Object%E2%80%93relational_mapping): ""is a programming technique for converting data between incompatible type systems using object-oriented programming languages""\n\nIn regards to the relational databases:\n\n  * Database as code\n  * Database abstraction\n  * Encapsulates SQL complexity\n  * Enables code review process\n  * Enables usage as a native OOP structure\n\nRemember: database fundamentals (ACID, indexing, normalization, query optimization) transfer across all database systems. Master the concepts, then learn vendor-specific syntax."	projects/knowledge/interview/databases/014-what-is-orm-what-benefits-it-provides-in-regards-t.txt
databases/27eabe1c1021	databases	easy	databases, query-optimization	What is database sharding and when is it needed?	Sharding is a horizontal partitioning — splitting data across multiple database instances, each holding a subset of rows. It is needed when a single database server cannot handle the write throughput, storage volume, or query load.\n\nRemember: sharding = horizontal partitioning across multiple database instances. Each shard holds a subset of data. Key challenge: choosing the shard key to avoid hotspots.\n\nGotcha: cross-shard JOINs and transactions become very expensive or impossible. Design your shard key so most queries hit a single shard.\n\nExample: shard by user_id — all of user 42's data lives on one shard. Queries for user 42 hit one database, not all of them.\n\nRemember: sharding = horizontal partitioning across multiple database instances. Each shard holds a subset of data. Key challenge: choosing the shard key to avoid hotspots.	projects/knowledge/interview/databases/004-what-is-sharding.txt
databases/2a385176e8e3	databases	medium	databases, sql, replication	What types of data tables have you used?	* Primary data table: main data you care about\n  * Details table: includes a foreign key and has one to many relationship\n  * Lookup values table: can be one table per lookup or a table containing all the lookups and has one to many relationship\n  * Multi reference table\n\nRemember: database fundamentals (ACID, indexing, normalization, query optimization) transfer across all database systems. Master the concepts, then learn vendor-specific syntax.	projects/knowledge/interview/databases/013-what-types-of-data-tables-have-you-used.txt
databases/2d9b60ef32c0	databases	hard	databases, sql, nosql, query-optimization	What does it mean when a database is ACID compliant?	ACID ensures reliable database transactions:\n\n- **Atomicity**: Transactions succeed or fail as a whole — no partial updates\n- **Consistency**: Every transaction moves the DB from one valid state to another, enforced by constraints\n- **Isolation**: Concurrent transactions don't see each other's intermediate states\n- **Durability**: Committed data survives crashes (written to non-volatile storage)\n\nSQL databases are ACID by nature. Some NoSQL databases (e.g., MongoDB 4.0+) support ACID transactions, but most NoSQL systems trade ACID guarantees for performance and scalability.\n\nRemember: ACID = Atomicity (all or nothing), Consistency (valid state), Isolation (concurrent transactions don't interfere), Durability (committed data survives crashes). Mnemonic: 'All Changes Isolated Durably.'	projects/knowledge/interview/databases/003-what-does-it-mean-when-a-database-is-acid-complian.txt
databases/37139196318c	databases	medium	databases, sql, schema, query-optimization	Explain Normalization	Data that is used multiple times in a database should be stored once and referenced with a foreign key. \nThis has the clear benefit of ease of maintenance where you need to change a value only in a single place to change it everywhere.\n\nRemember: normalization forms: 1NF (atomic values), 2NF (no partial dependencies), 3NF (no transitive dependencies). Mnemonic: 'The key, the whole key, and nothing but the key.'	projects/knowledge/interview/databases/011-explain-normalization.txt
databases/4091bf6fd9c9	databases	easy	databases, connection-pooling	What is a connection leak?	A connection leak is a situation where database connection isn't closed after being created and is no longer needed.\n\nDebug clue: monitor connection count in pg_stat_activity (PostgreSQL) or SHOW PROCESSLIST (MySQL). A steadily growing number of idle connections is a leak.\n\nGotcha: connection leaks eventually exhaust max_connections, causing new connections to fail with 'too many connections' errors — often during peak traffic.\n\nExample: in Python, always use context managers: with conn.cursor() as cur: ... or connection pool libraries that auto-reclaim connections.\n\nRemember: choose your database based on access patterns, not brand loyalty. Relational for transactions, document for flexible schemas, graph for relationships, time-series for metrics.\n\nGotcha: 'NoSQL' does not mean 'no schema' — it means schema is enforced at the application level instead of the database level. This shifts responsibility to developers.	projects/knowledge/interview/databases/007-what-is-a-connection-leak.txt
databases/4572d648676a	databases	medium	databases, sql, replication, query-optimization	Explain Primary Key and Foreign Key	Primary Key: each row in every table should a unique identifier that represents the row. \nForeign Key: a reference to another table's primary key. This allows you to join table together to retrieve all the information you need without duplicating data.\n\nRemember: database fundamentals (ACID, indexing, normalization, query optimization) transfer across all database systems. Master the concepts, then learn vendor-specific syntax.	projects/knowledge/interview/databases/012-explain-primary-key-and-foreign-key.txt
databases/47f79688ae5f	databases	easy	databases, performance, connection-pooling, query-optimization	What is a connection pool?	Connection Pool is a cache of database connections and the reason it's used is to avoid an overhead of establishing a connection for every query done to a database.\n\nNumber anchor: a PostgreSQL connection costs ~10MB of RAM. Without pooling, 100 microservice replicas each opening 10 connections = 1000 connections = 10GB of PostgreSQL memory just for connections.\n\nExample: PgBouncer (PostgreSQL) and ProxySQL (MySQL) are popular external connection poolers that multiplex thousands of app connections over a few dozen database connections.\n\nRemember: database fundamentals (ACID, indexing, normalization, query optimization) transfer across all database systems. Master the concepts, then learn vendor-specific syntax.	projects/knowledge/interview/databases/006-what-is-a-connection-pool.txt
databases/4f6dd8406715	databases	easy	databases, sql, performance	What is a Data Warehouse?	A data warehouse is a subject-oriented, integrated, time-variant and non-volatile collection of data in support of organisation's decision-making process\n\nRemember: choose your database based on access patterns, not brand loyalty. Relational for transactions, document for flexible schemas, graph for relationships, time-series for metrics.\n\nGotcha: 'NoSQL' does not mean 'no schema' — it means schema is enforced at the application level instead of the database level. This shifts responsibility to developers.	projects/knowledge/interview/databases/009-what-is-a-data-warehouse.txt
databases/5ec2e25664fa	databases	hard	databases, performance, sql, connection-pooling	Your database performs slowly than usual. More specifically, your queries are taking a lot of time. What would you do?	* Query for running queries and cancel the irrelevant queries\n* Check for connection leaks (query for running connections and include their IP)\n* Check for table locks and kill irrelevant locking sessions\n\nRemember: database fundamentals (ACID, indexing, normalization, query optimization) transfer across all database systems. Master the concepts, then learn vendor-specific syntax.	projects/knowledge/interview/databases/008-your-database-performs-slowly-than-usual-more-spec.txt
databases/67ff237e167b	databases	easy	databases, sql, schema	What is a relational database?	* Data Storage: system to store data in tables\n  * SQL: programming language to manage relational databases\n  * Data Definition Language: a standard syntax to create, alter and delete tables\n\nName origin: 'Relational' comes from E.F. Codd's 1970 paper introducing the relational model, where data is organized into 'relations' (tables).\n\nExample: popular relational databases — PostgreSQL (open source, feature-rich), MySQL/MariaDB (web-scale), SQLite (embedded), Oracle (enterprise), SQL Server (Microsoft).\n\nRemember: database fundamentals (ACID, indexing, normalization, query optimization) transfer across all database systems. Master the concepts, then learn vendor-specific syntax.	projects/knowledge/interview/databases/002-what-is-a-relational-database.txt
databases/821826791252	databases	easy	databases, sql, indexing	What is an index in a database?	A database index is a data structure that improves the speed of operations in a table. Indexes can be created using one or more columns, providing the basis for both rapid random lookups and efficient ordering of access to records.\n\nRemember: an index is a data structure (usually B-tree) that speeds up lookups at the cost of slower writes and extra storage. Like a book's index — you don't read every page to find a topic.	projects/knowledge/interview/databases/010-what-is-an-index-in-a-database.txt
databases/bc211f25daa3	databases	easy	databases, sql, nosql	What type of databases are you familiar with?	Relational (SQL)\nNoSQL\nTime series\n\nRemember: the major categories — Relational (PostgreSQL, MySQL), Document (MongoDB), Key-Value (Redis, DynamoDB), Column-family (Cassandra), Graph (Neo4j), Time-series (InfluxDB, TimescaleDB).\n\nInterview tip: mention specific databases you have used and why — e.g., 'PostgreSQL for ACID transactions, Redis for caching, TimescaleDB for metrics.'\n\nRemember: choose your database based on access patterns, not brand loyalty. Relational for transactions, document for flexible schemas, graph for relationships, time-series for metrics.\n\nGotcha: 'NoSQL' does not mean 'no schema' — it means schema is enforced at the application level instead of the database level. This shifts responsibility to developers.	projects/knowledge/interview/databases/001-what-type-of-databases-are-you-familiar-with.txt
databases/ccccfc4edfdd	databases	easy	databases, sql, schema	What is DDL (Data Definition Language) in SQL?	"[Wikipedia](https://en.wikipedia.org/wiki/Data_definition_language): ""In the context of SQL, data definition or data description language (DDL) is a syntax for creating and modifying database objects such as tables, indices, and users.""\n\nRemember: database fundamentals (ACID, indexing, normalization, query optimization) transfer across all database systems. Master the concepts, then learn vendor-specific syntax."	projects/knowledge/interview/databases/015-what-is-ddl.txt
databases/fbe7e971be85	databases	medium	databases, sql, indexing, performance	What is Time Series database?	A database designed specifically for time series based data. Time series data consists of data points indexed by time, such as metrics, logs, and sensor readings.\n\nKey optimizations:\n\n* **Time-based partitioning**: Data is automatically partitioned by time intervals, enabling efficient queries over specific time ranges\n* **Compression**: Specialized compression algorithms that exploit the temporal nature of data (e.g., delta encoding, run-length encoding)\n\nRemember: database fundamentals (ACID, indexing, normalization, query optimization) transfer across all database systems. Master the concepts, then learn vendor-specific syntax.	projects/knowledge/interview/databases/016-what-is-time-series-database.txt
databases/8a37feb236bb	databases	medium	databases,replication	Explain database replication and when you would use it in production.	Database replication copies data from one database server (primary/master) to one or more replicas. Use cases:\n\n* Read scaling: route read queries to replicas to reduce primary load\n* High availability: promote a replica if the primary fails\n* Disaster recovery: maintain replicas in another region/AZ\n* Reporting: run heavy analytics queries against a replica\n\nReplication modes:\n* Synchronous: primary waits for replica acknowledgment (stronger consistency, higher latency)\n* Asynchronous: primary does not wait (lower latency, risk of data loss on failover)\n* Semi-synchronous: primary waits for at least one replica	projects/knowledge/interview/databases/017-database-replication.txt
databases/5e232f018099	databases	hard	databases,cap	What is the CAP theorem and how does it apply to database selection?	The CAP theorem states that a distributed system can guarantee at most two of three properties simultaneously:\n\n* Consistency: every read returns the most recent write\n* Availability: every request receives a response\n* Partition tolerance: the system continues operating despite network partitions\n\nIn practice, network partitions are unavoidable, so the trade-off is between consistency and availability during a partition.\nCP: etcd, ZooKeeper, HBase\nAP: Cassandra, DynamoDB (eventual consistency)\nCA: single-node PostgreSQL (no partition tolerance)\n\nRemember: CAP = Consistency, Availability, Partition tolerance. Pick 2 of 3 during a network partition. CP = consistent but may reject requests. AP = available but may return stale data.	projects/knowledge/interview/databases/018-cap-theorem.txt
databases/7d63495529b8	databases	medium	databases,backup	What strategies exist for database backup and recovery?	Common strategies:\n\n* Logical backups (pg_dump, mysqldump): SQL-level export, portable but slow for large DBs\n* Physical backups (pg_basebackup, xtrabackup): file-level copy, faster restore\n* WAL/binlog archiving: continuous archiving enables point-in-time recovery (PITR)\n* Snapshots: storage-level snapshots (EBS, ZFS) for near-instant backup\n\nBest practices:\n* Test restores regularly\n* Store backups off-site or in a different region\n* Monitor backup job success/failure\n* Document RTO and RPO targets\n\nRemember: database fundamentals (ACID, indexing, normalization, query optimization) transfer across all database systems. Master the concepts, then learn vendor-specific syntax.	projects/knowledge/interview/databases/019-backup-recovery.txt
databases/841c2e224edb	databases	hard	databases,troubleshooting	How do you diagnose and fix a slow query in PostgreSQL?	Diagnostic steps:\n\n1. Enable slow query logging: set log_min_duration_statement (e.g., 500ms)\n2. Use EXPLAIN ANALYZE to see the actual execution plan\n3. Look for sequential scans on large tables, nested loops, inaccurate row estimates\n4. Check pg_stat_statements for aggregate query performance\n\nCommon fixes:\n* Add missing indexes (check pg_stat_user_tables for seq_scan counts)\n* Rewrite queries to avoid unnecessary joins or subqueries\n* Run ANALYZE to update planner statistics\n* Tune work_mem, shared_buffers, effective_cache_size\n* Consider partitioning for very large tables\n\nRemember: database fundamentals (ACID, indexing, normalization, query optimization) transfer across all database systems. Master the concepts, then learn vendor-specific syntax.	projects/knowledge/interview/databases/020-slow-query-postgres.txt
databases/772fe0664989	databases	medium	databases,migrations	How should database schema migrations be handled in a DevOps workflow?	Best practices:\n\n* Use a migration tool (Flyway, Alembic, Liquibase) to version schema changes\n* Store migrations in version control alongside application code\n* Make migrations idempotent and reversible when possible\n* Run migrations as part of CI/CD pipeline, before application deployment\n* Avoid destructive changes in production (prefer additive: add column, backfill, remove old)\n* Test migrations against a copy of production data\n* Use transactions for DDL where supported (PostgreSQL does, MySQL mostly does not)\n\nRemember: database fundamentals (ACID, indexing, normalization, query optimization) transfer across all database systems. Master the concepts, then learn vendor-specific syntax.	projects/knowledge/interview/databases/021-schema-migrations.txt
databases/edbeb02a0625	databases	medium	databases,monitoring	What database metrics should you monitor in production?	Key metrics by category:\n\n* Connections: active connections, pool utilization, connection wait time\n* Performance: query latency (p50/p95/p99), queries per second, slow query count\n* Resources: CPU, memory/buffer cache hit ratio, disk I/O (IOPS, latency)\n* Replication: replica lag, replication slot status\n* Storage: table size growth, WAL/binlog size, disk space remaining\n* Locks: lock wait time, deadlock count, long-running transactions\n\nTools: pg_stat_statements, Performance Schema (MySQL), CloudWatch RDS, postgres_exporter\n\nRemember: database fundamentals (ACID, indexing, normalization, query optimization) transfer across all database systems. Master the concepts, then learn vendor-specific syntax.	projects/knowledge/interview/databases/022-database-monitoring.txt
databases/514d096adc3c	databases	hard	databases,ha	Describe a high-availability PostgreSQL setup and its failure modes.	Typical HA setup:\n\n* Primary + synchronous standby (or Patroni-managed cluster)\n* Connection pooler (PgBouncer) in front\n* Automated failover via Patroni, repmgr, or cloud-managed (RDS Multi-AZ)\n\nFailure modes:\n* Primary crash: standby promoted, clients reconnect (brief downtime)\n* Network partition: split-brain risk if fencing misconfigured, use quorum/witness\n* Replica lag: synchronous replication prevents data loss but adds latency\n* Connection pooler failure: run multiple instances behind a VIP or DNS\n* Storage failure: RAID or cloud volume replication protects against single-disk failure\n\nRemember: database fundamentals (ACID, indexing, normalization, query optimization) transfer across all database systems. Master the concepts, then learn vendor-specific syntax.	projects/knowledge/interview/databases/023-ha-postgres.txt
databases/746919d9dcc0	databases	easy	databases,nosql	When would you choose a NoSQL database over a relational database?	Choose NoSQL when:\n\n* Schema is highly variable or evolving rapidly (document stores like MongoDB)\n* You need horizontal write scaling (Cassandra, DynamoDB)\n* Data is naturally key-value or graph-shaped (Redis, Neo4j)\n* You need very low-latency lookups at massive scale\n* Eventual consistency is acceptable\n\nStick with relational when:\n* You need strong consistency and complex transactions (ACID)\n* Your data has clear relationships and you need JOINs\n* You need mature tooling for reporting and analytics\n\nRemember: SQL = structured, ACID, joins, schema. NoSQL = flexible schema, horizontal scaling, eventual consistency. Use SQL for transactions, NoSQL for scale + flexibility.	projects/knowledge/interview/databases/024-nosql-vs-relational.txt
databases/b8cdaf5a38e1	databases	hard	databases,transactions	Explain transaction isolation levels and when each is appropriate.	SQL standard defines four levels (weakest to strongest):\n\n* Read Uncommitted: can see uncommitted changes (dirty reads). Rarely used.\n* Read Committed: only sees committed data. Default in PostgreSQL. Good for most OLTP.\n* Repeatable Read: snapshot at transaction start. Default in MySQL InnoDB. Good for reporting.\n* Serializable: transactions behave as if sequential. Highest overhead. Use for financial calculations or inventory where correctness is critical.\n\nHigher isolation = more locking/overhead = lower throughput.\n\nRemember: transaction isolation levels (low to high): Read Uncommitted, Read Committed, Repeatable Read, Serializable. Higher isolation = fewer anomalies but more locking/overhead.	projects/knowledge/interview/databases/025-isolation-levels.txt

<!-- wiki:related:start -->
---

## Wiki Navigation

### Related Content

- SQL Flashcards *(CLI)* (flashcard_deck, L1) — SQL
- [SQL Fundamentals](../../../../library/topics/sql-fundamentals/index.md) (Topic Pack, L0) — SQL

<!-- wiki:related:end -->
