---
tags:
- devops
- l1
- flashcard-deck
- postmortem-slo
---
<!-- wiki:breadcrumb:start -->
[Portal](../../../../library/portal/index.md) | **Level:** [L1: Foundations](../../../../library/portal/levels.md) | **Topics:** [Postmortems & SLOs](../../../../library/portal/topics.md) | **Domain:** DevOps & Tooling
<!-- wiki:breadcrumb:end -->

id	category	difficulty	tags	question	answer	source_path
devops/001368aa8db6	devops	easy	automation, ci-cd, devops	What is a git commit and what information does it store?	* In Git, a commit is a snapshot of your repo at a specific point in time.\n* The git commit command will save all staged changes, along with a brief description from the user, in a “commit” to the local repository.	projects/knowledge/interview/devops/010-what-is-a-commit.txt
devops/012c4e38a38d	devops	medium	devops, version-control	What are some practical implementations or practices of GitOp?	* Store Infra files in a version control repository (like Git)\n* Apply review/approval process for changes\n\nExample: ArgoCD and Flux watch Git repos and auto-sync cluster state. Terraform configs and Helm values stored in Git — every change goes through code review.	projects/knowledge/interview/devops/038-what-are-some-practical-implementations-or-practic.txt
devops/043edf42c7ef	devops	hard	automation, ci-cd, culture, devops	How does a web server work?	\n    \nWe can understand web servers using two view points, which is:\n    \n    (i) Hardware (ii) Software\n\n(i)   A web server is nothing but a remote computer which stores website's component files(HTML,CSS and Javascript files) and web server's software.A web server connects to\n      the Internet and supports physical data interchange with other devices connected to the web.\n	projects/knowledge/interview/devops/024-how-does-a-web-server-work.txt
devops/0d387569db58	devops	easy	ci-cd, devops, version-control	What is a merge in git and how does it combine branches?	* Merging is Git's way of putting a forked history back together again. The git merge command lets you take the independent lines of development created by git branch and integrate them into a single branch.	projects/knowledge/interview/devops/011-what-is-a-merge.txt
devops/1116bbe265d7	devops	hard	devops, incident-response, operations, reliability	When should you NOT fix a production issue immediately?	"When the fix has a larger blast radius than the problem, or when the cause is unclear.\n\nDon't fix when:\n\n1. During partial outage with unclear cause\n   - 10% of users affected\n   - Root cause unknown\n   - ""Fix"" might affect other 90%\n   - Worse: mask the real problem\n\n2. When fix has larger blast radius\n   - Problem: one microservice degraded"\n\nRemember: cure must not be worse than disease. 10% affected + risky fix hitting the other 90% = wait, gather evidence, coordinate.	projects/knowledge/interview/devops/104-when-not-to-fix.txt
devops/199cafce1175	devops	medium	architecture, devops	Explain stateless vs. stateful	Stateless applications don't store any data in the host which makes it ideal for horizontal scaling and microservices.\nStateful applications depend on the storage to save state and data, typically databases are stateful applications.\n\nExample: REST API = stateless. PostgreSQL = stateful. K8s uses Deployments for stateless, StatefulSets for stateful workloads.	projects/knowledge/interview/devops/022-explain-stateless-vs-stateful.txt
devops/1b0beac30e26	devops	medium	ci-cd, devops, jenkins, monitoring	How to add a new worker node in Jenkins ?	Log into the Jenkins master and navigate to Manage Jenkins > Manage Nodes > New Node. Enter a name for the new node and select Permanent Agent. Configure SSH and click on Launch.\n\nGotcha: ensure agent has Java and network access to Jenkins master. Test with curl http://jenkins:8080 before configuring.	projects/knowledge/interview/devops/058-how-to-add-a-new-worker-node-in-jenkins.txt
devops/2b98db61cea3	devops	easy	architecture, devops	What is caching? How does it work? Why is it important?	Caching is fast access to frequently used resources which are computationally expensive or IO intensive and do not change often. There can be several layers of cache that can start from CPU caches to distributed cache systems. Common ones are in memory caching and distributed caching. Caches are typically data structures that contains some data, such as a hashtable or dictionary.\n\nRemember: cache layers — CPU L1/L2 (ns), RAM/Memcached (us), Redis (ms), CDN (ms), Disk (ms). Each trades freshness for speed.	projects/knowledge/interview/devops/021-what-is-caching-how-does-it-work-why-is-it-importa.txt
devops/2bbc8112c51b	devops	easy	blameless, devops, incident-response	What is the core value often put forward when talking about postmortem?	Blamelessness. \nPostmortems need to be blameless and this value should be remided at the beginning of every postmortem. This is the best way to ensure that people are playing the game to find the root cause and not trying to hide their possible faults.	projects/knowledge/interview/devops/049-what-is-the-core-value-often-put-forward-when-talk.txt
devops/2f9bb3e3cdaa	devops	easy	automation, devops, version-control	What is a merge conflict?	* A merge conflict is an event that occurs when Git is unable to automatically resolve differences in code between two commits. When all the changes in the code occur on different lines or in different files, Git will successfully merge commits without your help.\n\nGotcha: conflicts happen when the SAME lines change in two branches. Git marks them with <<<<<<< / ======= / >>>>>>> markers.	projects/knowledge/interview/devops/012-what-is-a-merge-conflict.txt
devops/302c017c969b	devops	easy	devops, monitoring, sre	What is the role of monitoring in SRE?	"Google: ""Monitoring is one of the primary means by which service owners keep track of a system’s health and availability""\n\nRead more about it [here](https://sre.google/sre-book/introduction)"\n\nRemember: Google's four golden signals — Latency, Traffic, Errors, Saturation (LTES). The minimum effective monitoring for any production service.	projects/knowledge/interview/devops/045-what-is-the-role-of-monitoring-in-sre.txt
devops/30c76f9e12b1	devops	medium	automation, ci-cd, devops, feedback-loops	What is shared modules in Jenkins ?	Shared modules in Jenkins refer to a collection of reusable code and resources that can be shared across multiple Jenkins jobs. This allows for easier maintenance, reduced duplication, and improved consistency across multiple build processes.\n\n\nExample: Git repo with vars/deployToK8s.groovy. Pipelines use @Library('mylib') to import. Eliminates copy-paste across Jenkinsfiles.	projects/knowledge/interview/devops/055-what-is-shared-modules-in-jenkins.txt
devops/35c625aa24ed	devops	medium	devops, sre	What are the two main SRE KPIs	Service Level Indicators (SLI) and Service Level Objectives (SLO).\n\nRemember: SLI measures, SLO targets, SLA contracts. Indicator → Objective → Agreement. SLI is the number, SLO is the goal, SLA has business consequences.	projects/knowledge/interview/devops/046-what-are-the-two-main-sre-kpis.txt
devops/3916aa4439e2	devops	medium	automation, ci-cd, devops, monitoring	Can you explain the CICD process in your current project ? or Can you talk about any CICD process that you have implemented ?	In the current project we use the following tools orchestrated with Jenkins to achieve CICD.\n   - Maven, Sonar, AppScan, ArgoCD, and Kubernetes\n   \n   Coming to the implementation, the entire process takes place in 8 steps\n    \n    1. Code Commit: Developers commit code changes to a Git repository hosted on GitHub.\n\n\nExample: Git commit → Build → Test → Scan (SonarQube/AppScan) → Artifact → Deploy staging → Integration test → Deploy prod (ArgoCD to K8s).	projects/knowledge/interview/devops/050-can-you-explain-the-cicd-process-in-your-current-p.txt
devops/3beb21536d56	devops	medium	ci-cd, culture, devops, sre	What are the benefits of DevOps? What can it help us to achieve?	* Collaboration\n  * Improved delivery\n  * Security\n  * Speed\n  * Scale\n  * Reliability\n\nRemember: CSSRS — Collaboration, Security, Speed, Reliability, Scale. DevOps = culture + automation + measurement + sharing (CAMS model).	projects/knowledge/interview/devops/002-what-are-the-benefits-of-devops-what-can-it-help-u.txt
devops/3e0f9145d718	devops	medium	ci-cd, devops, version-control	What are the anti-patterns of DevOps?	A couple of examples:\n\n* One person is in charge of specific tasks. For example there is only one person who is allowed to merge the code of everyone else into the repository.\n* Treating production differently from development environment. For example, not implementing security in development environment\n* Not allowing someone to push to production on Friday ;)\n\nGotcha: "Only Bob can deploy" = bus factor 1. "No Friday deploys" = deploy process too scary to trust. Both worth fixing.	projects/knowledge/interview/devops/003-what-are-the-anti-patterns-of-devops.txt
devops/418df89f3127	devops	medium	ci-cd, devops, testing	What is latest version of Jenkins or which version of Jenkins are you using ?	This is a very simple question interviewers ask to understand if you are actually using Jenkins day-to-day, so always be prepared for this.\n\nGotcha: interviewers ask this to check you actually USE Jenkins daily. Know your version — check bottom-right of Jenkins UI.	projects/knowledge/interview/devops/054-what-is-latest-version-of-jenkins-or-which-version.txt
devops/49fa78e92302	devops	easy	automation, ci-cd, culture, devops	What is DevOps and what problems does it solve?	"The definition of DevOps from selected companies:\n\n**Amazon**:\n\n""DevOps is the combination of cultural philosophies, practices, and tools that increases an organization’s ability to deliver applications and services at high velocity: evolving and improving products at a faster pace than organizations using traditional software development and infrastructure management processes."	projects/knowledge/interview/devops/001-what-is-devops.txt
devops/54a61dc0ce37	devops	medium	ci-cd, deployment, devops	"Would you prefer a ""configuration->deployment"" model or ""deployment->configuration""? Why?"	"Both have advantages and disadvantages.\nWith ""configuration->deployment"" model for example, where you build one image to be used by multiple deployments, there is less chance of deployments being different from one another, so it has a clear advantage of a consistent environment."\n\nExample: config→deploy = one image, many envs (12-factor). deploy→config = per-env images. Config→deploy is generally preferred for consistency.	projects/knowledge/interview/devops/014-would-you-prefer-a-configuration-deployment-model-.txt
devops/593e00a87e43	devops	medium	automation, devops, toil	What benefits does infrastructure-as-code have?	- fully automated process of provisioning, modifying and deleting your infrastructure\n- version control for your infrastructure which allows you to quickly rollback to previous versions\n- validate infrastructure quality and stability with automated tests and code reviews\n- makes infrastructure tasks less repetitive	projects/knowledge/interview/devops/028-what-benefits-does-infrastructure-as-code-have.txt
devops/5960ebb1b541	devops	easy	blameless, ci-cd, devops, incident-response	What is a postmortem ?	The postmortem is a process that should take place following an incident. It’s purpose is to identify the root cause of an incident and the actions that should be taken to avoid this kind of incidents from happening again.	projects/knowledge/interview/devops/048-what-is-a-postmortem.txt
devops/5b25c8867e73	devops	medium	ci-cd, culture, devops, feedback-loops	How would you describe a successful DevOps engineer or a team?	The answer can focus on:\n\n* Collaboration\n* Communication\n* Set up and improve workflows and processes (related to testing, delivery, ...)\n* Dealing with issues\n\nThings to think about:\n\n* What DevOps teams or engineers should NOT focus on or do?\n* Do DevOps teams or engineers have to be innovative or practice innovation as part of their role?\n\nRemember: great DevOps teams automate toil, share on-call, deploy frequently with confidence, treat incidents as learning opportunities.	projects/knowledge/interview/devops/004-how-would-you-describe-a-successful-devops-enginee.txt
devops/5be7b322e8d4	devops	medium	automation, ci-cd, devops	What are the different ways to trigger jenkins pipelines ?	"This can be done in multiple ways,\n   To briefly explain about the different options,\n   ```\n     - Poll SCM: Jenkins can periodically check the repository for changes and automatically build if changes are detected. \n                 This can be configured in the ""Build Triggers"" section of a job.\n"\n\nExample: Poll SCM, webhook, cron schedule, remote API (curl), upstream job trigger — five main Jenkins trigger mechanisms.	projects/knowledge/interview/devops/051-what-are-the-different-ways-to-trigger-jenkins-pip.txt
devops/5c3f61cfe188	devops	medium	automation, ci-cd, devops, monitoring	What's your philosophy on automation?	Automate the boring, repeatable, and dangerous. Not everything.\n\n**Good candidates for automation**:\n* Repetitive tasks (deployments, provisioning)\n* Error-prone manual steps\n* Dangerous operations (safer when scripted)\n* Toil that doesn't require judgment\n\n**Bad candidates for automation**:\n* One-off tasks (time to automate > time to do)\n\nRemember: automate the boring, repeatable, and dangerous. Not one-off tasks where automation time exceeds manual time.	projects/knowledge/interview/devops/063-whats-your-philosophy-on-automation.txt
devops/5e7170c7a437	devops	medium	ci-cd, devops	"What do you think about the following statement: ""100% is the only right availability target for a system"""	Wrong. No system can guarantee 100% availability as no system is safe from experiencing zero downtime.\nMany systems and services will fall somewhere between 99% and 100% uptime (or at least this is how most systems and services should be).\n\nFun fact: five nines (99.999%) = 5.26 min/year downtime. Each additional nine is exponentially harder and more expensive. 100% is theoretical impossibility.	projects/knowledge/interview/devops/043-what-do-you-think-about-the-following-statement-10.txt
devops/5ed568ce9ceb	devops	medium	devops, career, architecture	What's the biggest mistake senior engineers make?	"Over-engineering before constraints are real.\n\n**Symptoms**:\n* Building for scale you don't have\n* Abstract frameworks for single use cases\n* ""We might need this later""\n* Complexity without corresponding value\n* Not invented here syndrome\n\n**Why it happens**:"\n\nRemember: YAGNI — You Aren't Gonna Need It. Build for today's constraints, not imagined future scale. Premature abstraction is technical debt.	projects/knowledge/interview/devops/065-biggest-mistake-seniors-make.txt
devops/60c45f945d5c	devops	easy	ci-cd, culture, devops	What is Version Control?	* Version control is the system of tracking and managing changes to software code.\n* It helps software teams to manage changes to source code over time.\n* Version control also helps developers move faster and allows software teams to preserve efficiency and agility as the team scales to include more developers.	projects/knowledge/interview/devops/009-what-is-version-control.txt
devops/63e936c9cc60	devops	medium	blameless, career, devops	What's the biggest mistake junior engineers make?	"Changing things before understanding blast radius.\n\n**Symptoms**:\n* ""I'll just restart this service"" (in production)\n* Running commands from Stack Overflow without understanding\n* ""It worked in dev"" mentality\n* Not checking current state before changing it\n\n**The fix**:\n* Read before write - understand current state"\n\nRemember: "Read before write." Understand current state before changing it. Check blast radius before acting. Ask what happens if this fails.	projects/knowledge/interview/devops/064-biggest-mistake-juniors-make.txt
devops/672f551415ad	devops	hard	devops, career, engineering, leadership	What makes a principal/staff Linux engineer different from a senior engineer?	"Principal engineers anticipate failure modes and design systems that limit blast radius.\n\nKey differentiators:\n\n1. Anticipates failure modes\n   - Doesn't just fix problems, predicts them\n   - ""What happens when X fails?"" before X fails\n   - Designs for failure, not just success\n   - Thinks in failure domains and cascades\n\n2. Designs blast-radius limits\n   - Isolation between components"\n\nRemember: seniors fix problems. Principals prevent them with blast-radius limits, failure domains, and cascading failure protections.	projects/knowledge/interview/devops/103-principal-engineer-traits.txt
devops/6e56d0832afd	devops	hard	architecture, automation, devops, monitoring	Explain mutable vs. immutable infrastructure	In mutable infrastructure paradigm, changes are applied on top of the existing infrastructure and over time\nthe infrastructure builds up a history of changes. Ansible, Puppet and Chef are examples of tools which\nfollow mutable infrastructure paradigm.\n\n\nRemember: Mutable = update in place (Ansible). Immutable = replace entirely (Terraform/Packer). Immutable prevents snowflake servers that drift over time.	projects/knowledge/interview/devops/015-explain-mutable-vs-immutable-infrastructure.txt
devops/725daebe0753	devops	medium	automation, ci-cd, devops, monitoring	Explain Declarative and Procedural styles. The technologies you are familiar with (or using) are using procedural or declarative style?	Declarative - You write code that specifies the desired end state \nProcedural - You describe the steps to get to the desired end state\n\nDeclarative Tools - Terraform, Puppet, CloudFormation, Ansible \nProcedural Tools - Chef\n\nTo better emphasize the difference, consider creating two virtual instances/servers.\n\n\nRemember: Declarative = WHAT (Terraform, Puppet). Procedural = HOW (Chef). Declarative is naturally idempotent — run twice, same result.	projects/knowledge/interview/devops/033-explain-declarative-and-procedural-styles-the-tech.txt
devops/741e347db006	devops	medium	ci-cd, devops, jenkins	can you use Jenkins to build applications with multiple programming languages using different agents in different stages ?	Yes, Jenkins can be used to build applications with multiple programming languages by using different build agents in different stages of the build process.\n\nJenkins supports multiple build agents, which can be used to run build jobs on different platforms and with different configurations.	projects/knowledge/interview/devops/056-can-you-use-jenkins-to-build-applications-with-mul.txt
devops/74da0c621cfe	devops	medium	devops, kernel, text-processing	"Are you familiar with ""The Cathedral and the Bazaar models""? Explain each of the models"	* Cathedral - source code released when software is released\n* Bazaar - source code is always available publicly (e.g. Linux Kernel)\n\nName origin: Eric Raymond's 1997 essay. Cathedral = closed releases. Bazaar = open continuous development (Linux kernel model).	projects/knowledge/interview/devops/020-are-you-familiar-with-the-cathedral-and-the-bazaar.txt
devops/81fc58c18bde	devops	medium	ci-cd, devops, jenkins	How to add a new plugin in Jenkins ?	"Using the CLI, \n   `java -jar jenkins-cli.jar install-plugin <PLUGIN_NAME>`\n  \n  Using the UI,\n\n   1. Click on the ""Manage Jenkins"" link in the left-side menu.\n   2. Click on the ""Manage Plugins"" link."\n\nGotcha: test plugins in staging first. Plugin conflicts can break pipelines. Pin plugin versions in production.	projects/knowledge/interview/devops/059-how-to-add-a-new-plugin-in-jenkins.txt
devops/83602c6ac05f	devops	hard	ci-cd, devops, jenkins	What is JNLP and why is it used in Jenkins ?	"In Jenkins, JNLP is used to allow agents (also known as ""slave nodes"") to be launched and managed remotely by the Jenkins master instance. This allows Jenkins to distribute build tasks to multiple agents, providing scalability and improving performance.\n\n   When a Jenkins agent is launched using JNLP, it connects to the Jenkins master and receives build tasks, which it then executes. The results of the build are then sent back to the master and displayed in the Jenkins user interface."	projects/knowledge/interview/devops/060-what-is-jnlp-and-why-is-it-used-in-jenkins.txt
devops/84f97aea0b9e	devops	easy	automation, devops, sre, toil	What is Toil in the context of SRE and DevOps?	Google: Toil is the kind of work tied to running a production service that tends to be manual, repetitive, automatable, tactical, devoid of enduring value, and that scales linearly as a service grows\n\nRead more about it [here](https://sre.google/sre-book/eliminating-toil/)\n\nRemember: TOIL = Manual, Repetitive, Automatable, Tactical, No enduring value, Linearly scaling. If work checks all six boxes, automate it.	projects/knowledge/interview/devops/047-what-is-toil.txt
devops/8df45f7ccfe5	devops	hard	automation, devops, iac	How to deal with a configuration drift?	Configuration drift can be avoided with desired state configuration (DSC) implementation. Desired state configuration can be a declarative file that defined how a system should be. There are tools to enforce desired state such a terraform or azure dsc. There are incremental or complete strategies.	projects/knowledge/interview/devops/032-how-to-deal-with-a-configuration-drift.txt
devops/8ee1c9061966	devops	medium	ci-cd, culture, devops	Two engineers in your team argue on where to put the configuration and infra related files of a certain application. One of them suggests to put it in the same repo as the application repository and the other one suggests to put to put it in its own separate repository. What's your take on that?	One might say we need more details as to what these configuration and infra files look like exactly and how complex the application and its CI/CD pipeline(s), but in general, most of the time you will want to put configuration and infra related files in their own separate repository and not in the repository of the application for multiple reasons:\n\n	projects/knowledge/interview/devops/039-two-engineers-in-your-team-argue-on-where-to-put-t.txt
devops/8fce2d7515d8	devops	medium	ci-cd, devops, jenkins	What are some of the common plugins that you use in Jenkins ?	Be prepared for answer, you need to have atleast 3-4 on top of your head, so that interview feels you use jenkins on a day-to-day basis.\n\nExample: Pipeline, Git, Blue Ocean, Credentials Binding, Docker Pipeline, Kubernetes, Slack Notification. Name at least 4 confidently.	projects/knowledge/interview/devops/061-what-are-some-of-the-common-plugins-that-you-use-i.txt
devops/95a08d548d29	devops	hard	automation, ci-cd, devops, monitoring	How do you store/secure/handle secrets in Jenkins ?	Again, there are multiple ways to achieve this, \n   Let me give you a brief explanation of all the posible options.\n```  \n   - Credentials Plugin: Jenkins provides a credentials plugin that can be used to store secrets such as passwords, API keys, and certificates.\n\nGotcha: never echo secrets in logs. Use withCredentials block and external vaults (HashiCorp Vault, AWS Secrets Manager) for production.	projects/knowledge/interview/devops/053-how-do-you-storesecurehandle-secrets-in-jenkins.txt
devops/9732f0a4f708	devops	medium	ci-cd, deployment, devops	What deployment strategies are you familiar with or have used?	There are several deployment strategies:\n    * Rolling\n    * Blue green deployment\n    * Canary releases\n    * Recreate strategy\n\nRemember: RBCC — Rolling (gradual), Blue-Green (instant switch), Canary (small % test), Recreate (stop all, start new).	projects/knowledge/interview/devops/030-what-deployment-strategies-are-you-familiar-with-o.txt
devops/a3392f0caf65	devops	easy	ci-cd, culture, devops	"One of your team members suggests to set a goal of ""deploying at least 20 times a day"" in regards to CD. What is your take on that?"	"A couple of thoughts:\n\n1. Why is it an important goal? Is it affecting the business somehow? One of the KPIs? In other words, does it matters?\n2. This might introduce risks such as losing quality in favor of quantity\n3. You might want to set a possibly better goal such as ""be able to deploy whenever we need to deploy"""	projects/knowledge/interview/devops/005-one-of-your-team-members-suggests-to-set-a-goal-of.txt
devops/a528d31acadf	devops	easy	automation, devops, iac	"What is ""infrastructure as code""? What implementation of IAC are you familiar with?"	IAC (infrastructure as code) is a declarative approach of defining infrastructure or architecture of a system. Some implementations are ARM templates for Azure and Terraform that can work across multiple cloud providers.\n\nExample: Terraform (multi-cloud), CloudFormation (AWS), Pulumi (real code), Ansible (config + provisioning). IaC = reproducible, version-controlled infra.	projects/knowledge/interview/devops/027-what-is-infrastructure-as-code-what-implementation.txt
devops/a75490ea64a4	devops	hard	devops, philosophy, operations, reliability	What's the most dangerous assumption in Linux/infrastructure engineering?	"The system is doing what I think it is.\n\nThis assumption kills because:\n\n1. Storage tells the truth\n   - Drives report success, data not written\n   - RAID says healthy, silent corruption\n   - Backup job green, restore fails\n   - ""It said OK"" means nothing\n\n2. Monitoring sees everything\n   - Monitoring shows what you measure"\n\nRemember: "Trust but verify." The system does what you OBSERVE, not what you THINK. Every monitoring gap is an unverified assumption.	projects/knowledge/interview/devops/102-most-dangerous-assumption.txt
devops/aaff8e8874d5	devops	hard	devops, backup, snapshots, disaster-recovery	Why are snapshot-based backups dangerous?	"Snapshots capture crash-consistent state, not application-consistent state.\n\nThe illusion:\n- ""Snapshots are instant backups""\n- ""Point-in-time recovery""\n- ""Zero-downtime backups""\n\nThe reality:\n\n1. Crash-consistent vs app-consistent\n   - Crash-consistent: what disk looks like if power cut\n   - App-consistent: what disk looks like after clean shutdown"\n\nRemember: crash-consistent != application-consistent. Snapshot captures mid-write disk state. Always quiesce apps before snapshotting.	projects/knowledge/interview/devops/101-snapshot-backups-dangerous.txt
devops/ab279ff20508	devops	medium	devops, version-control	What are some of the advantages of applying GitOps?	* It introduces limited/granular access to infrastructure\n* It makes it easier to trace who makes changes to infrastructure\n* Declarative desired state in git enables drift detection and auto-remediation\n* Pull requests provide a review and approval workflow for infrastructure changes\n* Git history provides a complete audit trail of every change\n\nRemember: GitOps = 'Git as the single source of truth for infrastructure.' Declarative state in a repo, reconciliation loop keeps reality matching the repo.\n\nExample: ArgoCD and Flux are popular GitOps tools for Kubernetes — they watch a git repo and automatically apply changes to the cluster.	projects/knowledge/interview/devops/036-what-are-some-of-the-advantages-of-applying-gitops.txt
devops/b194b6c50a7d	devops	medium	devops, monitoring	What are MTTF (mean time to failure) and MTTR (mean time to repair)? What these metrics help us to evaluate?	* MTTF (mean time to failure) other known as uptime, can be defined as how long the system runs before if fails.\n    * MTTR (mean time to recover) on the other hand, is the amount of time it takes to repair a broken system.\n    * MTBF (mean time between failures) is the amount of time between failures of the system.	projects/knowledge/interview/devops/044-what-are-mttf-mean-time-to-failure-and-mttr-mean-t.txt
devops/b26dddc7776f	devops	hard	ci-cd, devops, incident-response, sre	What is an error budget?	"Atlassian: ""An error budget is the maximum amount of time that a technical system can fail without contractual consequences.""\n\nRead more about it [here](https://www.atlassian.com/incident-management/kpis/error-budget)"\n\nExample: 99.9% SLO = 43.8 min/month error budget. Error budget = 1 - SLO. Burn it on outages → freeze feature deploys until it refills.	projects/knowledge/interview/devops/042-what-is-an-error-budget.txt
devops/b29a019673eb	devops	easy	devops, version-control	What is a Software Repository?	"Wikipedia: ""A software repository, or “repo” for short, is a storage location for software packages. Often a table of contents is stored, as well as metadata.""\n\nRead more [here](https://en.wikipedia.org/wiki/Software_repository)"\n\nExample: Docker Hub (containers), PyPI (Python), npm (JavaScript), Maven Central (Java), Crates.io (Rust). Each ecosystem has its own package registry.	projects/knowledge/interview/devops/018-what-is-a-software-repository.txt
devops/b87f2467707b	devops	hard	automation, ci-cd, devops, monitoring	How to setup auto-scaling group for Jenkins in AWS ?	Here is a high-level overview of how to set up an autoscaling group for Jenkins in Amazon Web Services (AWS):\n```\n    - Launch EC2 instances: Create an Amazon Elastic Compute Cloud (EC2) instance with the desired configuration and install Jenkins on it. This instance will be used as the base image for the autoscaling group.\n	projects/knowledge/interview/devops/057-how-to-setup-auto-scaling-group-for-jenkins-in-aws.txt
devops/b9573d54ee71	devops	hard	devops, sre	What is a configuration drift? What problems is it causing?	Configuration drift happens when in an environment of servers with the exact same configuration and software, a certain server\nor servers are being applied with updates or configuration which other servers don't get and over time these servers become\nslightly different than all others.\n\nThis situation might lead to bugs which hard to identify and reproduce.\n\nExample: server A gets an emergency patch B and C miss. Now A behaves differently under load. terraform plan detects drift before it causes outages.	projects/knowledge/interview/devops/031-what-is-a-configuration-drift-what-problems-is-it-.txt
devops/c05993c17257	devops	hard	devops, backup, disaster-recovery, operations	Why do backups often succeed but restores fail?	"Backups test the write path, not the read path or application consistency.\n\nCommon restore failures:\n\n1. No restore testing\n   - Backup job: green checkmarks for years\n   - First restore attempt: during disaster\n   - ""We've never actually tried restoring""\n\n2. Permission/ownership issues\n   - Backup as root, restore files owned by root\n   - Application runs as different user"\n\nRemember: untested backups are Schrodinger's backups. Test restores monthly. First restore attempt should never be during a real disaster.	projects/knowledge/interview/devops/100-backups-succeed-restores-fail.txt
devops/c210809b391e	devops	hard	ci-cd, culture, devops	A team member of yours, suggests to replace the current CI/CD platform used by the organization with a new one. How would you reply?	Things to think about:\n\n* What we gain from doing so? Are there new features in the new platform? Does the new platform deals with some of the limitations presented in the current platform?\n* What this suggestion is based on? In other words, did he/she tried out the new platform? Was there extensive technical research?\n* What does the switch from one platform to another will require from the o\n\nRemember: migration cost includes retraining, pipeline rewriting, integration reconfiguration, and potential downtime. Evaluate carefully before switching.	projects/knowledge/interview/devops/008-a-team-member-of-yours-suggests-to-replace-the-cur.txt
devops/c5c188d877ee	devops	medium	blameless, devops, monitoring	What do you take into consideration when choosing a tool/technology?	A few ideas to think about:\n\n  * mature/stable vs. cutting edge\n  * community size\n  * architecture aspects - agent vs. agentless, master vs. masterless, etc.\n  * learning curve\n\nRemember: MCAL — Maturity, Community, Architecture (agent vs agentless), Learning curve. Evaluate against your team's skills, not feature lists.	projects/knowledge/interview/devops/006-what-do-you-take-into-consideration-when-choosing-.txt
devops/c82ddeaa480d	devops	easy	culture, devops, sre	What is Reliability? How does it fit DevOps?	Reliability, when used in DevOps context, is the ability of a system to recover from infrastructure failure or disruption. Part of it is also being able to scale based on your organization or team demands.\n\nExample: reliability = redundancy (multi-AZ) + monitoring (Prometheus) + auto-scaling + circuit breaking + chaos engineering.	projects/knowledge/interview/devops/023-what-is-reliability-how-does-it-fit-devops.txt
devops/cc68ac78736b	devops	medium	ci-cd, devops	Why are there multiple software distributions? What differences they can have?	Different distributions can focus on different things like: focus on different environments (server vs. mobile vs. desktop), support specific hardware, specialize in different domains (security, multimedia, ...), etc. Basically, different aspects of the software and what it supports, get different priority in each distribution.	projects/knowledge/interview/devops/017-why-are-there-multiple-software-distributions-what.txt
devops/d1b6c5721be2	devops	hard	devops, reliability, sre, high-availability, philosophy	"Why do ""five nines"" (99.999%) systems fail catastrophically when they do fail?"	The same optimizations that achieve high availability create conditions for catastrophic failure.\n\nThe paradox:\n\n1. Rare paths never tested\n   - 99.999% uptime = 5 minutes downtime/year\n   - Failure recovery code runs 5 min/year\n   - That code has bugs nobody found\n   - When it runs, it fails\n\n2. Humans out of practice\n   - On-call never pages for this system\n\nFun fact: recovery code running 5 min/year has untested bugs. When it finally runs during an outage, IT fails too.	projects/knowledge/interview/devops/105-five-nines-fail-catastrophically.txt
devops/d6ce5e2624a0	devops	medium	automation, ci-cd, devops, monitoring	You need to install periodically a package (unless it's already exists) on different operating systems (Ubuntu, RHEL, ...). How would you do it?	There are multiple ways to answer this question (there is no right and wrong here):\n\n* Simple cron job\n* Pipeline with configuration management technology (such Puppet, Ansible, Chef, etc.)\n...	projects/knowledge/interview/devops/026-you-need-to-install-periodically-a-package-unless-.txt
devops/d855f9990668	devops	hard	architecture, ci-cd, culture, devops, monitoring	How do you decide build vs buy?	Simple framework:\n\n**Build if**:\n* It's core value - competitive advantage\n* Unique requirements that tools don't solve\n* Long-term ownership makes sense\n* You have the team to maintain it\n\n**Buy if**:\n* It's plumbing - everyone needs it\n\nRemember: Build if core value. Buy if plumbing. Ownership cost often exceeds purchase cost. Most teams should buy observability, build differentiators.	projects/knowledge/interview/devops/062-how-do-you-decide-build-vs-buy.txt
devops/da90b0f1657d	devops	medium	ci-cd, devops, version-control	How do you manage build artifacts?	Build artifacts are usually stored in a repository. They can be used in release pipelines for deployment purposes. Usually there is retention period on the build artifacts.\n\nExample: Artifactory, Nexus, GitHub Packages store versioned outputs (Docker images, JARs, npm packages) with retention policies and access control.	projects/knowledge/interview/devops/029-how-do-you-manage-build-artifacts.txt
devops/de7313de317c	devops	medium	ci-cd, culture, devops, monitoring	What SRE team is responsible for?	"Google: ""the SRE team is responsible for availability, latency, performance, efficiency, change management, monitoring, emergency response, and capacity planning of their services""\n\nRead more about it [here](https://sre.google/sre-book/introduction)"\n\nRemember: SRE owns ALPMECEC — Availability, Latency, Performance, Monitoring, Emergency response, Change mgmt, Efficiency, Capacity planning.	projects/knowledge/interview/devops/041-what-sre-team-is-responsible-for.txt
devops/df5c3eea2164	devops	medium	automation, devops, monitoring, version-control	What best practices are you familiar with regarding version control?	* Use a descriptive commit message\n* Make each commit a logical unit\n* Incorporate others' changes frequently\n* Share your changes frequently\n* Coordinate with your co-workers\n* Don't commit generated files\n* Don't commit binary files\n\nGotcha: never commit generated files — they cause merge conflicts and bloat. Build artifacts belong in .gitignore, rebuilt by CI/CD.	projects/knowledge/interview/devops/013-what-best-practices-are-you-familiar-with-regardin.txt
devops/e4ea93c13fec	devops	medium	ci-cd, devops, testing, version-control	"When a repository refereed to as ""GitOps Repository"" what does it means?"	A repository that doesn't holds the application source code, but the configuration, infra, ... files that required to test and deploy the application.\n\nExample: app repo = source + Dockerfile. GitOps repo = Helm charts + Kustomize overlays + env values. ArgoCD watches the GitOps repo.	projects/knowledge/interview/devops/037-when-a-repository-refereed-to-as-gitops-repository.txt
devops/e51fadd2d311	devops	medium	ci-cd, devops, testing	What types of tests are you familiar with?	Styling, unit, functional, API, integration, smoke, scenario, ...\n\nYou should be able to explain those that you mention.\n\nRemember: testing pyramid — Unit (fast, many) → Integration (medium) → E2E (slow, few). Pyramid shape = fast feedback loop.	projects/knowledge/interview/devops/025-what-types-of-tests-are-you-familiar-with.txt
devops/e54608bac40e	devops	hard	devops, testing	Do you have experience with testing cross-projects changes? (aka cross-dependency)	Note: cross-dependency is when you have two or more changes to separate projects and you would like to test them in mutual build instead of testing each change separately.	projects/knowledge/interview/devops/034-do-you-have-experience-with-testing-cross-projects.txt
devops/e6c21eac004d	devops	medium	ci-cd, devops, sre	What are the differences between SRE and DevOps?	"Google: ""One could view DevOps as a generalization of several core SRE principles to a wider range of organizations, management structures, and personnel.""\n\nRead more about it [here](https://sre.google/sre-book/introduction)"\n\nRemember: "SRE implements DevOps" — adds concrete practices (error budgets, SLOs, toil budgets) on top of DevOps cultural principles.	projects/knowledge/interview/devops/040-what-are-the-differences-between-sre-and-devops.txt
devops/e83892835982	devops	medium	automation, ci-cd, devops, monitoring	Can you describe which tool or platform you chose to use in some of the following areas and how?	This is a more practical version of the previous question where you might be asked additional specific questions on the technology you chose\n\n  * CI/CD - Jenkins, Circle CI, Travis, Drone, Argo CD, Zuul\n  * Provisioning infrastructure - Terraform, CloudFormation\n  * Configuration Management - Ansible, Puppet, Chef\n  * Monitoring & alerting - Prometheus, Nagios\n  * Logging - Logstash, Graylog	projects/knowledge/interview/devops/007-can-you-describe-which-tool-or-platform-you-chose-.txt
devops/ec14116932dd	devops	medium	automation, ci-cd, devops	How to backup Jenkins ?	Backing up Jenkins is a very easy process, there are multiple default and configured files and folders in Jenkins that you might want to backup.\n```  \n  - Configuration: The `~/.jenkins` folder. You can use a tool like rsync to backup the entire directory to another location.\n\n\nRemember: back up JENKINS_HOME — config.xml, jobs/, plugins/, secrets/. Use rsync or the ThinBackup plugin.	projects/knowledge/interview/devops/052-how-to-backup-jenkins.txt
devops/f20c753b8793	devops	medium	ci-cd, devops, monitoring	"Explain ""Software Distribution"""	"Read [this](https://venam.nixers.net/blog/unix/2020/03/29/distro-pkgs.html) fantastic article on the topic.\n\nFrom the article: ""Thus, software distribution is about the mechanism and the community that takes the burden and decisions to build an assemblage of coherent software that can be shipped."""	projects/knowledge/interview/devops/016-explain-software-distribution.txt
devops/f7ea00f46fa2	devops	medium	automation, ci-cd, devops	What ways are there to distribute software? What are the advantages and disadvantages of each method?	* Source - Maintain build script within version control system so that user can build your app after cloning repository. Advantage: User can quickly checkout different versions of application. Disadvantage: requires build tools installed on users machine.\n	projects/knowledge/interview/devops/019-what-ways-are-there-to-distribute-software-what-ar.txt

<!-- wiki:related:start -->
---

## Wiki Navigation

### Related Content

- [Postmortem & SLO Drills](../../../../library/drills/postmortem_slo_drills.md) (Drill, L2) — Postmortems & SLOs
- Postmortem SLO Flashcards *(CLI)* (flashcard_deck, L1) — Postmortems & SLOs
- [Postmortems & SLOs](../../../../library/topics/postmortem-slo/index.md) (Topic Pack, L2) — Postmortems & SLOs
- [SRE Practices](../../../../library/topics/sre-practices/index.md) (Topic Pack, L2) — Postmortems & SLOs
- [Skillcheck: Postmortems & SLOs](../../../../library/skillchecks/postmortem-slo.skillcheck.md) (Assessment, L2) — Postmortems & SLOs

<!-- wiki:related:end -->
