Reliability & Resilience

Reliability and resilience focus on ensuring systems continue delivering business value despite failures, outages, unexpected traffic patterns, and operational challenges. Reliable systems perform correctly over time, while resilient systems recover gracefully when failures inevitably occur.

Overview

Failures are not rare events in distributed systems. Hardware fails, software contains defects, networks become unreliable, dependencies become unavailable, and humans occasionally make mistakes.

Architects frequently face questions such as:

What happens if a service goes down?
What happens if a database becomes unavailable?
How quickly can the system recover?
How much downtime is acceptable?
How much data loss can the business tolerate?
How do we prevent one failure from impacting everything?

Reliability and resilience provide a framework for designing systems that continue operating despite failures and recover quickly when problems occur.

Key Insight:
Failures are inevitable. The goal of architecture is not preventing every failure. The goal is minimizing business impact when failures occur.

A Running Example

Consider a healthcare diagnostics platform supporting clinicians, laboratories, patients, and healthcare partners worldwide.

Patient Service
Order Service
Billing Service
Notification Service
Search Service

Patients and clinicians depend on the platform for critical healthcare workflows.

Create Diagnostic Order
View Patient Record
Retrieve Laboratory Results
Generate Billing Information
Send Patient Notifications

Now consider several potential scenarios.

Billing Service Becomes Unavailable
Database Fails
Traffic Suddenly Doubles
A Cloud Region Experiences An Outage
A Deployment Introduces Defects

The primary architectural question becomes:

Can The Platform Continue Delivering Critical Business Functions?

This diagnostics platform will be used throughout the page to explore reliability and resilience concepts.

Why Reliability & Resilience Matter

Users, businesses, and regulators expect systems to remain available when needed.

Failures can impact:

  • Patient Care
  • Business Operations
  • Revenue
  • Regulatory Compliance
  • Customer Trust

A minor technical failure can quickly evolve into a major business problem.

Service Failure
↓
User Impact
↓
Business Impact

Reliability and resilience practices help reduce the likelihood, duration, and impact of such failures.

Architect Perspective:
Business stakeholders rarely care which component failed. They care whether users can continue accomplishing their objectives.

Reliability vs Resilience

Although often discussed together, reliability and resilience address different concerns.

Concept Primary Question
Reliability Can The System Continue Operating Correctly?
Resilience Can The System Recover When Failure Occurs?

A reliable system performs expected functions consistently over time.

Correct Results
Consistent Behavior
Expected Availability

A resilient system responds effectively when something goes wrong.

Failure Occurs
↓
System Adapts
↓
Recovery Achieved

A system may be highly reliable but difficult to recover when failures occur. Conversely, a system may experience occasional failures but recover extremely quickly.

Interview Insight:
Reliability focuses on preventing disruption. Resilience focuses on recovering from disruption.

Understanding Failure

One of the most important mindset shifts in architecture is recognizing that failures are normal.

Many systems are designed as though failures are rare events. In reality, failures occur continuously across infrastructure, software, networks, dependencies, and operations.

Common failure sources include:

Failure Source Examples
Hardware Server Failures, Disk Failures
Software Bugs, Memory Leaks, Deployment Defects
Network Latency, Packet Loss, Connectivity Issues
Dependencies Third-Party Outages, Database Failures
Human Error Misconfigurations, Operational Mistakes

Architects should assume failures will eventually happen.

Failure Is Not An Exception
↓
Failure Is An Expected Condition

Systems designed with this mindset typically recover more effectively and create less business disruption.

Architect Perspective:
The question is rarely whether a component will fail. The question is when it will fail and how the system will respond.

Single Points Of Failure

A Single Point Of Failure (SPOF) is any component whose failure can disrupt the entire system.

Architects actively identify and eliminate these weaknesses.

Consider a platform with a single database.

Applications
↓
Single Database

If the database becomes unavailable, the entire platform may stop functioning.

Other examples include:

  • Single Application Instance
  • Single Database
  • Single Cache Cluster
  • Single Network Path
  • Single Cloud Region

A useful review question is:

If This Component Fails
↓
Can The Business Continue Operating?

If the answer is no, the component may represent a single point of failure.

Interview Insight:
One of the fastest ways to strengthen an architecture is identifying and removing single points of failure.

Redundancy

Redundancy reduces risk by ensuring alternative resources are available when failures occur.

Instead of relying on a single component, multiple instances are deployed.

Application Instance A
Application Instance B
Application Instance C

If one instance fails, remaining instances continue serving requests.

Common redundancy approaches include:

Component Redundancy Approach
Application Multiple Instances
Database Replicas
Network Multiple Paths
Infrastructure Multiple Availability Zones
Deployment Multiple Regions

Redundancy improves availability but also increases cost and operational complexity.

Architect Perspective:
Redundancy is one of the most effective reliability investments because it reduces the impact of inevitable failures.

Failover Strategies

Redundancy alone is not enough. Systems must also determine how traffic moves when failures occur.

This process is known as failover.

Active-Passive
Primary System
↓
Handles Traffic
↓
Failure Occurs
↓
Passive System Activated

Benefits include:

  • Simpler Architecture
  • Lower Operational Complexity
  • Easier Consistency Management

Tradeoffs include:

  • Unused Resources During Normal Operation
  • Potential Recovery Delay
Active-Active
System A
↔
System B
↓
Both Handle Traffic

Benefits include:

  • Better Resource Utilization
  • Faster Recovery
  • Higher Availability

Tradeoffs include:

  • More Complex Coordination
  • Higher Operational Complexity
Strategy Strength
Active-Passive Simplicity
Active-Active Availability
Architect Perspective:
The failover strategy should align with business recovery requirements rather than technology preferences.

Health Checks & Failure Detection

Systems cannot recover from failures they do not detect.

Reliable architectures continuously monitor component health.

Liveness Checks

Liveness checks determine whether a component is still running.

Application Running?
↓
Yes Or No
Readiness Checks

Readiness checks determine whether a component is capable of serving traffic.

Application Running
↓
Database Available?
Cache Available?
Dependencies Available?
Monitoring & Detection

Failures may be detected through:

  • Health Checks
  • Metrics
  • Logs
  • Alerts
  • Tracing Systems

Early detection reduces recovery time and limits business impact.

Failure Occurs
↓
Detected Quickly
↓
Recovery Begins
Architect Perspective:
A failure that goes unnoticed for hours is often more damaging than the failure itself.

Timeouts

One of the simplest and most effective resilience techniques is ensuring requests do not wait indefinitely.

Without timeouts, failures can spread through the system as resources become consumed waiting for responses that may never arrive.

Request Sent
↓
Dependency Slow
↓
Wait Forever

With timeouts, systems can fail predictably and begin recovery sooner.

Request Sent
↓
Timeout Reached
↓
Fallback Or Recovery Action

Timeouts help:

  • Protect Resources
  • Reduce Cascading Failures
  • Improve Recovery Behavior
  • Prevent Resource Exhaustion

Timeout values should be based on business requirements, user expectations, and dependency behavior.

Architect Perspective:
Every network call should assume the dependency may never respond.

Retries

Some failures are temporary and may succeed if attempted again.

Retries provide a mechanism for recovering from transient failures.

Request Fails
↓
Retry
↓
Success

Retries can be valuable when dealing with:

  • Network Interruptions
  • Temporary Service Unavailability
  • Short-Term Resource Constraints
  • Transient Infrastructure Issues

However, retries also introduce risk.

Service Under Stress
↓
Clients Retry Aggressively
↓
Service Receives More Load

This situation is known as a retry storm.

Architects should define:

  • Retry Limits
  • Backoff Strategies
  • Retry Eligibility Rules
  • Failure Escalation Approaches
Interview Insight:
Retries can improve resilience, but poorly designed retries can become part of the problem rather than the solution.

Circuit Breakers

When a dependency is failing repeatedly, continuously sending requests often provides little value.

Circuit breakers protect systems by temporarily stopping requests to unhealthy dependencies.

Repeated Failures
↓
Circuit Opens
↓
Requests Blocked

This prevents additional pressure from being placed on a dependency that is already struggling.

A typical circuit breaker lifecycle follows:

Closed
↓
Failures Increase
↓
Open
↓
Recovery Attempt
↓
Closed Again

Benefits include:

  • Reduced Cascading Failures
  • Dependency Protection
  • Faster Failure Detection
  • Improved Recovery Behavior
Architect Perspective:
A dependency that is already failing rarely benefits from receiving more traffic.

Bulkheads

Bulkheads isolate failures so that problems in one area do not spread throughout the entire system.

The concept originates from ship design where compartments prevent flooding from affecting the entire vessel.

In software systems, isolation limits the blast radius of failures.

Notification Service Failure
↓
Notification Workload Impacted
↓
Billing Continues Operating
Orders Continue Processing

Common isolation strategies include:

  • Separate Service Resources
  • Dedicated Thread Pools
  • Independent Queues
  • Independent Infrastructure

Without isolation, a single overloaded component can consume resources required by unrelated workloads.

Architect Perspective:
Failures should be contained rather than allowed to spread freely through a platform.

Graceful Degradation

In many situations, partial functionality is preferable to complete unavailability.

Graceful degradation allows systems to continue providing critical capabilities even when some components are unavailable.

Consider an e-commerce platform.

Recommendations Unavailable
↓
Checkout Continues Working

For the diagnostics platform:

Notification Service Down
↓
Orders Continue Processing

Architects should identify:

  • Critical Business Functions
  • Optional Features
  • Acceptable Degraded States
  • Recovery Priorities

This approach reduces business disruption during outages.

Interview Insight:
Users often tolerate reduced functionality better than total service unavailability.

Rate Limiting & Load Shedding

Not all failures originate from broken components. Sometimes systems fail because demand exceeds their capacity.

Rate limiting controls how much work the system accepts.

Request Volume Exceeds Threshold
↓
Additional Requests Rejected

Load shedding intentionally drops low-priority work to preserve critical business functions.

Traffic Spike
↓
Low Priority Requests Rejected
↓
Critical Requests Protected

Benefits include:

  • Resource Protection
  • Improved Stability
  • Predictable Behavior During Peaks
  • Higher Availability For Critical Workloads

Architects should identify which workloads receive priority during periods of extreme demand.

Patient Record Access
Higher Priority
Analytics Refresh Requests
Lower Priority
Architect Perspective:
A system that attempts to serve every request during overload conditions may ultimately fail all requests.

Disaster Recovery

Not every failure is small. Some failures affect entire environments, regions, or critical data stores.

Disaster Recovery (DR) focuses on restoring business operations after major disruptions.

Potential disaster scenarios include:

  • Regional Outages
  • Database Corruption
  • Ransomware Attacks
  • Infrastructure Failures
  • Operational Mistakes
  • Data Center Outages

A typical disaster recovery flow may look like:

Major Failure
↓
Failure Detected
↓
Recovery Plan Activated
↓
Services Restored

Common disaster recovery capabilities include:

  • Backups
  • Database Replication
  • Secondary Environments
  • Multi-Region Deployments
  • Automated Recovery Procedures

Architects should ask:

What Happens If An Entire Environment Is Lost?
↓
Can The Business Recover?
Architect Perspective:
Backups are not disaster recovery. Recovery is only proven when restoration has been tested successfully.

RTO & RPO

Reliable recovery begins with understanding business expectations.

Two important recovery metrics are Recovery Time Objective (RTO) and Recovery Point Objective (RPO).

Recovery Time Objective (RTO)

RTO defines how quickly a service must be restored after a failure.

Failure Occurs
↓
System Restored
↓
Time Elapsed = RTO
Recovery Point Objective (RPO)

RPO defines how much data loss is acceptable.

Failure Occurs
↓
Restore Data
↓
Acceptable Data Loss = RPO
Metric Question Answered
RTO How Fast Must We Recover?
RPO How Much Data Can We Lose?

Different systems often have very different recovery expectations.

System Typical Expectations
Patient Care Platform Very Low RTO & RPO
Billing Platform Low RTO & RPO
Internal Reporting Higher RTO & RPO
Interview Insight:
The business determines RTO and RPO. Technology decisions are then made to meet those targets.

Observability & Monitoring

Systems cannot be operated reliably if engineers cannot understand what is happening within them.

Observability provides visibility into system behavior and health.

Metrics

Metrics provide numerical insights into system performance.

Latency
Error Rate
Request Volume
Resource Utilization
Logs

Logs provide detailed records of system activity.

Tracing

Distributed tracing helps track requests as they move across services.

User Request
↓
API Service
↓
Order Service
↓
Database
Alerts

Monitoring becomes actionable when abnormal conditions generate alerts.

Error Rate Increases
↓
Alert Triggered
↓
Investigation Begins

Observability reduces detection time and accelerates recovery.

Architect Perspective:
A failure that cannot be detected cannot be managed effectively.

Chaos Engineering

Many organizations assume systems are resilient without ever validating those assumptions.

Chaos engineering introduces controlled failures to evaluate system behavior under real-world conditions.

Simulate Failure
↓
Observe Behavior
↓
Improve Resilience

Examples include:

  • Terminating Application Instances
  • Disabling Dependencies
  • Introducing Network Latency
  • Simulating Database Failures

The objective is learning.

Assumption
↓
Experiment
↓
Validation

Effective resilience often comes from repeatedly testing failure scenarios before real outages occur.

Architect Perspective:
Systems are often less resilient than teams believe until failures are intentionally tested.

Reliability Failure Scenarios

Reliable systems anticipate common failure patterns and define recovery approaches in advance.

Service Failure
Billing Service Down
↓
Fallback Applied
↓
Order Processing Continues
Database Failure
Primary Database Unavailable
↓
Failover To Replica
Traffic Spike
Traffic Exceeds Capacity
↓
Rate Limiting Applied
Region Failure
Primary Region Outage
↓
Traffic Redirected
Misconfiguration
Deployment Error
↓
Rollback Performed

Architects continuously evaluate these scenarios during design reviews because many outages originate from known failure patterns.

Interview Insight:
Reliability discussions become significantly stronger when architects can explain how specific failure scenarios are handled.

Reliability Tradeoffs

Architects often discuss reliability as though more reliability is always better. In practice, every reliability improvement introduces tradeoffs.

Higher reliability frequently requires additional infrastructure, operational processes, monitoring, testing, and cost.

Goal Typical Tradeoff
Higher Availability Higher Cost
Faster Recovery More Redundancy
Lower Data Loss Greater Complexity
Multi-Region Resilience Higher Operational Overhead
Extreme Availability Additional Architecture Complexity

For example, achieving 99.9% availability and achieving 99.999% availability are dramatically different engineering challenges.

Higher Reliability
↓
Higher Cost
↓
Higher Complexity

The appropriate level of reliability should be driven by business requirements rather than technology enthusiasm.

Architect Perspective:
The objective is not maximizing reliability at all costs. The objective is achieving the level of reliability the business actually needs.

Choosing The Right Reliability Strategy

Not every application requires the same level of resilience.

Architects should evaluate reliability investments based on business impact.

System Type Reliability Expectations
Patient Care Platform Very High
Billing System High
Customer Portal High
Internal Reporting Moderate
Development Tools Lower

A practical decision framework is:

How Important Is The Workload?
↓
What Is The Business Impact Of Failure?
↓
What Recovery Expectations Exist?
↓
What Reliability Investment Is Justified?

Reliability strategies should be proportional to the consequences of failure.

Interview Insight:
Reliability requirements should always be driven by business priorities rather than technical preferences.

Real-World Case Study

Consider the healthcare diagnostics platform supporting clinicians, laboratories, and patients across multiple locations.

Several critical business functions depend on platform availability.

Patient Access
Order Processing
Laboratory Results
Billing Operations
Notifications

The architecture team identified several reliability risks.

Single Database Dependency
Service Failures
Traffic Spikes
Deployment Risks

To improve reliability, the platform adopted:

  • Redundant Service Instances
  • Database Replication
  • Automated Failover
  • Circuit Breakers
  • Retry Policies
  • Health Monitoring
  • Automated Backups

If the Notification Service becomes unavailable, order processing continues through graceful degradation.

Notification Failure
↓
Orders Continue Processing
↓
Notifications Recover Later

The result was improved availability, reduced downtime, and lower operational risk without requiring every component to be perfectly available.

Architect Perspective:
The most successful reliability improvements often focus on reducing business impact rather than eliminating every technical failure.

Reliability & Resilience Review Checklist

The following checklist can be used during architecture reviews and production readiness assessments.

✅ Single Points Of Failure Identified
✅ Redundancy Strategy Defined
✅ Failover Approach Defined
✅ Timeout Policy Configured
✅ Retry Policy Defined
✅ Circuit Breakers Evaluated
✅ Bulkhead Isolation Considered
✅ Graceful Degradation Defined
✅ Monitoring Strategy Defined
✅ Alerting Strategy Defined
✅ Backup Strategy Defined
✅ Disaster Recovery Plan Defined
✅ RTO Established
✅ RPO Established
✅ Failure Testing Performed

Reliability & Resilience Canvas

The Reliability & Resilience Canvas summarizes major reliability decisions for a system.

Area Example
Critical Services Order Processing, Patient Access
Redundancy Strategy Multi-Instance Deployment
Failover Strategy Active-Passive
Timeout Policy 5 Seconds
Retry Strategy Exponential Backoff
Circuit Breakers Enabled
Monitoring Logs, Metrics, Traces
Backup Frequency Hourly
RTO 30 Minutes
RPO 5 Minutes
Primary Risk Regional Outage

Common Anti-Patterns

No Timeouts

Requests that wait indefinitely can create resource exhaustion and cascading failures.

Retry Storms

Aggressive retry behavior can overwhelm already struggling dependencies.

Single Region Everything

Deploying all critical resources in a single location increases outage risk.

No Failure Testing

Assuming systems are resilient without validating behavior often leads to unpleasant surprises during real incidents.

Monitoring As An Afterthought

Without visibility, diagnosis and recovery become difficult.

Assuming Cloud Means Reliable

Cloud providers offer resilient building blocks, but architectural decisions still determine system reliability.

Ignoring Recovery Requirements

Failing to define RTO and RPO often creates unrealistic expectations during outages.

Architect Perspective:
Many outages are not caused by unknown problems. They are caused by known risks that were never addressed.

How Concepts Connect

Reliability and resilience are deeply connected to many other architecture disciplines.

Requirements & Workloads
↓
Capacity & Performance
↓
Scalability & Elasticity
↓
Communication & APIs
↓
Messaging & Events
↓
Reliability & Resilience
↓
Business Continuity

As systems become more distributed, reliability increasingly depends on how components interact, recover, and respond to failures.

Key Takeaway

Reliability is the ability to continue operating correctly over time.

Resilience is the ability to recover when failures occur.

Failures are inevitable in distributed systems.

The objective of architecture is not preventing every failure.

The objective is minimizing business impact when failures occur.

Successful systems detect problems quickly, isolate failures, recover predictably, and continue delivering critical business functions.

The best architectures are not those that never fail.

The best architectures are those that fail gracefully, recover rapidly, and maintain user trust throughout the recovery process.