Consistency & Transactions

Consistency and Transactions ensure systems remain correct as multiple users, services, databases, and business processes interact. Architects use consistency models and transaction strategies to balance correctness, availability, performance, and scalability.

Overview

Modern systems rarely update a single database using a single request. Business operations often span multiple services, databases, APIs, and asynchronous workflows.

Architects therefore face questions such as:

What happens if payment succeeds but order creation fails?
How do multiple services remain consistent?
What happens when part of a workflow succeeds and another part fails?
When should strong consistency be required?
When can eventual consistency be acceptable?

Consistency and transaction strategies help systems preserve business correctness while operating at scale.

Key Insight:
Consistency is not about databases. It is about ensuring business rules remain true even when failures occur.

A Running Example

Consider a healthcare diagnostics platform used by clinicians and laboratories.

A typical workflow may involve:

Create Lab Order
↓
Schedule Diagnostic Test
↓
Generate Bill
↓
Notify Patient

Each step may be handled by a different service.

The challenge begins when part of the workflow succeeds and another part fails.

Lab Order Created ✅
Billing Completed ✅
Patient Notification Failed ❌

Should everything be rolled back? Should compensating actions be executed? Should the workflow continue?

These are consistency and transaction design decisions.

Why Consistency Matters

Business users expect systems to behave correctly even when failures occur.

Without proper consistency controls, systems can enter invalid states.

Examples include:

Order Created
Payment Missing
Inventory Reserved
Order Cancelled
Lab Scheduled
Patient Never Notified

These situations may create operational, financial, regulatory, and customer experience issues.

The purpose of consistency mechanisms is to ensure business processes remain trustworthy despite failures, retries, concurrency, and distributed communication challenges.

Architect Perspective:
Users care about business correctness, not transaction protocols. Consistency techniques exist to protect business outcomes.

Transaction Fundamentals

A transaction groups multiple operations into a logical unit of work.

The goal is simple:

Either Everything Succeeds
Or
Nothing Succeeds

Consider a laboratory order.

Create Order
Reserve Slot
Generate Bill

If bililng fails after the order is created, a transaction can prevent the system from remaining in a partially completed state.

Key transaction concepts include:

  • Commit
  • Rollback
  • Atomicity
  • Consistency

Transactions are one of the primary mechanisms used to preserve business correctness.

ACID Transactions

ACID is one of the most widely used approaches for maintaining data correctness within transactional systems.

Rather than viewing ACID as four definitions to memorize, architects should think of ACID as a set of guarantees that help protect critical business operations.

Atomicity

All operations within the transaction succeed or none of them succeed.

Create Lab Order
Reserve Appointment Slot
Generate Bill
↓
If Billing Fails
Everything Rolls Back
Consistency

Business rules remain valid before and after the transaction executes.

Lab Test Cannot Exist
Without A Valid Patient
Isolation

Concurrent transactions should not interfere with one another in unexpected ways.

Clinician A Updates Record
Clinician B Updates Record
↓
Data Remains Predictable
Durability

Once a transaction commits successfully, the results survive crashes, restarts, and failures.

Transaction Committed
↓
System Restarts
↓
Data Remains

ACID is extremely valuable when business correctness is more important than raw scalability.

Common Examples:
Payments
Patient Records
Orders
Inventory Management
Accounting Systems

When ACID Becomes Expensive

ACID provides strong guarantees, but those guarantees are not free.

As systems become larger, more distributed, and more globally available, maintaining strict consistency often introduces additional overhead.

Potential costs include:

  • Higher Latency
  • More Coordination
  • Reduced Throughput
  • Additional Locking
  • Scalability Constraints

Consider a globally distributed diagnostics platform.

North America
Europe
Asia Pacific

If every update must be coordinated across all regions before completing, response times may increase significantly.

More Consistency
↓
More Coordination
↓
More Latency

This does not mean ACID is bad. It means architects must evaluate whether every workload truly requires strict transactional guarantees.

Architect Perspective:
The goal is not maximizing consistency everywhere. The goal is applying the appropriate level of consistency to each business requirement.

BASE Systems

Many large-scale distributed systems relax strict consistency requirements in exchange for improved scalability, availability, and resilience.

BASE stands for:

Concept Meaning
Basically Available System Continues Responding
Soft State State May Change Over Time
Eventual Consistency Data Converges Over Time

Consider a recommendation system.

New Data Generated
↓
Recommendations Updated Later

A brief delay rarely creates significant business risk.

Common examples include:

  • Activity Feeds
  • Recommendation Systems
  • Analytics Dashboards
  • Search Suggestions
  • Telemetry Platforms

BASE systems prioritize continued operation and scalability while accepting temporary inconsistencies.

Update Occurs
↓
Replication Happens
↓
System Eventually Converges
Architect Perspective:
Eventual consistency does not mean incorrect forever. It means consistency arrives over time rather than immediately.

ACID vs BASE

One of the most common architecture discussions involves determining whether strong consistency or eventual consistency is appropriate.

Requirement Typical Choice
Payments ACID
Billing ACID
Patient Records ACID
Inventory Management ACID
Analytics Dashboards BASE
Recommendations BASE
Search Suggestions BASE
Activity Feeds BASE

A useful framework is:

Business Correctness Critical?
↓
Favor ACID
Can Temporary Inconsistency Be Tolerated?
↓
Consider BASE

Many modern systems use both models simultaneously, applying the appropriate consistency model to each workload.

Interview Insight:
Experienced architects rarely choose ACID or BASE for an entire platform. They choose the appropriate consistency model for each business capability.

Isolation Levels

Transactions rarely execute in isolation. In real systems, hundreds or thousands of users may update data simultaneously.

Isolation levels control how concurrent transactions interact with one another.

The goal is to balance data correctness with system throughput.

Dirty Reads

A transaction reads data that has not yet been committed.

Transaction A Updates Bill
↓
Transaction B Reads Change
↓
Transaction A Rolls Back

Transaction B has now used information that never actually existed.

Non-Repeatable Reads

A transaction reads the same record twice and receives different values.

Read Patient Balance = $100
↓
Another Transaction Updates Balance
↓
Read Again = $150
Phantom Reads

A transaction executes the same query twice and discovers new records.

Query Today’s Lab Orders
↓
Another Transaction Creates New Orders
↓
Same Query Returns Additional Records

Stronger isolation reduces these problems but increases coordination and concurrency costs.

Higher Isolation
↓
Greater Data Protection
↓
Reduced Concurrency
Architect Perspective:
Isolation levels exist because multiple users often interact with the same business data simultaneously.

Choosing Isolation Levels

One of the most common misconceptions is that every workload should use the highest isolation level available.

In reality, stronger isolation often increases locking, contention, and latency.

Workload Typical Isolation Need
Payments High
Inventory Management High
Patient Records High
Reporting Moderate
Analytics Queries Lower
Recommendation Systems Lower

For the diagnostics platform, patient updates may require a higher isolation level than analytical dashboards.

Critical Business Data
↓
Favor Higher Isolation
Reporting And Analytics
↓
Lower Isolation Often Acceptable

The appropriate isolation level depends on business risk, correctness requirements, and concurrency expectations.

Interview Insight:
The best isolation level is not necessarily the strongest one. It is the one that delivers the required level of correctness without unnecessary performance penalties.

Distributed Transactions

Traditional transactions work well when all operations occur within a single database.

Modern systems often involve multiple services and multiple databases.

Consider the diagnostics platform.

Order Service
↓
Scheduling Service
↓
Billing Service
↓
Notification Service

A single business workflow may span several independent systems.

The challenge becomes:

How Do We Keep Multiple Systems Consistent?

If billing succeeds but scheduling fails, the platform may enter an invalid business state.

Distributed transactions attempt to coordinate updates across multiple participants while preserving consistency.

Unlike local transactions, distributed transactions must handle:

  • Network Failures
  • Timeouts
  • Partial Success
  • Independent Databases
  • Independent Services
Key Insight:
The difficulty of distributed transactions is not the transaction itself. The difficulty is coordinating multiple systems that can fail independently.

Why Distributed Transactions Are Hard

Distributed systems introduce failure scenarios that do not exist within a single database.

Consider the following workflow:

Create Lab Order ✅
Generate Bill ✅
Schedule Test ❌

The system must now answer difficult questions.

Should Billing Be Reversed?
Should The Order Be Removed?
Should The Workflow Retry?
Should An Operator Be Notified?

Network failures make these situations even more complex.

Service A Sends Commit Request
↓
Network Failure Occurs
↓
Did Service B Commit?

Sometimes the system cannot immediately determine what actually happened.

Common distributed transaction challenges include:

  • Partial Failures
  • Duplicate Requests
  • Retries
  • Network Partitions
  • Service Timeouts
  • Recovery Coordination

These challenges are the reason modern architectures often look for alternatives to traditional distributed transactions.

Architect Perspective:
Distributed transactions force architects to think about failure scenarios first and success scenarios second.

Two-Phase Commit (2PC)

Two-Phase Commit is a coordination protocol designed to maintain strong consistency across multiple participating systems.

The protocol introduces a coordinator responsible for ensuring all participants either commit or abort together.

2PC operates in two stages.

Phase 1: Prepare
Coordinator
↓
Prepare Request
↓
Order Service
Billing Service
Scheduling Service

Each participant verifies whether it can successfully complete the transaction.

Can Commit?
↓
YES / NO
Phase 2: Commit

If all participants agree, the coordinator instructs everyone to commit.

All Participants Agree
↓
Commit Request
↓
Changes Persisted

If any participant rejects the request, the coordinator instructs all participants to roll back.

Any Participant Rejects
↓
Rollback Everywhere

The primary benefit of 2PC is strong consistency across multiple systems.

Architect Perspective:
2PC attempts to make multiple independent systems behave like a single transaction.

Why Architects Often Avoid 2PC

Although Two-Phase Commit provides strong consistency, many modern architectures avoid it for large-scale distributed systems.

The challenge is that strong coordination introduces significant operational costs.

Common concerns include:

  • Higher Latency
  • Coordinator Dependency
  • Reduced Availability
  • Blocking Behavior During Failures
  • Additional Scalability Constraints

Consider the healthcare diagnostics platform.

North America
Europe
Asia Pacific

If every transaction must wait for all regions to acknowledge success, user response times may increase substantially.

More Coordination
↓
More Waiting
↓
Higher Latency

Additionally, if the coordinator becomes unavailable at the wrong moment, participants may remain blocked while waiting for instructions.

This is one reason many modern microservice architectures favor eventually consistent approaches rather than distributed ACID transactions.

Interview Insight:
A common interview discussion is not “How does 2PC work?” but rather “Why might you avoid 2PC in large distributed systems?”

Saga Pattern

The Saga Pattern is a common alternative to distributed ACID transactions in microservice architectures.

Instead of treating multiple services as a single atomic transaction, Saga breaks a business workflow into a sequence of smaller local transactions.

Create Lab Order
↓
Schedule Diagnostic Test
↓
Generate Bill
↓
Notify Patient

Each step commits independently.

If a later step fails, compensating actions are executed to reverse previously completed work.

Consider the following scenario.

Create Lab Order ✅
Schedule Diagnostic Test ✅
Generate Bill ✅
Notify Patient ❌

Instead of rolling back a distributed transaction, Saga performs compensating actions.

Notification Failed
↓
Cancel Bill
↓
Release Appointment Slot
↓
Cancel Lab Order

This approach improves scalability and availability while accepting eventual consistency.

Benefits include:

  • Lower Coupling
  • Higher Scalability
  • Improved Availability
  • Works Well With Microservices

Challenges include:

  • Compensation Complexity
  • Eventual Consistency
  • Operational Visibility
  • Error Handling Complexity
Architect Perspective:
Saga does not prevent failures. It provides a structured way to recover from failures while preserving business correctness.

Saga Orchestration vs Choreography

Saga workflows are typically implemented using either orchestration or choreography.

Orchestration

A central coordinator manages the workflow.

Saga Orchestrator
↓
Order Service
↓
Billing Service
↓
Notification Service

Benefits:

  • Centralized Visibility
  • Easier Workflow Management
  • Simplified Troubleshooting

Tradeoffs:

  • Coordinator Dependency
  • Potential Central Bottleneck
Choreography

Services react to events without a central controller.

Order Created Event
↓
Billing Service Reacts
↓
Billing Completed Event
↓
Notification Service Reacts

Benefits:

  • Loose Coupling
  • Higher Autonomy
  • Natural Event-Driven Design

Tradeoffs:

  • More Difficult Troubleshooting
  • Less Process Visibility
  • Complex Event Flows
Characteristic Orchestration Choreography
Control Centralized Distributed
Visibility High Lower
Coupling Moderate Low
Troubleshooting Easier Harder

Idempotency Considerations

Distributed systems frequently retry operations because timeouts, network failures, and temporary service outages are common.

Retries introduce an important risk:

Operation Executed
↓
Response Lost
↓
Client Retries

If the operation is not idempotent, duplicate business actions may occur.

Example:

Charge Credit Card
↓
Retry
↓
Charge Again

Idempotent operations guarantee that repeated execution produces the same outcome.

Charge Request #123
↓
Already Processed
↓
No Duplicate Charge

Idempotency is especially important for Saga workflows because retries commonly occur during distributed processing.

Architect Perspective:
Many distributed consistency problems are actually retry problems. Idempotency provides a critical safety mechanism for handling retries correctly.

2PC vs Saga

One of the most common architecture discussions involves choosing between Two-Phase Commit and Saga-based coordination.

Both approaches attempt to maintain business correctness across multiple systems, but they do so using very different strategies.

Characteristic 2PC Saga
Consistency Strong Eventual
Coordination Centralized Distributed
Availability Lower Higher
Scalability Limited Better
Failure Handling Rollback Compensation
Microservice Suitability Lower Higher

A simple decision framework is:

Single Database
↓
Traditional ACID Transaction
Multiple Services
Strong Consistency Required
↓
Consider 2PC
Multiple Services
High Availability And Scalability Required
↓
Consider Saga

Modern microservice platforms typically favor Saga because it aligns better with distributed architectures and independent service ownership.

Interview Insight:
A common senior-level answer is not “always use Saga” or “always use 2PC.” The correct answer depends on consistency requirements, business risk, and operational constraints.

Consistency Decision Framework

Consistency decisions should begin with business requirements rather than technology preferences.

Architects should ask several fundamental questions.

What Happens If Data Is Temporarily Incorrect?
Can Corrections Be Applied Later?
Is Financial Loss Possible?
Is Regulatory Risk Involved?
How Much Latency Is Acceptable?

The answers often lead to different consistency models.

Business Requirement Typical Approach
Payments Strong Consistency
Billing Strong Consistency
Patient Records Strong Consistency
Inventory Management Strong Consistency
Reporting Eventual Consistency
Recommendations Eventual Consistency
Analytics Eventual Consistency

Many successful systems use multiple consistency models simultaneously.

Patient Records
↓
Strong Consistency

Analytics Dashboard
↓
Eventual Consistency

The objective is to apply consistency where it creates business value rather than enforcing it universally.

Real-World Case Study

Consider the healthcare diagnostics platform processing a lab order workflow.

Create Lab Order
↓
Schedule Diagnostic Test
↓
Generate Bill
↓
Notify Patient

Initially, the architecture team attempted to coordinate everything using tightly coupled transactional behavior.

As services became independently deployed and scaled, this approach became increasingly difficult to manage.

Common challenges included:

  • Long Running Transactions
  • Service Timeouts
  • Availability Dependencies
  • Operational Complexity

The architecture evolved toward Saga-based coordination.

If patient notification failed:

Cancel Billing
↓
Release Appointment Slot
↓
Cancel Lab Order

Compensating actions restored business correctness without requiring a global distributed transaction.

The result was improved scalability, better availability, and clearer ownership between services while preserving the integrity of the overall workflow.

Architect Perspective:
The goal was not perfect technical consistency. The goal was maintaining correct business outcomes while operating at scale.

Consistency & Transactions Review Checklist

The following checklist can be used during architecture reviews and solution design discussions.

✅ Consistency Requirements Defined
✅ Transaction Boundaries Identified
✅ ACID vs BASE Evaluated
✅ Isolation Level Selected
✅ Distributed Transaction Need Validated
✅ 2PC Tradeoffs Evaluated
✅ Saga Strategy Defined
✅ Compensation Actions Identified
✅ Failure Scenarios Reviewed
✅ Retry Strategy Defined
✅ Idempotency Considered
✅ Monitoring And Observability Planned

Consistency & Transactions Canvas

The Consistency & Transactions Canvas helps summarize design decisions and transaction strategies.

Area Example
Consistency Model Strong Consistency
Transaction Scope Single Service
Isolation Level Read Committed
Distributed Strategy Saga
Consistency Pattern Eventual Consistency
Compensation Strategy Enabled
Retry Strategy Exponential Backoff
Idempotency Required
Primary Risk Partial Failure

Common Anti-Patterns

Strong Consistency Everywhere

Attempting to maximize consistency for every workload often reduces scalability, availability, and performance.

Distributed ACID For Every Workflow

Using distributed transactions everywhere introduces significant coordination costs and operational complexity.

Saga Without Compensation

Executing a Saga without defining compensating actions leaves workflows vulnerable to partial failures.

Ignoring Idempotency

Retries are common in distributed systems. Without idempotency, duplicate business actions may occur.

Assuming Network Calls Always Succeed

Distributed transactions must assume service failures, timeouts, and network interruptions will happen.

Using Eventual Consistency Without Understanding Impact

Eventual consistency introduces temporary inconsistencies that business stakeholders must be willing to accept.

Architect Perspective:
Many consistency problems originate from unrealistic assumptions about failures rather than flaws in the consistency model itself.

How Concepts Connect

Consistency and Transactions are closely related to many other architectural topics.

Business Requirements
↓
Data Architecture
↓
Consistency Requirements
↓
Transaction Design
↓
Distributed Systems
↓
Scalability & Reliability
↓
Business Correctness

As systems become more distributed, transaction design, consistency models, failure handling, and recovery mechanisms become increasingly important architectural concerns.

Key Takeaway

Consistency is not about protecting databases.

It is about protecting business correctness.

Transactions are not valuable because they commit or roll back.

They are valuable because they help ensure business rules remain true when failures occur.

ACID transactions provide strong guarantees for critical workloads.

BASE systems trade immediate consistency for scalability and availability.

Distributed transactions introduce additional complexity because multiple systems can fail independently.

Modern architectures frequently use Saga patterns, compensation actions, retries, and idempotency to maintain correctness without requiring global distributed transactions.

The job of an architect is not to maximize consistency everywhere.

The job is to apply the right level of consistency, coordination, and transactional protection to each business capability while balancing scalability, availability, performance, and operational complexity.