Overview
Capacity and Performance determine whether a system can meet user expectations under real-world conditions.
A system may provide all required functionality, yet still fail if it cannot process requests fast enough, support expected user volumes, or maintain acceptable response times during peak demand.
Architects frequently face questions such as:
What response times should users expect?
How much traffic can the platform handle?
Where are potential bottlenecks?
When will capacity limits be reached?
How much headroom should be planned?
Answering these questions requires understanding workload characteristics, performance objectives, capacity limits, and system behavior under stress.
Whether designing a new platform, reviewing an existing architecture, troubleshooting production issues, or preparing for a system design interview, capacity and performance analysis help architects make informed decisions based on measurable data rather than assumptions.
A Running Example
Throughout this page we will continue using the healthcare diagnostics platform from the previous section.
The system supports:
100,000 Diagnostic Orders Per Day
5,000,000 Result Searches Per Day
99.95% Availability Requirement
< 2 Second Search Response Target
These numbers describe the workload.
Capacity and Performance determine whether the architecture can support those expectations reliably.
Why Capacity & Performance Matter
Many systems initially appear successful because functionality works correctly.
Problems often appear later when real users arrive.
Why Capacity & Performance Matter
Many systems initially appear successful because functionality works correctly.
Problems often appear later when real users arrive.
↓
Users Increase
↓
Response Times Increase
↓
Failures Begin Appearing
Capacity and Performance help architects identify these risks before systems reach production.
The goal is not to build the largest possible system. The goal is to build a system capable of meeting expected demand with acceptable performance and reasonable cost.
Capacity Planning
Capacity planning translates business demand into technical requirements.
For example:
Architects immediately begin asking:
- How many application instances are required?
- How many database connections are necessary?
- How much memory is needed?
- How much storage growth should be expected?
Capacity planning prevents infrastructure decisions from becoming guesswork.
A workload estimate without a capacity plan is simply a prediction.
Translating Workloads Into Capacity
One of the most valuable skills in system design is converting business demand into measurable system load.
Suppose:
250 Searches Per Day Per Clinician
Estimated workload:
= 5,000,000 Searches Per Day
Average requests per second:
≈ 58 Requests Per Second
However, average traffic is rarely sufficient for planning capacity.
Peak traffic often matters more than average traffic.
Throughput Modeling
Throughput represents the amount of work a system can complete during a specific period.
A useful way to think about throughput is:
↓
Processing Occurs
↓
Response Returned
If a system processes 500 requests per second, then its throughput is:
Architects compare expected demand against available throughput to determine whether a system will saturate under load.
Peak vs Average Load
One of the most common design mistakes is planning for average traffic rather than peak traffic.
Example:
Peak Load = 580 Requests Per Second
The platform may appear healthy under average conditions while failing during peak business hours.
Good architects design around realistic peaks, not averages.
Infrastructure failures occur during peaks, not averages.
Latency Analysis
Throughput explains how much work a system can process. Latency explains how long users must wait.
From a user perspective, latency is usually more visible than throughput.
Consider a clinician searching for diagnostic results.
↓
System Processing
↓
Results Returned
Even if the system supports thousands of requests per second, users may still experience poor performance if response times are high.
Architects therefore measure both throughput and latency when evaluating systems.
Visual Latency Flow
Latency is rarely caused by a single component. It is often the sum of many small delays across a request path.
↓
Network = 200ms
↓
API Gateway = 150ms
↓
Application Service = 300ms
↓
Database = 750ms
↓
Response Rendering = 100ms
↓
Total = 1500ms
Architects use latency decomposition to identify where the majority of time is spent.
Optimization efforts should focus on the largest contributors rather than every component equally.
How To Measure Latency
Latency can be measured at multiple levels.
| Measurement | Description |
|---|---|
| End-to-End Latency | User Request Until User Sees Response |
| Application Latency | Request Arrival Until Response Generated |
| Database Latency | Query Submission Until Query Completion |
| Network Latency | Time Spent Moving Data Between Systems |
End-to-end latency is usually the metric users care about most because it reflects their actual experience.
P50, P95 and P99 Latency
Average latency alone can hide serious performance issues.
Architects often monitor latency percentiles.
| Metric | Meaning |
|---|---|
| P50 | 50% Of Requests Were Faster Than This Value |
| P95 | 95% Of Requests Were Faster Than This Value |
| P99 | 99% Of Requests Were Faster Than This Value |
Example:
P95 = 1.5 Seconds
P99 = 4 Seconds
This means most users are experiencing acceptable performance, while a small percentage may be experiencing significant delays.
Performance Budgets
Performance goals should be allocated across system components.
Suppose the search experience must complete within two seconds.
API Gateway = 150ms
Application Service = 350ms
Database = 800ms
Security & Logging = 100ms
Buffer = 400ms
Total Budget = 2 Seconds
Performance budgets help teams understand how much latency each component is allowed to consume.
Capacity Headroom
Architects rarely design systems to operate at maximum capacity.
Suppose the platform experiences:
Designing for exactly 580 requests per second leaves little room for unexpected events.
A more realistic approach may be:
Headroom accommodates growth, seasonal spikes, operational failures, and changing business needs.
Queuing Concepts
A performance problem often appears when requests arrive faster than they can be processed.
Example:
Processing Capacity = 80 Per Second
The result is a growing queue.
↓
Latency Growth
↓
Timeouts & Failures
This principle appears repeatedly in large-scale distributed systems.
Bottleneck Analysis
A system performs only as well as its slowest critical component.
Consider:
Business Logic = 100ms
Database = 700ms
The database dominates overall response time.
Even if engineers optimize the API from 50ms to 20ms, user experience may barely improve.
Measure first, identify bottlenecks, then optimize.
Setting Realistic Performance Goals
Poor performance goals are often vague and impossible to validate.
Example:
Better:
Strong performance goals are:
- Business Relevant
- Specific
- Measurable
- Testable
Technology should support the user experience rather than define it.
Performance Benchmarking
Performance assumptions should be validated through testing.
Typical benchmarks include:
- Response Time
- Throughput
- Error Rate
- CPU Utilization
- Memory Utilization
Example:
500 Concurrent Users
5,000 Concurrent Users
Benchmarking reveals how performance changes as workload increases.
Performance Review Checklist
A structured review helps teams identify potential issues before production.
✅ User Volume Estimated
✅ Peak Traffic Identified
✅ Growth Expectations Documented
Capacity
✅ Peak RPS Calculated
✅ Capacity Headroom Defined
✅ Storage Growth Estimated
Performance
✅ Response Time Goals Defined
✅ Latency Budget Established
✅ Potential Bottlenecks Identified
Validation
✅ Load Testing Planned
✅ Monitoring Strategy Defined
✅ Benchmark Criteria Established
Performance Troubleshooting Workflow
Architects typically investigate performance issues using a structured approach.
↓
Measure
↓
Identify Bottleneck
↓
Optimize
↓
Benchmark
↓
Monitor
This workflow prevents teams from optimizing components that are not contributing significantly to delays.
Real-World Case Study
Consider the diagnostic search platform.
Goal:
Observed Production Metrics:
P95 Response Time = 3.2 Seconds
Latency Breakdown:
Application Service = 120ms
Database = 2.8 Seconds
Bottleneck:
Actions Taken:
- Query Optimization
- Index Improvements
- Read Replica Introduction
Result:
The lesson is simple: measurement should drive optimization.
Capacity & Performance Canvas
| Area | Example |
|---|---|
| User Volume | 20,000 Clinicians |
| Transactions Per Day | 5 Million Searches |
| Average RPS | 58 |
| Peak RPS | 580 |
| Response Goal | < 2 Seconds |
| Availability Goal | 99.95% |
| Bottleneck | Database |
| Capacity Headroom | ~70% Target Utilization |
| Validation Method | Load Testing |
This artifact is useful during architecture reviews, capacity planning sessions, and system design interviews.
Common Anti-Patterns
| Anti-Pattern | Better Approach |
|---|---|
| Design For Average Load | Design For Peak Demand |
| Optimize Before Measuring | Measure First |
| No Performance Goals | Define Measurable Targets |
| Ignore Bottlenecks | Identify Constraints Early |
| Assume Unlimited Cloud Capacity | Perform Capacity Planning |
How Concepts Connect
↓
Capacity Planning
↓
Throughput Modeling
↓
Latency Analysis
↓
Bottleneck Identification
↓
Performance Goals
↓
Benchmarking
↓
Monitoring & Continuous Improvement
Key Takeaway
They are about ensuring systems can meet business expectations under realistic workloads while maintaining acceptable response times, reliability, and cost.
Successful architects estimate demand, measure behavior, identify bottlenecks, plan headroom, validate assumptions, and continuously monitor outcomes rather than relying on intuition alone.