Infrastructure architecture · HA / DR & resilience
Resilience, HA and DR — designed per workload, not assumed.
Most estates carry one HA/DR posture across every workload, regardless of what each one is actually worth to the business. 1722 designs resilience, high-availability and disaster-recovery architecture workload by workload: RTO/RPO objectives set with the business, a standby and replication pattern to match, DR runbooks, and a test strategy so the plan is proven, not assumed. This is design work, handed over as documents your team or MSP runs.
The architecture layer sits above day-to-day operations. Your team or your MSP holds the environment and the on-call responsibility that goes with it.
The objectives
RTO and RPO: two numbers, set per workload.
Every design decision that follows — standby pattern, replication method, runbook detail — is derived from these two objectives, agreed with the people who own the workload, not assumed by the infrastructure team.
RTO (recovery time objective) is the maximum acceptable downtime for a workload before the business impact becomes unacceptable. RPO (recovery point objective) is the maximum acceptable data loss, measured as a span of time, between the last recoverable copy and the point of failure. A near-zero RTO/RPO order-management system and a same-day-recovery internal reporting tool are not the same engineering problem, and pricing the same standby pattern across both wastes budget on one and leaves the other exposed. The RTO/RPO matrix sets both numbers per workload before any pattern is chosen.
The pattern
Standby patterns, chosen to match the objective.
The RTO/RPO for a workload determines which standby pattern is proportionate — not the other way round.
| Active-active | Multiple live sites serving traffic concurrently. Near-zero RTO/RPO; highest cost and operational complexity. |
|---|---|
| Active-passive (hot standby) | A fully provisioned, continuously replicated standby ready to take traffic. Low RTO/RPO; standing infrastructure cost. |
| Warm standby | Scaled-down standby infrastructure running, scaled up on failover. Moderate RTO; lower steady-state cost. |
| Pilot-light | Core data replicated and minimal infrastructure held ready; the rest is built out on failover. Higher RTO; lowest steady-state cost of the standing-infrastructure patterns. |
| Backup & restore | No standing standby; recovery from backup on failure. Highest RTO/RPO; the cheapest pattern where the objective allows it. |
A common question
HA vs DR — do we need both?
Usually yes, scoped differently. HA is designed for the failures you should expect — a node, a disk, an availability zone — and aims for continuous operation with no visible interruption. DR is designed for the failures you plan for and rarely see — a site, a region or a platform lost outright — and aims for a bounded, tested recovery within the agreed RTO/RPO. A workload can have a strong HA design and still have no credible DR plan, or the reverse; the matrix makes the gap visible per workload rather than assuming one posture fits the whole estate.
What you get
Deliverables you own.
Documents handed over at the end of the engagement, not advice retained in someone’s head.
- RTO/RPO matrix — objectives agreed per critical workload with the business owners who carry the impact
- Standby & replication design — the pattern chosen per workload, with the reasoning and the trade-offs made explicit
- DR runbooks — trigger conditions, recovery sequence, roles and validation steps for each workload in scope
- DR test strategy — what to test, how often, and to what evidence standard, so the plan is proven rather than assumed
- Handover pack — documentation your team or MSP uses to run and exercise the plan without us
Scope
Where the design ends.
The architecture layer sits above day-to-day operations. The engagement produces the RTO/RPO matrix, the standby and replication design, the runbooks and the test strategy — the plan a workload owner and an infrastructure team can act on. Executing that plan — standing up replication, running failover drills, holding on-call responsibility for a live incident — is operational work that stays with your team or your MSP. Where an MSP is engaged for that operational layer, this design is the brief they work from, not a contract they compete with.
Put a number on your recovery objectives.
Tell us which workloads carry the risk. A short discovery call, then a scoped design plan.
FAQ
HA, DR and resilience design, answered.
What is the difference between HA and DR?
High availability is the design that keeps a workload running through a component failure inside a site or region — redundant nodes, load balancing, automated failover, no data loss and little to no downtime. Disaster recovery is the design that recovers a workload after a larger event takes out a whole site, region or platform — a plan to stand the workload up elsewhere within agreed time and data-loss limits. Most estates need both, scoped differently by workload: HA for the failures you expect weekly, DR for the ones you plan for and rarely see.
What are RTO and RPO, in plain terms?
RTO (recovery time objective) is how long a workload can be down before the business impact is unacceptable. RPO (recovery point objective) is how much data loss, measured in time, is acceptable — the gap between the last good copy and the point of failure. Neither is a technical constant: they are business decisions per workload, and the standby pattern, replication method and runbook are all derived from them, not set first.
Do you take on our DR operations once the design is done?
No. This is design work: the RTO/RPO matrix, the standby and replication design, the runbooks and the test strategy, handed over as documents you own. The architecture layer sits above day-to-day operations — your team or your MSP is the one that executes the runbooks and owns the DR test cadence.
Is the design specific to one cloud or vendor?
No. We do not sell infrastructure, DR tooling or replication licences, and we take no vendor commissions, so the standby pattern and replication design are chosen against your RTO/RPO objectives and existing estate — on-premises, single cloud, multi-cloud or hybrid — not against a preferred platform.
How is a DR runbook different from a general operations runbook?
A DR runbook is scoped to one recovery scenario for one workload: the trigger conditions, the failover or recovery sequence, the roles involved, the validation steps and the fallback if a step fails. It is written to be followed under pressure by someone who did not design the system, and it is only credible once it has a test strategy behind it — a runbook that has never been exercised is a draft, not a plan.
How is an engagement scoped?
Discovery starts from your critical workload list and the current backup, replication and failover state. Output is an RTO/RPO matrix agreed with the business owners of each workload, a standby/replication design, DR runbooks for the workloads in scope, and a DR test strategy defining cadence and method. Priced day-rate, fixed-scope or retainer against a written statement of work — discuss scope and workload count on a discovery call.
Do you help us test the DR plan once it is written?
The test strategy — what to test, how often, and to what evidence standard — is part of the design deliverable. Running the tests is operational work for your team or MSP; we can be engaged separately to review results and refine the design in light of what a test surfaces, but execution is not the design engagement.