A multi-cloud setup runs parts of your infrastructure on more than one provider, for example application servers on a VPS provider, backups in an object storage service from another company, and DNS with a third. Done well, it removes single points of failure and keeps you free to move. Done badly, it doubles your operational work without improving availability. This guide explains when multi-cloud is worth it, compares the common patterns, and walks through a concrete active-passive setup with DNS failover and off-provider backups.
Prerequisites
To follow the practical parts of this guide you need:
- Two servers at different providers running the same stateless web application, for example a CubePath VPS as primary and a VM at another provider as secondary. Both should run Ubuntu 24.04 LTS.
- A domain whose DNS you can move to Amazon Route 53, and an AWS account with the AWS CLI v2 configured (
aws configure). - A health endpoint on both servers, such as
/health, that returns HTTP 200 only when the app and its dependencies work. - Basic familiarity with DNS records and TTLs.
When multi-cloud makes sense
Multi-cloud has real costs: two sets of consoles, APIs, network quirks and bills, plus the hard problem of keeping data consistent across providers. It pays off in a few specific situations:
- Your availability target is higher than what one provider can deliver. A full provider or region outage takes everything with it, however many servers you run there.
- You need backups that survive losing a provider account. A suspended account, a billing mistake or a compromised API key can wipe out both production and backups if they live in the same place.
- Regulation or customer contracts require data in specific jurisdictions that no single provider covers.
- You want an exit strategy. Keeping infrastructure portable means you can move when prices or service change.
For most small projects, one well-run provider plus backups stored with a second provider gives most of the benefit at a fraction of the complexity. Start there and add more only when you have a concrete requirement.
Common patterns
| Pattern | How it works | Good for | Main difficulty |
|---|---|---|---|
| Off-provider backups | Production on one provider, backups copied to another | Everyone, as a baseline | Testing restores regularly |
| Active-passive | Primary serves traffic, a warm standby at another provider takes over through DNS failover | Stateless apps, APIs, sites with a replicable database | Keeping the standby's data and config current |
| Active-active | Both providers serve traffic at the same time behind DNS or a global load balancer | High traffic, global audiences | Writes to shared data, conflict handling |
| Tiered (best of breed) | Compute on a VPS, object storage, DNS or CDN from other vendors | Using managed services without moving everything | Egress costs and many dependencies |
The hardest part of every pattern is state. Stateless web servers are easy to duplicate; databases, uploaded files and sessions are not. Before choosing a pattern, answer two questions: how much data can you afford to lose (your recovery point objective) and how long can you be down (your recovery time objective). Asynchronous database replication to another provider typically gives a few seconds of possible data loss; nightly backups give up to a day.
Step 1 - Making the application portable
Failover only works if the secondary server is a real copy of the primary. That is much easier when servers are built from code rather than by hand.
- Keep application config and deployment in Git, and deploy to both servers with the same tool (Ansible, Docker Compose, or a CI pipeline).
- Avoid provider-specific features in the application path. For example, use a standard S3 API client for object storage so you can switch endpoints, and keep secrets in environment variables rather than in a single provider's secret manager.
- Store sessions in a shared backend (Redis, the database) or use signed cookies, so users are not logged out when traffic moves.
Check that both servers serve the same response to the health endpoint before going further:
curl -s -o /dev/null -w "%{http_code}\n" -H "Host: app.your_domain" http://primary_server_ip/health
curl -s -o /dev/null -w "%{http_code}\n" -H "Host: app.your_domain" http://secondary_server_ip/health
200
200
Step 2 - Creating a health check for the primary
DNS failover needs a monitor outside both providers that decides whether the primary is healthy. Route 53 health checks run from several AWS locations and mark an endpoint unhealthy after a number of consecutive failures.
Create a health check against the primary server. Replace primary_server_ip and app.your_domain:
aws route53 create-health-check \
--caller-reference "app-primary-$(date +%s)" \
--health-check-config '{
"IPAddress": "primary_server_ip",
"Port": 443,
"Type": "HTTPS",
"ResourcePath": "/health",
"FullyQualifiedDomainName": "app.your_domain",
"RequestInterval": 30,
"FailureThreshold": 3
}'
The output includes an Id. Save it, you need it in the next step. With these settings the primary is marked unhealthy after about 90 seconds of failures.
Confirm the checkers see it as healthy:
aws route53 get-health-check-status --health-check-id your_health_check_id
Each checker region should report a Success status after a minute or two.
Step 3 - Creating failover DNS records
Route 53 failover routing uses two records with the same name: a PRIMARY tied to the health check and a SECONDARY that is served only when the primary is unhealthy. A low TTL makes clients pick up the change quickly.
Save the change set to a file:
nano failover.json
{
"Changes": [
{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "app.your_domain",
"Type": "A",
"SetIdentifier": "primary",
"Failover": "PRIMARY",
"TTL": 60,
"HealthCheckId": "your_health_check_id",
"ResourceRecords": [{ "Value": "primary_server_ip" }]
}
},
{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "app.your_domain",
"Type": "A",
"SetIdentifier": "secondary",
"Failover": "SECONDARY",
"TTL": 60,
"ResourceRecords": [{ "Value": "secondary_server_ip" }]
}
}
]
}
Apply it to your hosted zone:
aws route53 change-resource-record-sets --hosted-zone-id your_zone_id --change-batch file://failover.json
Verify that the name resolves to the primary:
dig +short app.your_domain
primary_server_ip
Step 4 - Testing the failover
A failover you have never tested is a guess. Stop the web server on the primary (or make /health return an error) during a quiet period:
sudo systemctl stop nginx
After the health check has failed three times and the 60-second TTL has expired, the name should resolve to the secondary:
dig +short app.your_domain
secondary_server_ip
Start the service again with sudo systemctl start nginx. Route 53 switches back automatically once the health check passes. Note how long each direction took; that is your real recovery time for this layer.
Notesome resolvers and clients cache records longer than the TTL. Expect most traffic to move within a few minutes, not all of it instantly.
Step 5 - Keeping backups with a second provider
DNS failover handles an unreachable server. Off-provider backups handle the worse cases: deleted data, ransomware or losing access to an account. Use a tool that encrypts on the client and works with any S3-compatible storage, such as restic, and point it at a bucket at a different company from the one running production.
The key rules:
- Use credentials that can write backups but live in a different account from production, so one compromised key cannot delete both.
- Encrypt before upload; the storage provider should never see plaintext.
- Restore a backup to a scratch server on a schedule. A backup that has never been restored is not a backup.
The Backblaze B2 server backup guide in this Cloud Integration section shows a complete restic setup with a systemd timer that you can reuse with any S3-compatible provider.
Step 6 - Watching every provider from one place
When infrastructure is spread across providers, monitoring each one only from its own console hides problems. Use one monitoring system that checks all servers from outside, for example Prometheus with node_exporter on every host (reached over a private VPN, not the public internet) plus external HTTP checks against the public hostname. Alert on the health endpoint, disk space, certificate expiry and backup age.
Keep an eye on the cost side too: data transfer between providers is usually billed as egress, so replicating large volumes of data across clouds can cost more than the servers themselves.
Which approach should you choose?
- Small site or single app: one provider for production plus encrypted backups at a second provider. Rebuild from code if the provider fails.
- Business-critical stateless service: active-passive with DNS failover, a warm standby at a second provider and database replication or frequent backups to it.
- Large, global service with an engineering team to run it: active-active across providers, with a data layer designed for it from the start.
Conclusion
A useful multi-cloud strategy starts with a clear reason, portable infrastructure and data you can restore anywhere. With a health check, failover records and tested off-provider backups you get most of the resilience benefits without running two full production stacks.
As next steps, automate the deployment of both servers with Ansible or Terraform, set up replication for your database to the standby, and schedule a quarterly failover and restore test.
