The Complete Backup Strategy for a Solo-Run VPS: Five Layers, One Runbook
I have written about backups on this site more than once, restic and R2, monitoring with Healthchecks, what happens when Cloudflare itself has an outage, certificates quietly getting shorter. Each of those posts solves one real problem. None of them, on its own, is actually a backup strategy. A backup strategy is what happens when all of those pieces work together, and more importantly, what happens the day one of your servers is simply gone and you have to prove the whole thing actually works under pressure, not just in theory.
This is the piece that ties it together. I am going to walk through the full stack as five layers, then give you the actual sequence to follow the day something breaks, since that is the part nobody writes down until they need it and it is too late.
Why one backup destination is never actually enough
A single copy of your data in a single place is not a backup. It is a bet that nothing happens to that one place before you need it back. I would treat that as true no matter how reliable the provider is. Here is Cloudflare's own reminder of why a cheap, reliable-sounding destination like R2 still needs a real strategy wrapped around it, straight from when they launched it.
Announcing Cloudflare R2 Storage: Rapid and Reliable Object Storage, minus the egress fees. https://t.co/U2dVhVSb2U #BirthdayWeek🎂
— Cloudflare (@Cloudflare) September 28, 2021
Rapid and reliable is genuinely true most of the time. Most of the time is not the same as always, which is exactly what we found when we looked at Cloudflare's own outage history and what it means for anyone relying on R2 alone. The fix is not to avoid R2. It is to make sure R2 is one layer in a stack, not the entire stack.
The five layers of an actual backup strategy
Layer 1: Decide what actually needs to survive
Before any tool matters, know what you are protecting. For a typical Laravel or WordPress app on a single VPS, that is your application files, your uploaded media, and a fresh database dump taken at backup time, not a live database file copied mid-write. Everything else, cache directories, compiled assets, log files, can be rebuilt or regenerated and does not need to survive a disaster at all. Backing up things you can regenerate just wastes storage and slows down every restore you will ever run.
Layer 2: Where it actually goes
This is the part most people think of as "the backup," and it is genuinely the foundation. We covered the full setup in detail in our guide to backing up a VPS to Cloudflare R2 with restic, including the real monthly cost for a typical small app, usually under a dollar. Restic itself is worth understanding beyond just the commands. Here is a talk from Alexander Neumann, restic's own creator, explaining the actual design problems restic was built to solve, deduplication, encryption, and trusting a backup on infrastructure you do not fully control.
Layer 3: Proof it actually ran
A backup script that fails silently is worse than having no backup script at all, since it gives you false confidence right up until the moment you actually need to restore. We built a self-hosted dead man's switch for exactly this in our guide to monitoring restic backups with self-hosted Healthchecks. The one rule from that post worth repeating here: never host the monitor on the same server it is watching, since a total server failure would silently take down your alerting along with your backup.
Layer 4: Protect against the provider itself having a bad day
Even a well-run backup pointed at a single provider is exposed to that provider's own outages. This is not hypothetical. We walked through Cloudflare's actual 2025 and 2026 outage history in our post on why your R2 backups need a second, independent copy, along with the exact restic commands to run a second backup to a completely separate provider. The real cost of this layer is a few extra dollars a month. The real cost of skipping it is discovering, during an actual outage, that your only copy is the one you cannot reach.
Layer 5: The front door has to survive too
A perfect backup does not help if nobody can actually reach your restored site over HTTPS. Certificate lifetimes are shrinking across the industry, and we covered both sides of that in our posts on Let's Encrypt's move to shorter certificates and the rate limit mistake that shorter renewal cycles make more common. A disaster recovery plan that forgets certificates ends with a fully restored server that visitors' browsers refuse to trust.
The actual sequence: what to do the day a VPS is gone
This is the part I have not written down anywhere else on this site. Here is the order that actually matters when a server is well and truly gone, not just misbehaving.
- Provision a new VPS first, before touching anything else. Get the base server running and reachable over SSH before you think about data at all.
- Install restic and restore the latest snapshot. Point it at your primary repository first. If that provider is the one having the outage, this is exactly when your second, independent copy from Layer 4 earns its keep.
- Restore the database from the dump inside that snapshot, not from a live copy. Import it fresh rather than trying to reconstruct a running database state.
- Re-point DNS to the new server's IP address. This step has a real, unavoidable delay while DNS propagates, so start it as early as you safely can rather than leaving it until everything else is done.
- Issue a fresh TLS certificate on the new server. Do not assume an old certificate or key material survived the disaster in any usable form. Start clean.
- Verify the actual application works, not just that the server responds. Log in, submit a real form, check that uploaded media is actually there. A 200 response on the homepage is not the same as a working app.
- Re-point your backup monitoring at the new server. The old check is now watching a server that no longer exists, which means a second failure could go unnoticed right after you have already recovered from the first one.
Step 7 is the one people skip most often, and I think it is the most costly one to skip, since it leaves you blind again immediately after you just proved why being blind is dangerous.
What this entire stack actually costs
| Component | Typical monthly cost | Covered in |
|---|---|---|
| Primary backup storage (R2) | Under $0.50 for a typical small app | restic + R2 guide |
| Second, independent backup copy | $1 to $3, depending on provider and size | Outage resilience guide |
| Self-hosted monitoring (separate small VPS) | $4 to $6 if not already running one elsewhere | Healthchecks guide |
| TLS certificates | $0, Let's Encrypt is free | Certificate lifetime guides |
A genuinely resilient backup strategy for a small, solo-run app lands somewhere between $2 and $10 a month beyond your existing VPS cost, depending mainly on whether you already have a second small server that can double as your monitoring host.
A quick self-audit
Answer these honestly for your own current setup.
- Does your backup destination have a second, independent copy on a different provider, or is one outage away from being your only copy?
- If your backup script silently stopped running tomorrow, would anything actually tell you, or would you only find out during a real restore?
- Have you ever actually run a full restore, start to finish, or only ever confirmed the backup step completes?
- Is your certificate renewal automation tuned for its actual current lifetime, or still assuming a 90-day buffer from years ago?
If any of those made you pause, that is the specific layer to fix next, not a reason to rebuild everything at once.
FAQ
Do I really need all five layers for a small personal project?
For something low-stakes, layers 1 through 3 cover the real risk, having a backup at all and knowing it actually ran. Layers 4 and 5 matter more as the cost of downtime or data loss grows, which is a judgment call only you can make for your specific project.
How often should I actually test a full restore, not just a backup?
Monthly is a reasonable habit for a small, actively used project. The goal is not perfection, it is making sure the first time you ever run a full restore is not during an actual emergency.
What is the single most common mistake in a self-hosted backup setup?
Treating a successful backup command as proof the whole system works. A backup that runs is not the same as a backup you can actually restore from, which is exactly why layers 3 and the restore drill in the self-audit both exist.
Should I automate the entire disaster recovery sequence with a script?
For the pieces that are safe to automate, like restic restore and DNS updates through your provider's API, yes, that reduces the chance of a mistake under pressure. I would still walk through the full sequence manually at least once yourself first, so you actually understand what the script is doing on your behalf.
Does using a managed database service change any of this?
It removes layer 1's database dump step from your own responsibility, since the managed provider handles that backup layer for you, but layers 2 through 5 for your application files, monitoring, provider resilience, and certificates still apply exactly the same way.
Is a second backup provider really necessary if my primary provider has a good uptime record?
A good historical uptime record describes the past, not a guarantee about the exact moment you need to restore something. The cost of a second copy is small enough that I would not skip it based on a provider's track record alone.
Bottom line
A real backup strategy is not one tool. It is five separate, boring decisions working together, what to capture, where it goes, proof it ran, a second copy in case the first provider fails, and a front door that still works when you need it. None of these pieces is hard on its own. The failure mode is always the same, treating any single layer as if it were the whole plan, and finding out otherwise on the one day it actually matters.
Sources: restic official website; restic official documentation; Cloudflare R2 official documentation; Healthchecks official documentation; Let's Encrypt official documentation. Compiled from Cloudhim's own prior reporting and verified official sources on September 24, 2026.
Comments 0
Be the first to comment.
Leave a comment