← all writing

My Postgres backups silently weren't backing up

A backup you never test is not a backup. A backup that fails silently while reporting success is worse than no backup, because the no-backup version at least gives you the dignity of knowing what you don't have.

I set up Postgres backups in my homelab cluster two weeks ago. CloudNativePG running a Postgres 16 cluster, MinIO running on the same cluster as an S3-compatible object store, the standard backup primitive (a ScheduledBackup CRD) wired up to point at the bucket. The cluster came up. The bucket existed. The credentials worked. The first scheduled backup fired right on schedule.

And nothing landed in the bucket.

The cluster reported "continuous archiving working" as true. Then, twenty minutes later, false. Then back to true. The backup status said running and stayed there. No errors in any of the dashboards I was watching. If I'd trusted the green-vs-red coloring I'd never have looked at the pod logs.

I looked at the pod logs.

#The actual error

About forty lines into the Postgres pod's log, in among normal startup chatter, this:

ERROR: Barman cloud WAL archiver exception:
An error occurred (405) when calling the CreateBucket operation:
Method Not Allowed

HTTP 405 from MinIO. Method Not Allowed. On a CreateBucket call. To a bucket that, I'd checked twice, definitely existed.

A 405 from any S3-shaped service is unusual enough to be a clue. It almost always means: you sent a request the server didn't recognize as valid for the resource you sent it to. Not 403 (forbidden), not 404 (no such thing), not 409 (conflict). The server understood you, didn't recognize what you were asking, and got polite about it.

Which raised the obvious question: why was the backup tool trying to create a bucket I'd already created?

#What barman-cloud actually does

CNPG's backup machinery, under the hood, calls a tool called barman-cloud-backup for full backups and barman-cloud-wal-archive for streaming write-ahead-log files. Both tools are AWS-S3-aware Python scripts. Neither one assumes the destination bucket exists. So both of them call CreateBucket on every invocation, just to be sure.

In AWS, CreateBucket on a bucket you already own returns HTTP 200 with a body that essentially says "yes, you already have this." The tool reads that as success and moves on. No problem.

In MinIO, CreateBucket on a bucket that exists but is owned by a different account than the one the request came in on returns 405 Method Not Allowed. The bucket exists. The caller can't create it because someone else already did. That's apparently the most polite thing MinIO has to say about the situation.

In my setup, the bucket was owned by the MinIO root user (which I'd used during initial setup), and the actual caller was a service account I'd created later, with a deliberately narrow policy. The narrow policy was the problem.

#The least-privilege trap

Here's the policy I'd given the service account, ham-fisted-but-deliberate:

{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Action": ["s3:*"],
    "Resource": [
      "arn:aws:s3:::backups",
      "arn:aws:s3:::backups/*"
    ]
  }]
}

Translated: "this account can do anything to the backups bucket or anything inside it, and nothing else anywhere else." I felt good about that policy when I wrote it. Minimum access. Crisp. Resource-scoped.

It is also subtly wrong, because CreateBucket has a quirk: even though the resource is the bucket itself, the action is service-level, not bucket-level. To call CreateBucket, you need permission to enumerate buckets at all (s3:ListAllMyBuckets) and to ask "where is this bucket located" (s3:GetBucketLocation) at the service level (Resource: "arn:aws:s3:::*"), not the bucket level.

My policy gave the service account every action on the backups bucket. It did not give the service account permission to find out the backups bucket existed before trying to create it. So MinIO's view of the call was: an unprivileged account is asking to create backups; the account has no way to know backups exists; MinIO is not going to tell it backups exists either; the request is malformed for the resource; 405.

#The fix

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "ServiceLevel",
      "Effect": "Allow",
      "Action": [
        "s3:ListAllMyBuckets",
        "s3:GetBucketLocation",
        "s3:CreateBucket"
      ],
      "Resource": "arn:aws:s3:::*"
    },
    {
      "Sid": "BucketLevel",
      "Effect": "Allow",
      "Action": ["s3:*"],
      "Resource": [
        "arn:aws:s3:::backups",
        "arn:aws:s3:::backups/*"
      ]
    }
  ]
}

Two statements instead of one. The first one grants three specific service-level actions across all buckets (which sounds scary until you remember the account still can't read or write to any bucket without the second statement). The second one is the original bucket-scoped grant.

After applying the new policy, the next WAL archive attempt succeeded. barman-cloud-wal-archive called CreateBucket, MinIO replied "you already own it," the tool shrugged and moved on to actually archiving the WAL file. The first full backup landed about thirty seconds later. The cluster's "last backup succeeded" status flipped from false to true and stayed there.

#What I'd do differently

The signal of "wrong thing is happening" was buried four log-levels deep in the Postgres pod. The dashboard above it said "running." If I hadn't reflexively checked logs I'd have lived with this for weeks before noticing.

So now the cluster has an alert: if the most recent successful backup is more than two hours old, page me. That's not a fix for the original bug. It's a fix for the pretending-to-be-fine part of the bug, which is the part that actually scared me. The single biggest thing I learned out of this whole adventure isn't about IAM policies or barman-cloud or MinIO. It's that "I set it up, it didn't error, I moved on" is the way you end up with a backup story that doesn't exist.

The smaller thing I learned is about least-privilege policies in AWS-shaped IAM systems. Action and Resource are not always the same axis they appear to be. Some actions have an implicit service-level component that no amount of resource scoping can replace. CreateBucket, ListAllMyBuckets, GetBucketLocation, and a few others all sit in that category. If your "minimum privileges" policy doesn't acknowledge them, you'll get errors that look like authorization failures but read like protocol failures, because that's what they technically are. The server understood you. The server can't tell you what's wrong without revealing things you don't have permission to know.

Now I know. And every IAM policy I write from here on gets a quick pass through "does anything in this list require service-level access I didn't grant?" before I ship it.

Even, especially, the ones I write at home.


← all writing