Backup and Restore
What a backup of the API Service has to contain, how long to keep it, and the order a restore has to happen in.
Deliberately tool-agnostic: whatever takes your Postgres backups today takes these. What follows is what has to be in the set and what the service assumes about it.
What has to be backed up
| What | Why it cannot be reconstructed |
|---|---|
The plugins table | The registry itself: which plugin versions exist, their config, their command line, their recorded checksums. Nothing else holds this. |
The audit_log partitions | The audit trail. Enterprise sells it; it exists nowhere else and there is no second copy. |
The goose_db_version table | Which migrations have run. Restoring data without it makes the next startup either re-run migrations or refuse to start. |
| The object storage bucket | Plugin archives. The local cache is a cache — it is evicted — so the bucket is the only copy of the binaries. |
DB_POSTGRES_DSN, AUTH_WRITE_TOKENS, S3 credentials | Secrets are not in the database and are usually not in the cluster backup either. |
LICENSE_KEY and LICENSE_PUBLIC_KEYS | Recoverable from the licence registry, but not by you at 3am. Without them the restored service runs in community mode: no audit, four workers, ten plugins. |
The database and the bucket have to be recoverable to roughly the same point.
They are not transactional with each other, and the direction of the skew is what
matters: a plugins row without its archive fails generation with
UNAVAILABLE / STORAGE_UNAVAILABLE — the download fails — and the log names
the missing key. An archive with no row is merely orphaned. So prefer a bucket snapshot slightly newer than the database
one — never older.
"Newer" is safe only while archives are never overwritten. If an archive was
re-pushed with --force between the two snapshots, the newer bucket holds an
object whose sha256 is not the one the restored row recorded, and that plugin
fails with plugin archive checksum mismatch on its first download. Two ways to
rule it out: never overwrite a registered version (publish a new one instead),
or enable object versioning on the bucket and restore the object version that
existed at the database snapshot's time.
How much history to keep
config.audit.retentionMonths (12 by default) is enforced by dropping whole
partitions on a schedule. That is a real delete: after it, the only copy of that
month is in a backup taken while it still existed.
So the backup retention has to exceed the audit retention, not match it. Matching
them means the month drops out of the database and out of the archive at
approximately the same time, which leaves nothing anywhere. If you are asked to
keep audit history for a compliance window, that window is a constraint on the
backups, and retentionMonths is only the size of the working set.
An RPO of hours is fine for plugins — plugin registration is a deliberate,
infrequent act and is easily repeated. For audit_log the RPO is the size of
the gap in the audit trail, so it should be measured in minutes if audit is
being relied on.
Restoring
-
Restore the database first, then the bucket. The reverse order leaves the service briefly able to serve plugins it has no rows for.
-
Do not run migrations by hand. The service applies them itself at startup and serialises that across replicas. Start one replica and let it.
-
Check
goose_db_versionagainst the binary you are restoring onto. A database restored from a newer release than the binary starts anyway — this is not caught. goose objects only to migrations missing below the database's highest version; one it has never heard of, above that mark, leaves it with nothing to apply and no complaint. Pinned byTestRollbackOntoAnOlderBinary.So the limit on running an older binary is not a startup check, it is time. The older binary has no partition maintainer for anything migration
00002introduced, so once the months that migration pre-created are used up, audit rows land inaudit_log_default— and a non-empty default partition blocks creating the month that would overlap it, which is a problem you inherit on the way forward. Treat a rollback as bounded byconfig.audit.preCreateMonths(three by default), and prefer restoring onto the matching release. -
Expect the plugin cache to be empty. It is a cache, and with
persistence.enabled=falseit is empty after every restart anyway. The first request for each plugin re-downloads it. Watcheasyp_plugin_cache_bytesclimb; nothing needs doing. -
Verify audit continuity before declaring done. Partition maintenance runs every
config.audit.partitionCheckInterval, so a restore landing in a month whose partition was never created writes into the default partition —easyp_audit_default_partition_usedgoes to 1 and stays there. See Runbooks.
Verifying
A backup that has never been restored is a hypothesis. The cheap version of the
test: restore into a scratch namespace, start the service against it, and check
that easyp_business_plugins_total matches what production reports and that the
newest audit_log row is inside the RPO you think you have.