Runbooks
One procedure per alert the API Service ships: what the alert means, what is actually broken, and what to do about it.
One section per alert in deploy/charts/easyp-service/templates/prometheusrule.yaml
and in deploy/observability/mimir/rules/anonymous/. Each alert's runbook_url
annotation points at the matching anchor here, so the heading text is
load-bearing: renaming one breaks the link from the alert.
Most of what follows is environment-independent: configuration keys, PromQL, and what a number means. Where a procedure needs to read logs it is written twice, because the two deployments this service has are a Kubernetes release and a compose stack, and only one of them has ever run in anger:
| Kubernetes | kubectl in the release's namespace, $REL as the release name |
| compose | ssh to the host, docker against the container |
The host section at the end is compose-only by nature — it is about the machine rather than the service, and on Kubernetes the cluster owns that ground.
EasypLicenceExpiringSoon
The licence stops being valid within prometheusRule.licenceExpiryWarningDays.
Nothing breaks at expiry — the token also carries a grace period — but at the end of that, the tier drops to community: audit stops, workers cap at 4, plugin registration starts refusing past 10. All of it silent.
- Confirm what the service actually loaded, rather than what you think it has:
kubectl logs deploy/$REL | grep -i licen, ordocker logs easyp-api-enterprise 2>&1 | grep -i licen - Check
easyp_license_expiry_timestamp_secondsagainsttime(). - Get a renewed token, put it in the secret as
LICENSE_KEY(on compose, in~/easyp/.env), and restart —kubectl rollout restart deploy/$REL, ordocker compose ... up -d --force-recreate service-enterprise. The licence is read at startup.
If the token is present but ignored, the usual cause is a missing or mismatched
LICENSE_PUBLIC_KEYS entry: a token is only as good as the key it verifies against, and
the key id in the token footer has to be one of the keys configured. See
config.license.publicKeys.
EasypLicenceInGrace
Expiry has passed and the service is running on the grace period the token granted. This is the last warning before the tier drops.
Same procedure as above, without the slack.
EasypGenerationsRejected
Generation requests are being refused with ResourceExhausted (reason
SERVER_OVERLOADED). The service is doing this deliberately — it is not an
error, it is a full queue. The alert watches
easyp_pool_generations_rejected_total: every generation slot was busy and
config.workerPool.queueSize requests were already waiting.
easyp_pool_generations_activeagainstconfig.workerPool.maxConcurrentGenerations, andeasyp_pool_generations_waitingagainstqueueSize. The lookup side (easyp_pool_active_workers,easyp_pool_queue_depth,easyp_pool_rejected_total) refuses with the same reason and is worth a look when cache misses are frequent.- Decide which limit is actually binding:
maxConcurrentGenerations,workers, or the per-callerrateLimit.maxConcurrentPerIP(which shows up ineasyp_concurrency_rejected_total, not here). On Community, generations are capped at 16 and workers at 4 whatever the configuration says. - Raise the binding one — but check the memory arithmetic first. Peak buffers
are
maxConcurrentGenerations × maxOutputSize × 2, and the chart refuses an install where that exceeds the memory limit. Raising concurrency without raising memory turns rejection into OOMKill, which is strictly worse: an OOMKill looks like a crash, and nothing counts it.
If the rejections come from one caller, the per-IP limits are working as intended. Note that CI runners behind one NAT share an address.
EasypGenerationQueueSaturated
Work has been queued continuously for ten minutes. Same knobs as above; the
difference is that this is sustained rather than bursty, so it is a capacity
statement, not a spike. Add capacity rather than a deeper queue — a longer queue
only moves the latency around. Capacity here means a bigger pod: more CPU and a
higher maxConcurrentGenerations (with the memory to match). More replicas on a
shared cache volume are not supported.
EasypPluginCacheAtLimit
Unpacked plugins on disk have reached config.registry.cacheMaxBytes.
This is not by itself a fault: the cache is supposed to reach its limit and
evict. Watch easyp_plugin_cache_evictions_total. Steady eviction with steady
generation latency means it is working.
It matters when eviction churns — the same plugins evicted and re-downloaded
repeatedly, which shows up as generation latency and S3 egress. In that case the
working set does not fit: raise cacheMaxBytes, and the storage behind it. Both
storage paths are checked at install time, so the chart will refuse a
cacheMaxBytes that does not fit its volume.
EasypAuditEventsLost
Audit records are being dropped. Audit is an Enterprise commitment, so any
sustained loss here is a gap in what was promised. The reason label says which
failure this is, and they have nothing in common but the counter:
enqueue_timeout— the writer is not draining fast enough and callers gave up waiting for room. Look ateasyp_audit_queue_depthand at database write latency. Raiseconfig.audit.bufferSizeto absorb bursts;config.audit.enqueueTimeoutto trade request latency for fewer losses.save_failed— writes are being rejected outright, aftermaxSaveRetries. This is a database problem, not a tuning problem. Checkeasyp_audit_save_failures_total, then the database itself. A common cause is no partition for the current month — see EasypAuditMaintenanceStale.shutdown_timeout— a pod stopped before it finished draining. Occasional single-digit losses during a rollout are the queue remainder; anything larger means the shutdown budget is too tight for the queue depth.
EasypAuditMaintenanceStale
Partition maintenance has not succeeded for a day. It runs every
config.audit.partitionCheckInterval (6h by default), so this means several
consecutive failures.
Consequence: future partitions stop being created. Once the newest existing partition is passed, audit rows land in the default partition (see below), and if there is no default either, writes fail outright.
kubectl logs deploy/$REL | grep -i partition, ordocker logs easyp-api-enterprise 2>&1 | grep -i partition— the failure is logged with the reason.- Most often a permissions problem: the role needs to CREATE and DROP tables in the schema, not merely write rows.
- Check
easyp_audit_partition_maintenance_last_success_secondsrecovers after the next interval. Do not wait a full 6h to find out — restart the deployment to force a run.
EasypAuditDefaultPartitionUsed
Audit rows have landed in audit_log_default — the catch-all partition — because
the partition for their month did not exist when they were written.
Nothing is lost. The rows are queryable and complete. Two things are wrong while they sit there:
- the partition for that month cannot be created at all, because its range overlaps rows already in the default;
- Postgres cannot prune partitions for any query filtering on
created_at, so every such query reads the whole audit history. This is what turns a bounded scan into a full one.
Fix the cause first — if maintenance is still failing, draining the default only buys time until the next month. See EasypAuditMaintenanceStale.
Then drain it. Postgres cannot move rows between partitions in place, so the
procedure is detach, create, copy, drop. Run it in a maintenance window: the
detach takes an ACCESS EXCLUSIVE lock on the parent, briefly blocking audit
writes. The audit writer retries, so a short block costs nothing; a long one is
counted as save_failed.
BEGIN;
-- 1. Take the default out of the parent so its rows stop blocking range checks.
ALTER TABLE audit_log DETACH PARTITION audit_log_default;
-- 2. Create the month that was missing. Repeat per month present in the default:
-- SELECT DISTINCT date_trunc('month', created_at) FROM audit_log_default;
CREATE TABLE audit_log_2026_08 PARTITION OF audit_log
FOR VALUES FROM ('2026-08-01') TO ('2026-09-01');
-- 3. Move the rows. They route to the right partition on insert.
INSERT INTO audit_log
SELECT * FROM audit_log_default
WHERE created_at >= '2026-08-01' AND created_at < '2026-09-01';
DELETE FROM audit_log_default
WHERE created_at >= '2026-08-01' AND created_at < '2026-09-01';
-- 4. Reattach only once it is empty. A non-empty default reattaches fine but
-- puts you back where you started.
ALTER TABLE audit_log ATTACH PARTITION audit_log_default DEFAULT;
COMMIT;Confirm easyp_audit_default_partition_used returns to 0 within one scrape.
The metric is SELECT EXISTS(...), not a count, so it flips the moment the last
row leaves — it will not tell you how far through you are. Use
SELECT count(*) FROM audit_log_default for that.
If the default holds enough rows that the transaction above is impractical, do it a month at a time with the detach and reattach as separate transactions.
EasypGenerationErrorRate
More than prometheusRule.generationErrorRatio of generations are failing.
easyp_generation_errors_totalbypluginanderror_type— this is almost always one plugin, not a general fault.error_typeistransient(the error text containedconnection refusedortemporary failureand was retried) orpermanent(everything else).- If the logs say
plugin archive checksum mismatch, the archive in object storage does not match what was recorded at registration. Do not "fix" it by re-registering: find out why it changed. output limit exceededmeans the plugin wrote more thanconfig.registry.maxOutputSize.- Otherwise, run the plugin by hand with the same request. Plugin failures are reported faithfully, so the plugin's own stderr is in the logs.
EasypPanics
The service recovered from a panic. It kept running — that is what the barrier is for — but a recovered panic is still a bug, and the counter exists so that it is not silent.
kubectl logs deploy/$REL | grep -A30 panic for the stack, or
docker logs easyp-api-enterprise 2>&1 | grep -A30 panic. The counter has no
labels; the log line around the stack trace says which unit of work panicked —
a gRPC handler, a pool worker, or the audit writer.
The alert carries a tier label; on the compose stack it names the container,
so read easyp-api-community from it when that is the one that panicked.
EasypAuthFailures
Write credentials are being rejected faster than
prometheusRule.authFailureRate per second.
By default reads are anonymous, so everything here is a mutating call:
CreatePlugin, UpdatePlugin, DeletePlugin. With
config.auth.requireAuthentication on, reads count too — and every easyp CLI
user is a failure, since the CLI sends no token. The reason label separates
no_credentials (nothing sent) from unknown_token (a token that matches no
digest).
- A rollout of CI credentials that did not land is the usual cause — check whether the rate started at a deploy.
AUTH_WRITE_TOKENSmapsname=sha256; the logs name the token that failed, not the secret.- If it is not a known caller, it is someone trying tokens against the registry. The tokens are hashed and the rate limiter applies, but this is worth knowing about.
The host
These six exist only on the compose stack, in
deploy/observability/mimir/rules/anonymous/host.yaml. The disk and memory
alerts come from prometheus.exporter.unix in config.alloy, the target and
discovery alerts from Alloy's own scraping, and the certificate alert from
Traefik. On Kubernetes the cluster's own node monitoring owns this
ground, so the chart does not ship them.
The commands assume ssh user@host and the stack in ~/easyp.
EasypHostDiskLow
A filesystem is under 15% free. $labels.mountpoint says which.
This is a threshold, so by the time it fires the situation already exists. If it
fired before EasypHostDiskFillingUp, the disk filled faster than a day —
look for something writing hard right now rather than something growing slowly.
df -hfor the shape of it, thendocker system df— on this host the answer has been the build cache every time, andACTIVE 0next to a large number means all of it is reclaimable.docker builder prune -afanddocker image prune -f. Reclaiming 3.4 GB this way took the disk from 73% to 49% once already.- If the build cache is not it,
du -xh --max-depth=2 / | sort -h | tail -20. The-xmatters: without it you walk into every container's overlay. - A build cache that keeps coming back means something is building images on
this host. It should be pulling the published image instead:
docker compose pull, not a local build.
EasypHostDiskFillingUp
Six hours of trend say $labels.mountpoint reaches zero within four days.
Nothing is broken. This is the alert that exists to be acted on while it is still cheap, and the one the earlier disk problem would have tripped days before anyone noticed it by hand.
- Same first step:
docker system df. Growth without a cause in the build cache is unusual here. - Check whether it is data rather than waste —
docker system df -vseparates volumes from images. Postgres and the telemetry buckets grow legitimately; the answer for those is retention, not deletion. - Telemetry retention is set per backend: Mimir
168h, Loki168h, Tempo and Pyroscope72h. Those live in object storage, not on this disk, so they are only the answer if the bucket is local.
If the extrapolation looks wrong, it usually is: a filesystem that dipped once and recovered produces a downward line. The alert is guarded to fire only under 40% free for exactly that reason, so a false one means real movement.
EasypHostMemoryLow
Under 10% of memory is available.
Available, not free. Free memory on a healthy Linux box is near zero because the
kernel spends it on page cache, so this is measured with MemAvailable, which
counts what a new allocation could actually get.
docker stats --no-streamfor the split by container.free -h— ifavailableis low whilebuff/cacheis large, the kernel will reclaim it and there is no problem to solve.- Sustained pressure ends at the OOM killer, which chooses by size: postgres or
a plugin process, not whatever caused the pressure.
dmesg -T | grep -i oomsays whether that has already happened.
registry.cache_max_bytes is not the place to look. It bounds unpacked plugins
on disk, not in memory — internal/config/config.go says so and this
section used to say the opposite, which sent the reader to the largest number in
the configuration for a problem it has nothing to do with.
EasypTargetDown
A scrape target exists and does not answer. $labels.job says which; the
service tiers arrive as prometheus.scrape.service, the rest are named after
the component.
The container is still there — that is what distinguishes this from
EasypServiceMissing. So the process is wedged, crashlooping, or no longer
listening on the port.
docker ps -a --filter name=easyp— a restart count climbing is a crashloop;Upwith a failing scrape is a wedge.docker logs --tail 100 <container>. For the service tiers the last line before it stopped answering is usually enough.- Scrape it by hand from inside the network:
docker exec easyp-grafana curl -s http://easyp-api-enterprise:8081/metrics | head. A connection refused and a hang mean different things — refused is a dead listener, a hang is a live process that cannot answer. docker restart <container>clears a wedge and tells you nothing about why. Capture the logs first.
EasypServiceMissing
A whole tier has disappeared from discovery. Not "down" — absent: no container
carries the service=easyp-api-<tier> label that Alloy finds them by, so there
is no target and no up series to be zero.
Two ways to get here, and they need different answers.
- The container is stopped.
docker ps -a --filter label=serviceshows it asExited;docker compose ... up -dbrings it back. - The container is running but lost its label — recreated from an edited
compose file, or started by hand with
docker run. Then the service is serving traffic and invisible to every metric and alert in this file, which is the worse case because nothing else will tell you.docker inspect <container> --format '{{ index .Config.Labels "service" }}'returns empty when this is what happened.
EasypCertificateExpiringSoon
The edge certificate expires within 21 days and Traefik has not replaced it.
Traefik attempts renewal at 30 days, so reaching 21 means an attempt has already failed. Note that on this stand issuance was proven once, at setup, and renewal has never run — the first attempt is due around 7 October 2026.
docker logs easyp-traefik 2>&1 | grep -i acme | tail -30. The reason is almost always the HTTP-01 challenge.- HTTP-01 needs port 80 reachable from the internet and answered by Traefik.
curl -sI http://<domain>/.well-known/acme-challenge/probefrom outside should reach Traefik rather than time out or hit something else. docker exec easyp-traefik cat /acme.json | head -c 200— an empty or truncated store means the volume was lost and the certificate will be requested fresh, which is fine but rate-limited by Let's Encrypt.- Rate limits are the one failure that waiting fixes: five failures per account per hostname per hour. Read the log before retrying.