EasyP

Observability

Metrics, traces, profiles, logs, health endpoints and the alert rules, with a runbook for each alert.

Endpoints

Port (binary / chart)PathContent
23411 / 8081/metricsPrometheus exposition. Always on.
23412 / 8082/liveLiveness. 200 once the process is up; checks nothing else.
23412 / 8082/Readiness. 200 while PostgreSQL answers, 503 otherwise and during start.

Point liveness probes at /live and readiness at /. Using readiness for liveness restarts the pod on every database outage.

Metrics

All names carry the easyp_ prefix. Most are registered at start; those marked lazy appear after their first event, and the cache metrics only when object storage is configured.

Requests

MetricTypeLabelsMeaning
easyp_api_grpc_server_started_totalcountergrpc_service, grpc_method, grpc_typeRPCs started.
easyp_api_grpc_server_handled_totalcounter+ grpc_codeRPCs completed, by status code.
easyp_api_grpc_server_handling_secondshistogramgrpc_service, grpc_method, grpc_typeRPC latency.
easyp_api_grpc_server_msg_received_total, _msg_sent_totalcountersameMessages.
easyp_operations_totalcounteroperation, statusEvery core operation, success or failure, on both tiers.
easyp_rate_limit_requests_totalcounterstatusRate-limiter decisions.
easyp_rate_limit_active_clientsgauge—Tracked client buckets.
easyp_concurrency_rejected_totalcounter—Refused by the per-client concurrency limit.
easyp_concurrency_active_clientsgauge—Clients with requests in flight.
easyp_auth_failures_total (lazy)counterreason = no_credentials | unknown_tokenRejected credentials.

Generation

MetricTypeLabelsMeaning
easyp_generated_plugin_code_totalcounterpluginSuccessful generations.
easyp_generation_duration_secondshistogrampluginPlugin run time, retries included.
easyp_generation_errors_total (lazy)counterplugin, error_type = transient | permanentFailed generations (not counting callers that went away).
easyp_generation_retries_total (lazy)counterpluginServer-side retries.
easyp_pool_generations_activegauge—Plugin processes running.
easyp_pool_generations_waitinggauge—Requests waiting for a generation slot.
easyp_pool_generations_rejected_totalcounter—Refused: generation queue full.
easyp_pool_active_workersgauge—Lookup workers busy.
easyp_pool_queue_depthgauge—Lookup jobs queued.
easyp_pool_jobs_totalcounter—Lookup jobs accepted.
easyp_pool_rejected_totalcounter—Refused: lookup queue full.

Plugin cache (object storage mode)

MetricTypeMeaning
easyp_plugin_cache_bytesgaugeUnpacked plugins on disk.
easyp_plugin_cache_limit_bytesgaugeregistry.cache_max_bytes.
easyp_plugin_cache_evictions_totalcounterEvicted version directories.

Licence

MetricTypeLabelsMeaning
easyp_license_validgauge—1 when a valid licence is loaded.
easyp_license_expiry_timestamp_secondsgauge—Expiry; 0 without a licence.
easyp_license_in_gracegauge—1 during the grace period.
easyp_license_feature_denied_totalcounterfeatureEnterprise features refused on Community.
easyp_config_service_tier_mismatchgauge—1 when telemetry.service_tier disagrees with the licence.

Audit

MetricTypeLabelsMeaning
easyp_audit_queue_depthgauge—Entries waiting for the writer.
easyp_audit_batch_sizehistogram—Rows per write.
easyp_audit_events_lost_totalcounterreason = enqueue_timeout | save_failed | shutdown_timeoutDropped entries.
easyp_audit_events_skipped_totalcounter—Not written because the licence has no audit.
easyp_audit_save_failures_totalcounter—Failed batch writes.
easyp_audit_partitions_currentgauge—Monthly partitions.
easyp_audit_partitions_created_total, _dropped_totalcounter—Partition maintenance.
easyp_audit_partition_maintenance_runs_totalcounterresultMaintenance runs.
easyp_audit_partition_maintenance_last_success_secondsgauge—Last successful run.
easyp_audit_default_partition_usedgauge—1 when rows sit in audit_log_default.

Database and inventory

MetricTypeLabelsMeaning
easyp_repo_call_duration_secondshistogramfuncQuery latency.
easyp_repo_errors_totalcounterfuncQuery errors.
easyp_db_open_connections, _idle_connectionsgauge—Pool state.
easyp_db_wait_count_total, _wait_duration_seconds_totalcounter—Waits for a connection.
easyp_business_plugins_totalgauge—Registered plugin versions.
easyp_business_plugins_by_groupgaugegroupVersions per group.
easyp_business_plugin_versions_countgaugegroup, nameVersions per plugin.
easyp_business_audit_log_totalgauge—Audit rows.
easyp_panics_totalcounter—Recovered panics, in handlers and background work.

Plus the standard go_* and process_* collectors.

Traces

Set telemetry.otlp_endpoint (or OTEL_EXPORTER_OTLP_ENDPOINT) to an OTLP gRPC collector. The core, the registry and each plugin run are wrapped in spans; pool.Get records the time spent queued. Empty endpoint: no exporter is built and nothing is attempted.

Profiles

Set telemetry.pyroscope_endpoint for continuous profiling with Pyroscope.

Logs

Structured JSON on stdout. Every start logs configuration resolved with the settings that differ from defaults, secrets redacted. Worth alerting on in your log system: gRPC server is running WITHOUT TLS, lowered to the licence tier's limit, licence expired, audit enqueue timed out.

Alerts

The Helm chart ships a PrometheusRule (prometheusRule.enabled=true, needs the Prometheus Operator CRDs). Every alert carries a runbook_url pointing at prometheusRule.runbookBaseUrl plus an anchor named after the alert.

AlertFires whenRunbook
EasypLicenceExpiringSoonLicence expires within licenceExpiryWarningDays (14)link
EasypLicenceInGraceeasyp_license_in_grace == 1 for 15mlink
EasypGenerationsRejectedrate(easyp_pool_generations_rejected_total[5m]) > 0 for 5mlink
EasypGenerationQueueSaturatedeasyp_pool_generations_waiting > 0 for 10mlink
EasypPluginCacheAtLimitcache bytes / limit > cacheUsageWarningRatio (0.95) for 30mlink
EasypAuditEventsLostrate(easyp_audit_events_lost_total[5m]) > 0 for 5mlink
EasypAuditMaintenanceStaleno successful partition maintenance for 24hlink
EasypAuditDefaultPartitionUsedeasyp_audit_default_partition_used == 1 for 15mlink
EasypGenerationErrorRateerror ratio > generationErrorRatio (0.05) for 10mlink
EasypPanicsincrease(easyp_panics_total[5m]) > 0link
EasypAuthFailuresrate(easyp_auth_failures_total[5m]) > authFailureRate (0.5) for 10mlink

The compose stack adds host and edge alerts, evaluated by its own Mimir from what Grafana Alloy collects — EasypHostDiskLow, EasypHostDiskFillingUp, EasypHostMemoryLow, EasypTargetDown, EasypServiceMissing, EasypCertificateExpiringSoon — in deploy/observability/mimir/rules/anonymous/host.yaml. They are also in Runbooks.

ServiceMonitor (serviceMonitor.enabled=true) scrapes /metrics every 30 s.

On this page