ka2a documentation

Metrics reference

Developer preview. Not yet production ready. This page is rendered from docs/metrics-reference.md of the ka2a repository at revision 96fb45e5e5f6693e779837998c3aa9581d6a37f2. It describes the behavior of that revision.

Contents

ka2a run --metrics-listen ADDRESS serves the metrics of the node at http://ADDRESS/metrics in the Prometheus text format, version 0.0.4. A library user mounts Node.MetricsHandler() (see ka2a.WritePrometheus). The listener accepts only loopback addresses, so scrape it from the same host, for example with a local Prometheus agent or a sidecar.

Each scrape reads the node status once. A scrape changes no state and starts no work. A failed status read answers HTTP 503 with the error code. A node that does not run serves no metrics: alert on the scrape itself (up == 0) in addition to the rules below.

A test (TestMetricsReferenceMatchesTheExposition, pkg/ka2a) keeps this reference complete: every family of the exposition has one row here, with its type, its labels and every fixed label value. A new, renamed or removed family fails the test until this reference follows.

Label policy

  • Only ka2a_node_info has free labels: domain and endpoint of the local endpoint, from the configuration.
  • Every other label (state, result) has a fixed set of values. Every value of the set appears in each scrape. A one-hot family has the value 1 for the current value and 0 for the other values.
  • No label holds an operation ID, a context ID, a run ID, a task ID, a message body, a key ID, a topic, a path or a broker address.
  • A timestamp gauge is in Unix seconds with millisecond precision. The value 0 means "never" or "not applicable".
  • A counter starts at 0 when the process starts. It is not persistent. Use rate() or increase().

Node and store

MetricTypeLabelsMeaningAlert on
ka2a_node_infogaugedomain, endpoint (free)The local endpoint of the node. The value is always 1.Join it to name the endpoint in an alert.
ka2a_node_runninggaugenone1 while Run works.0 for a node that must run.
ka2a_node_readygaugenone1 after the start of Run finished (the ready event). The node then accepts work.0 for more than 5 minutes while ka2a_node_running is 1.
ka2a_node_closedgaugenone1 after Close started.Information only. A stop is expected during a deployment.
ka2a_quarantine_activegaugenone1 while a restore quarantine holds every handler start.Any 1. Follow the runbook Quarantine after a restore.
ka2a_status_observed_timestamp_secondsgaugenoneTime of the store read of this scrape.time() - value larger than a few scrape intervals: the scrape returns old data.
ka2a_store_retained_bytesgaugenoneRetained bytes of the local store.Above 80% of ka2a_store_budget_bytes. Follow Storage pressure.
ka2a_store_budget_bytesgaugenoneRetained-state budget of the local store (Limits.RetainedBytes).Use it as the denominator of the retained-bytes alert.
ka2a_tasksgaugenoneTasks in the local task state.Trend only.
ka2a_rejectionsgaugenoneRetained rejection records. Retention purges old records.A fast rise: see ka2a_input_rejected_total.

Outbox, exchanges and dispatch

MetricTypeLabelsMeaningAlert on
ka2a_outbox_recordsgaugestate: queued, publishing, held, errorOutbox records by send state. held records wait for an operator; error records failed permanently.state="held" above 0. state="error" rising.
ka2a_outbox_oldest_queued_age_secondsgaugenoneAge of the oldest queued outbox record, or 0.Above 300 seconds: the node cannot publish. Check the broker state.
ka2a_exchangesgaugestate: awaiting, heldSent requests without a final outcome, by state. held: the outcome is unknown and needs an operator.state="held" above 0. Follow Held work.
ka2a_dispatchgaugestate: queued, started, heldReceived requests in the dispatcher, by state. held: the handler outcome is unknown and the node does not run the handler again.state="held" above 0. Follow Held work.

Broker, credentials and catalog

MetricTypeLabelsMeaningAlert on
ka2a_mailbox_checkgaugeresult: not_checked, verified, unverified, broker_unreachable, mismatch (one-hot)Result of the mailbox verification at the start of Run.result="mismatch" is 1. result="broker_unreachable" is 1 for more than 10 minutes.
ka2a_broker_stategaugestate: unknown, reachable, unreachable (one-hot)Live broker connectivity that the Kafka clients of the node observe. unknown: no recent answer and no failed attempt since the last answer.state="unreachable" is 1 for more than 2 minutes. Follow Broker outage.
ka2a_broker_last_contact_timestamp_secondsgaugenoneTime of the last broker answer, or 0.time() - value above 300 seconds on a running node.
ka2a_broker_last_failure_timestamp_secondsgaugenoneTime of the last failed broker contact attempt, or 0. The status names its class and hint (see the error reference).Correlate with ka2a_broker_state.
ka2a_credentials_last_reloadgaugeresult: none, reloaded, refused (one-hot)Result of the last broker credential reload.result="refused" is 1. The node keeps the previous credentials.
ka2a_credentials_last_reload_timestamp_secondsgaugenoneTime of the last broker credential reload, or 0.Information only.
ka2a_credentials_reloads_totalcounterresult: reloaded, refusedBroker credential reloads, by result.increase(...{result="refused"}[15m]) > 0.
ka2a_client_certificate_not_after_timestamp_secondsgaugenoneEnd of the validity of the client certificate in use, or 0 without one.Less than 14 days ahead (and not 0). Follow Refused or expiring client certificate.
ka2a_revocation_lists_next_update_timestamp_secondsgaugenoneEarliest next update of the revocation lists of the broker chain in use, or 0 without lists.Less than 24 hours ahead (and not 0). Follow Revocation list due.
ka2a_catalog_diskgaugestate: same, differs, refused (one-hot)The catalog file on disk compared with the running catalog.state="differs" or state="refused" is 1. Follow Catalog on disk differs from the running digest.

Counters

MetricTypeLabelsMeaningAlert on
ka2a_submit_accepted_totalcounternoneSubmits that admitted a new operation.Trend only.
ka2a_submit_reused_totalcounternoneSubmits that returned an earlier operation with the same key.Trend only. Retries of callers.
ka2a_submit_conflicts_totalcounternoneSubmits refused with identity_conflict.Any increase. Follow identity_conflict.
ka2a_submit_refused_totalcounternoneSubmits refused for another reason, for example storage_pressure or unauthorized.A sustained rate.
ka2a_publish_acknowledged_totalcounternoneRecords that the broker acknowledged.A rate of 0 while ka2a_outbox_records{state="queued"} is above 0.
ka2a_publish_retries_totalcounternonePublish attempts that the publisher scheduled again.A sustained rate: the broker refuses or times out.
ka2a_publish_unknown_totalcounternonePublish attempts with an unknown outcome. The node publishes the same bytes again.A sustained rate.
ka2a_publish_rejected_totalcounternoneRecords that the broker refused for sure. The exchange becomes held.Any increase.
ka2a_publish_expired_totalcounternoneRecords that expired before a publication.Any increase.
ka2a_requests_recorded_totalcounternoneReceived requests that the store recorded.Trend only.
ka2a_request_duplicates_totalcounternoneReceived duplicates of recorded requests. Redelivery after a restart or a rebalance.Trend only.
ka2a_request_conflicts_totalcounternoneReceived requests that reuse an ID for another request.Any increase.
ka2a_input_rejected_totalcounternoneReceived records that the node rejected. Each rejection has a reason (see the error reference).A sustained rate: a misconfigured or hostile peer.
ka2a_responses_recorded_totalcounternoneReceived responses that the store recorded.Trend only.
ka2a_response_duplicates_totalcounternoneReceived duplicates of recorded responses.Trend only.
ka2a_responses_rejected_totalcounternoneReceived responses that the node rejected.Any increase.
ka2a_admission_retries_totalcounternoneReceived records whose admission waited and was tried again, for example for storage pressure.A sustained rate.
ka2a_checkpoint_commits_totalcounternoneConsumer offset commits that the broker confirmed.A rate of 0 while records arrive.
ka2a_checkpoint_failures_totalcounternoneConsumer offset commits that failed.A sustained rate.
ka2a_consumer_restarts_totalcounternoneRestarts of the consumer after a consumer failure.More than 3 in 15 minutes.
ka2a_dispatch_started_totalcounternoneHandler starts.Trend only.
ka2a_dispatch_completed_totalcounternoneHandler results that the store recorded.Trend only.
ka2a_dispatch_held_totalcounternoneDispatches that the node held.Any increase.
ka2a_dispatch_interrupted_totalcounternoneDispatches that a stop or a deadline interrupted.A rate outside deployments.
ka2a_observer_delivered_totalcounternoneEvents that the observer received.Trend only.
ka2a_observer_dropped_totalcounternoneEvents that a full observer queue dropped. Messaging continues.A sustained rate: the observer is too slow.
ka2a_observer_panics_totalcounternonePanics of the observer.Any increase.

Alert rules

examples/alerts.yml has Prometheus alerting rules for the conditions above. Adapt the thresholds and the for durations to your service levels. A test (TestAlertRulesNameOnlyExposedMetrics, pkg/ka2a) checks that the file parses and that every ka2a_ metric that it names is in the exposition.

All ka2a documents