ka2a documentation
Metrics reference
Developer preview. Not yet production ready. This page is rendered from docs/metrics-reference.md of the ka2a repository at revision 96fb45e5e5f6693e779837998c3aa9581d6a37f2. It describes the behavior of that revision.
Contents
ka2a run --metrics-listen ADDRESS serves the metrics of the node at
http://ADDRESS/metrics in the Prometheus text format, version 0.0.4. A library user mounts
Node.MetricsHandler() (see ka2a.WritePrometheus). The listener accepts only loopback
addresses, so scrape it from the same host, for example with a local Prometheus agent or a
sidecar.
Each scrape reads the node status once. A scrape changes no state and starts no work. A
failed status read answers HTTP 503 with the error code. A node that does not run serves no
metrics: alert on the scrape itself (up == 0) in addition to the rules below.
A test (TestMetricsReferenceMatchesTheExposition, pkg/ka2a) keeps this reference complete:
every family of the exposition has one row here, with its type, its labels and every fixed
label value. A new, renamed or removed family fails the test until this reference follows.
Label policy
- Only
ka2a_node_infohas free labels:domainandendpointof the local endpoint, from the configuration. - Every other label (
state,result) has a fixed set of values. Every value of the set appears in each scrape. A one-hot family has the value 1 for the current value and 0 for the other values. - No label holds an operation ID, a context ID, a run ID, a task ID, a message body, a key ID, a topic, a path or a broker address.
- A timestamp gauge is in Unix seconds with millisecond precision. The value 0 means "never" or "not applicable".
- A counter starts at 0 when the process starts. It is not persistent. Use
rate()orincrease().
Node and store
| Metric | Type | Labels | Meaning | Alert on |
|---|---|---|---|---|
ka2a_node_info | gauge | domain, endpoint (free) | The local endpoint of the node. The value is always 1. | Join it to name the endpoint in an alert. |
ka2a_node_running | gauge | none | 1 while Run works. | 0 for a node that must run. |
ka2a_node_ready | gauge | none | 1 after the start of Run finished (the ready event). The node then accepts work. | 0 for more than 5 minutes while ka2a_node_running is 1. |
ka2a_node_closed | gauge | none | 1 after Close started. | Information only. A stop is expected during a deployment. |
ka2a_quarantine_active | gauge | none | 1 while a restore quarantine holds every handler start. | Any 1. Follow the runbook Quarantine after a restore. |
ka2a_status_observed_timestamp_seconds | gauge | none | Time of the store read of this scrape. | time() - value larger than a few scrape intervals: the scrape returns old data. |
ka2a_store_retained_bytes | gauge | none | Retained bytes of the local store. | Above 80% of ka2a_store_budget_bytes. Follow Storage pressure. |
ka2a_store_budget_bytes | gauge | none | Retained-state budget of the local store (Limits.RetainedBytes). | Use it as the denominator of the retained-bytes alert. |
ka2a_tasks | gauge | none | Tasks in the local task state. | Trend only. |
ka2a_rejections | gauge | none | Retained rejection records. Retention purges old records. | A fast rise: see ka2a_input_rejected_total. |
Outbox, exchanges and dispatch
| Metric | Type | Labels | Meaning | Alert on |
|---|---|---|---|---|
ka2a_outbox_records | gauge | state: queued, publishing, held, error | Outbox records by send state. held records wait for an operator; error records failed permanently. | state="held" above 0. state="error" rising. |
ka2a_outbox_oldest_queued_age_seconds | gauge | none | Age of the oldest queued outbox record, or 0. | Above 300 seconds: the node cannot publish. Check the broker state. |
ka2a_exchanges | gauge | state: awaiting, held | Sent requests without a final outcome, by state. held: the outcome is unknown and needs an operator. | state="held" above 0. Follow Held work. |
ka2a_dispatch | gauge | state: queued, started, held | Received requests in the dispatcher, by state. held: the handler outcome is unknown and the node does not run the handler again. | state="held" above 0. Follow Held work. |
Broker, credentials and catalog
| Metric | Type | Labels | Meaning | Alert on |
|---|---|---|---|---|
ka2a_mailbox_check | gauge | result: not_checked, verified, unverified, broker_unreachable, mismatch (one-hot) | Result of the mailbox verification at the start of Run. | result="mismatch" is 1. result="broker_unreachable" is 1 for more than 10 minutes. |
ka2a_broker_state | gauge | state: unknown, reachable, unreachable (one-hot) | Live broker connectivity that the Kafka clients of the node observe. unknown: no recent answer and no failed attempt since the last answer. | state="unreachable" is 1 for more than 2 minutes. Follow Broker outage. |
ka2a_broker_last_contact_timestamp_seconds | gauge | none | Time of the last broker answer, or 0. | time() - value above 300 seconds on a running node. |
ka2a_broker_last_failure_timestamp_seconds | gauge | none | Time of the last failed broker contact attempt, or 0. The status names its class and hint (see the error reference). | Correlate with ka2a_broker_state. |
ka2a_credentials_last_reload | gauge | result: none, reloaded, refused (one-hot) | Result of the last broker credential reload. | result="refused" is 1. The node keeps the previous credentials. |
ka2a_credentials_last_reload_timestamp_seconds | gauge | none | Time of the last broker credential reload, or 0. | Information only. |
ka2a_credentials_reloads_total | counter | result: reloaded, refused | Broker credential reloads, by result. | increase(...{result="refused"}[15m]) > 0. |
ka2a_client_certificate_not_after_timestamp_seconds | gauge | none | End of the validity of the client certificate in use, or 0 without one. | Less than 14 days ahead (and not 0). Follow Refused or expiring client certificate. |
ka2a_revocation_lists_next_update_timestamp_seconds | gauge | none | Earliest next update of the revocation lists of the broker chain in use, or 0 without lists. | Less than 24 hours ahead (and not 0). Follow Revocation list due. |
ka2a_catalog_disk | gauge | state: same, differs, refused (one-hot) | The catalog file on disk compared with the running catalog. | state="differs" or state="refused" is 1. Follow Catalog on disk differs from the running digest. |
Counters
| Metric | Type | Labels | Meaning | Alert on |
|---|---|---|---|---|
ka2a_submit_accepted_total | counter | none | Submits that admitted a new operation. | Trend only. |
ka2a_submit_reused_total | counter | none | Submits that returned an earlier operation with the same key. | Trend only. Retries of callers. |
ka2a_submit_conflicts_total | counter | none | Submits refused with identity_conflict. | Any increase. Follow identity_conflict. |
ka2a_submit_refused_total | counter | none | Submits refused for another reason, for example storage_pressure or unauthorized. | A sustained rate. |
ka2a_publish_acknowledged_total | counter | none | Records that the broker acknowledged. | A rate of 0 while ka2a_outbox_records{state="queued"} is above 0. |
ka2a_publish_retries_total | counter | none | Publish attempts that the publisher scheduled again. | A sustained rate: the broker refuses or times out. |
ka2a_publish_unknown_total | counter | none | Publish attempts with an unknown outcome. The node publishes the same bytes again. | A sustained rate. |
ka2a_publish_rejected_total | counter | none | Records that the broker refused for sure. The exchange becomes held. | Any increase. |
ka2a_publish_expired_total | counter | none | Records that expired before a publication. | Any increase. |
ka2a_requests_recorded_total | counter | none | Received requests that the store recorded. | Trend only. |
ka2a_request_duplicates_total | counter | none | Received duplicates of recorded requests. Redelivery after a restart or a rebalance. | Trend only. |
ka2a_request_conflicts_total | counter | none | Received requests that reuse an ID for another request. | Any increase. |
ka2a_input_rejected_total | counter | none | Received records that the node rejected. Each rejection has a reason (see the error reference). | A sustained rate: a misconfigured or hostile peer. |
ka2a_responses_recorded_total | counter | none | Received responses that the store recorded. | Trend only. |
ka2a_response_duplicates_total | counter | none | Received duplicates of recorded responses. | Trend only. |
ka2a_responses_rejected_total | counter | none | Received responses that the node rejected. | Any increase. |
ka2a_admission_retries_total | counter | none | Received records whose admission waited and was tried again, for example for storage pressure. | A sustained rate. |
ka2a_checkpoint_commits_total | counter | none | Consumer offset commits that the broker confirmed. | A rate of 0 while records arrive. |
ka2a_checkpoint_failures_total | counter | none | Consumer offset commits that failed. | A sustained rate. |
ka2a_consumer_restarts_total | counter | none | Restarts of the consumer after a consumer failure. | More than 3 in 15 minutes. |
ka2a_dispatch_started_total | counter | none | Handler starts. | Trend only. |
ka2a_dispatch_completed_total | counter | none | Handler results that the store recorded. | Trend only. |
ka2a_dispatch_held_total | counter | none | Dispatches that the node held. | Any increase. |
ka2a_dispatch_interrupted_total | counter | none | Dispatches that a stop or a deadline interrupted. | A rate outside deployments. |
ka2a_observer_delivered_total | counter | none | Events that the observer received. | Trend only. |
ka2a_observer_dropped_total | counter | none | Events that a full observer queue dropped. Messaging continues. | A sustained rate: the observer is too slow. |
ka2a_observer_panics_total | counter | none | Panics of the observer. | Any increase. |
Alert rules
examples/alerts.yml has Prometheus alerting rules for the conditions
above. Adapt the thresholds and the for durations to your service levels. A test
(TestAlertRulesNameOnlyExposedMetrics, pkg/ka2a) checks that the file parses and that
every ka2a_ metric that it names is in the exposition.