ka2a documentation

Runbooks

Developer preview. Not yet production ready. This page is rendered from docs/runbooks.md of the ka2a repository at revision 96fb45e5e5f6693e779837998c3aa9581d6a37f2. It describes the behavior of that revision.

Contents

Each runbook has a symptom, a check and the steps. The error reference explains each error code and reason, and the metrics reference each metric. Commands that change state need the owner lock: stop the owner first. They exit with 5 while an owner runs.

Read commands never change the store or the broker: ka2a observe, ka2a operation show, ka2a doctor, ka2a snapshot export, ka2a serve-ui, ka2a topics verify, ka2a keys public, ka2a catalog validate, ka2a catalog digest and ka2a topics acl-plan.

Held work

Symptom: ka2a doctor shows a warning in store.held_work. The metrics show held items. An Operation.Wait returns held.

ka2a holds work when it cannot know the outcome of a handler call: the handler returned a non-A2A error, panicked or timed out, or the process stopped while the handler ran. ka2a never runs held work again by itself.

  1. List the held dispatch rows: ka2a observe --state-dir DIR --list dispatch. The column IN is the inbound sequence number.
  2. Read one operation: ka2a operation show --state-dir DIR <operation>.
  3. Check the external effect of the handler in the system that the handler changes. Use the operation ID (--raw-ids shows it).
  4. Decide:
    • The earlier attempt had no effect, or the handler deduplicates the operation ID: run ka2a held release --state-dir DIR --dispatch IN --reason checked_no_effect. The next start calls the handler again with the same operation identity.
    • The effect happened, or you cannot know: run ka2a held abandon --state-dir DIR --dispatch IN --reason effect_unknown. The node sends no response. The requester sees no result.
  5. For a sent request that waits too long, run ka2a held abandon --state-dir DIR --operation <operation-id> --reason peer_gone. A queued request is then never published again, and Wait returns abandoned.
  6. Start the owner.

Never release a held row only to make the warning go away. A release can repeat an effect.

Quarantine after a restore

Symptom: the node logs a quarantine event with a reason token. ka2a doctor shows a problem in store.quarantine. No handler runs.

Reasons: a restore with ka2a restore, a backup copy opened as a store, a missing or foreign epoch witness, a database behind its witness, or a committed consumer offset beyond the accounted progress (broker_ahead).

While the store is in quarantine, senders still admit work, and the node still publishes frozen records. Receivers deduplicate the republished bytes.

  1. Stop the owner.
  2. Read the state: ka2a doctor --state-dir DIR and ka2a observe --state-dir DIR.
  3. Reconcile the work after the backup time with the peers and with the systems that the handlers change. Work after the backup can be lost (a sent request that the backup does not hold) or seen again (a received request that ran after the backup).
  4. Record the decision: ka2a quarantine release --state-dir DIR --operator NAME --note "TEXT". The command starts a new store generation and records an incident with your name and note.
  5. Resolve each held item with the runbook Held work. A quarantine release does not run, requeue or discard held work.
  6. Start the owner.

Broker outage

Symptom: Status.Broker.State is unreachable with a failure class, for example connection_refused, timeout or disconnected. The log line is ka2a broker unreachable class=<class>. ka2a run --ui-listen shows it in the UI.

During an outage the node admits work locally (stage locally_queued). It publishes the work when a broker answers. A Wait returns unknown_outcome at its deadline; the work continues.

  1. Check the class. authentication_failed and tls_failed are not an outage: use Refused broker connection.
  2. Check the brokers and the network from the host of the owner: ka2a topics verify --brokers ... --catalog ... --domain D --endpoint E with the connection flags of the owner.
  3. Do not restart the owner to fix an outage. The node reconnects by itself.
  4. Watch the retained-state budget. Admitted work stays in the store until a broker acknowledges it.
  5. After the outage, the state becomes reachable. The node publishes the queued work.

A node that starts without a broker starts without broker contact, unless Config.WaitForBroker is set. Its mailbox check is broker_unreachable.

Refused broker connection

Symptom: ka2a doctor or ka2a topics verify exits with 3. The check broker.connection is a problem, and the result of topics verify is connection_failed. A running node shows Status.Broker.FailureClass tls_failed or authentication_failed, and Status.Broker.FailureHint (also in the run UI after "cause:") holds the fixed cause of the table below, for example the revocation hint.

Cause in broker.connectionClassFix
the broker certificate does not chain to a trusted CAtls_failedGive the CA of the broker with --tls-ca (or Kafka.TLS.RootCAs).
the broker certificate is not valid for the server nametls_failedUse a broker address that the certificate names, or set --tls-server-name.
the broker certificate is not valid (for example expired or not yet valid)tls_failedRenew the broker certificate. Check the clock of the host.
the broker listener needs TLS, and the client connected without TLStls_failedAdd --tls.
the broker listener does not speak TLStls_failedUse the TLS listener of the broker.
SASL authentication failed (user name, password or mechanism)authentication_failedCheck --sasl-username, --sasl-password-file and --sasl-mechanism. Check the user on the broker.
the broker does not enable this SASL mechanismauthentication_failedUse a mechanism that the listener enables.
the broker listener needs a client certificate (mutual TLS), and the client sent none that a CA trusted by the broker issuedtls_failedGive --tls-cert and --tls-key. Use a certificate of a CA in the trust store of the broker listener.
the broker refused the client certificate (untrusted CA, expired, revoked or not valid for client authentication)tls_failedRun ka2a doctor with the same flags: broker.client_certificate shows the validity period. Renew an expired certificate. Check the trust store and the revocation list of the broker.
a certificate of the broker chain is revoked by the revocation list of the client (--tls-crl)tls_failedThe broker presents a certificate that a list of --tls-crl revokes. Install a new broker certificate. If the list is wrong, publish a corrected list of the CA and send SIGHUP (or restart). Never remove --tls-crl to connect to a revoked broker.
the TLS handshake failed (for example no common TLS version or cipher, or a missing or refused client certificate)tls_failedCheck the TLS versions of the listener (ka2a needs TLS 1.2 or 1.3), then the client certificate.

The output never holds the password. Fix the configuration, then run ka2a doctor again before you restart the owner. A running node with a refused connection keeps admitting work locally. Its start log says no broker answered with the class.

Revocation list due

Symptom: ka2a doctor --tls-crl ... reports broker.revocation_lists as a warning ("the earliest next update is ..., in less than 24 hours"), or ka2a_revocation_lists_next_update_timestamp_seconds is less than a day ahead. A start or a reload after that time fails with "--tls-crl: ... next update" and the node keeps the lists it has.

  1. Get the current lists of each CA of --tls-ca from your CA (the next list must be published before the next update of the current one).
  2. Write them to the path of --tls-crl (PEM blocks X509 CRL, complete lists only).
  3. Run ka2a doctor --tls-ca CA --tls-crl NEW .... Expect broker.revocation_lists ok with a later next update and no broker.connection problem (a list that revokes the broker certificate gives tls_failed with the revocation hint).
  4. Send SIGHUP to ka2a run. Expect "broker credentials reloaded" with "The revocation lists are due for their next update at ..." on stderr, and the new time in the run UI and in ka2a_revocation_lists_next_update_timestamp_seconds. A refused reload keeps the previous lists and the previous time; the line names the class.
  5. A running node keeps checking with its lists after their next update; only the next start or reload refuses them. Do not remove --tls-crl to get past a stale list.

Refused or expiring client certificate

Symptom: ka2a run, send or topics verify fails with "--tls-cert: the client certificate expired at ...", or ka2a doctor reports broker.client_certificate as a warning or a problem.

  1. Issue a new client certificate with the same subject, so the principal and its ACLs stay the same.
  2. Write the key with mode 0600, owned by the user that runs ka2a.
  3. Run ka2a doctor --tls-cert NEW --tls-key NEW-KEY .... Expect broker.client_certificate ok and no broker.connection problem.
  4. Reload now, without a restart: write the new files at the paths of --tls-cert and --tls-key, then send SIGHUP to ka2a run. Expect "broker credentials reloaded" with the new end of validity on stderr, and the new time in ka2a_client_certificate_not_after_timestamp_seconds.
  5. If the reload is refused, the node keeps the previous certificate. The line names the class in parentheses, for example tls_failed for a certificate of a CA that a broker does not trust. The node checks every broker of --brokers, so one broker with another trust store refuses the reload; fix the trust store of that broker first. Fix the files and send SIGHUP again. When the previous certificate has expired already, restart the owner instead: unfinished work stays durable across the restart.

An expired certificate stops new broker connections only. Open connections stay until they close, so reload before the end of the validity, not after it.

A key file with group or other permissions, or a key that does not match the certificate, is refused at start. The message names the flag, never the key.

Rotated SASL password

Symptom: the broker administrator changes the password of the SASL user of an endpoint.

  1. Write the new password to the file of --sasl-password-file (mode 0600).
  2. Send SIGHUP to ka2a run. Expect "broker credentials reloaded".
  3. A refusal with authentication_failed means that the broker does not accept the new password yet. Open connections stay authenticated with the old password, so the node keeps working. Fix the password and send SIGHUP again.
  4. Requests whose produce failed the authentication in the meantime stay locally queued (ka2a_publish_unknown_total counts the attempts). The node sends them again after the reload, and the handler runs once; only the horizon of a request ends its attempts.
  5. A refusal with invalid_credentials means that the file is empty, unsafe (group or other permissions) or unreadable. The node keeps the previous password.

Catalog on disk differs from the running digest

Symptom: ka2a_catalog_disk{state="differs"} is 1, the log says "ka2a catalog on disk differs from the running digest", or ka2a doctor reports it in catalog.configuration.

  1. Check that the new catalog is the reviewed one: ka2a catalog validate FILE and ka2a catalog digest FILE.
  2. Restart the owner to adopt it. The running node keeps the previous catalog until then.

With state="refused", a node refuses the file on disk. Do not restart: the start would fail. Fix or restore the file first.

Missing broker ACL

Symptom: ka2a doctor or ka2a topics verify exits with 3 with the problem broker.authorization and the result authorization_failed. A running node logs "ka2a broker authorization failed" with class=authorization_failed and a scope, and its status shows Broker.DenialScope. The broker stays reachable.

ScopeMissing ACLEffect
group_readRead on the group ka2a.owner.<domain>.<endpoint>The node consumes nothing. Requests to it wait in its mailbox.
topic_readRead on the own mailbox topicThe node consumes nothing.
topic_describeDescribe or DescribeConfigs on the own mailbox topicThe mailbox check stays unverified (no_permission).
topic_writeWrite on the mailbox topic of a peerThe request or the response fails with publish_rejected (publish_unauthorized). It is not retried.
clusterIdempotentWrite on the cluster (brokers before Kafka 3.0)Publications fail.
  1. Print the minimal set: ka2a topics acl-plan --catalog ... --domain ... --endpoint ... --principal User:<name>. With mutual TLS, the principal comes from the client certificate and ssl.principal.mapping.rules; with SASL it is User:<user name>.
  2. Apply the missing entries with an administrative principal. ka2a never changes ACLs.
  3. Run ka2a topics verify and ka2a doctor again. Expect no broker.authorization problem.
  4. The client retries a denied group or fetch with a backoff of at most 5 s, so the node continues without a restart. Send a request that failed with publish_rejected again; the rejected record is not retried.

Signing key expires soon

Symptom: ka2a catalog validate reports the warning key_expiring, or no_signing_key for an endpoint. An owner whose key window ended does not start: ka2a doctor and ka2a run fail with key_not_valid.

  1. Make the new key on the host of the owner: ka2a keys generate --endpoint DOMAIN/ENDPOINT --key-id NEW-ID --out DIR/NEW-ID.key.json --not-before <switch time>. The directory must have mode 0700 (or 0750).
  2. In a reviewed change of the catalog, add the printed entry. Set not_after of the old key after the switch time plus the clock skew, and set its replaced_by to NEW-ID.
  3. Run ka2a catalog validate NEW-CATALOG. It must exit with 0. Without the replaced_by link it reports overlapping_validity.
  4. Distribute the catalog. On each host, compare ka2a catalog digest with the reviewed value. Restart each owner.
  5. At the switch time, restart the owner with --signing-key DIR/NEW-ID.key.json.
  6. Keep the old key in the catalog until its records leave the retention (see Key rotation).

When a key leaked, set revoked of the key instead and rotate at once. A revoked key never verifies, also for retained records.

Catalog check fails

Symptom: ka2a catalog validate exits with 3, or a node refuses the catalog with invalid_catalog or unsafe_file.

ClassFix
unsafe_fileMake the file a regular file that root or the owner user owns, and that group and others cannot write (for example mode 0644).
syntax, format, limit, invalid_fieldFix the named member. The catalog accepts no unknown member.
duplicate_endpoint, duplicate_topicList each endpoint once. Give each endpoint its own mailbox topic.
duplicate_keyGive each key its own key ID and its own key pair. Make a new key with ka2a keys generate.
unknown_endpointAdd the endpoint to endpoints, or remove the key or the grant.
duplicate_grantMerge the operations of the source and destination pair into one grant.
rotationLet replaced_by name a later key of the same endpoint. Remove the cycle.
overlapping_validityLink the keys with replaced_by, end the old window before the new one starts, or set revoked.

Fix the catalog in the reviewed source and validate it again before you distribute it.

Storage pressure

Symptom: Submit and ka2a send fail with storage_pressure. The receiver sends an admission_paused event and admits nothing more from the partition. ka2a doctor shows store.storage near the budget.

ka2a never removes unresolved state to make room.

  1. Read the retained bytes and the budget: ka2a doctor --state-dir DIR.
  2. Find what keeps the state: ka2a observe --state-dir DIR. Look for held work, failed outbox records and old unanswered requests.
  3. Resolve held work with the runbook Held work. Abandon sent requests that will never get a response.
  4. Wait for retention. Completed state leaves after its retention period.
  5. If the load needs more space, stop the owner and start it with a larger Limits.RetainedBytes. Check the free disk space first. See Retention budget sizing.

identity_conflict

Symptom: Submit, ka2a send or a peer answer fails with identity_conflict.

The operation key (SubmitOptions.OperationKey, --operation-id) is already used by another request: the body, the operation, the recovery class or the horizon differs. ka2a never sends the second request.

  1. Do not reuse a key for a new request. Make a new key (a lowercase UUIDv4) for each new request.
  2. To repeat the original request, send the identical request with the same key. You get the original operation and its result, with no new bytes on the broker.
  3. In a context, a principal can send messages only to its own contexts. A message that names the context of another principal also fails with identity_conflict. Use a new context.

Unknown outcomes and exit code 6

Symptom: ka2a send, task get, tasks list or task cancel exits with 6. Wait returns unknown_outcome.

Exit code 6 is not a failure. The request is admitted and durable. ka2a does not know the outcome yet: no response came within --wait, or the exchange is held. A timeout is never a remote failure.

  1. Do not send a new request with a new key. That can repeat the effect.
  2. Read the stages: ka2a operation show --state-dir DIR <operation>. The stages are locally_queued, broker acknowledged, remote recorded and the response.
  3. Repeat the command with the same --operation-id and the same message. It returns the original operation and waits again. A later start of the node publishes the request again if necessary and records the response.
  4. If the exchange is held, use the runbook Held work.
  5. If the peer is gone for good, abandon the request: ka2a held abandon --state-dir DIR --operation <operation-id> --reason peer_gone.

Store quarantined: broker_ahead

Symptom: the node logs "ka2a store quarantined: the broker committed input that the store never accounted for". ka2a doctor shows store.quarantine with the reason broker_ahead and a store.partitions problem: "the broker committed an offset beyond the accounted progress". ka2a_quarantine_active is 1.

When a partition is assigned, the node compares the committed offset of its consumer group with the progress that the store accounted for. The node commits only accounted input. A committed offset beyond that progress means that the store lost input that it once accounted for. These are the usual causes:

  • The state directory is new or empty (ka2a init after a loss, a new container without its volume), but the consumer group has committed offsets.
  • The state directory is an older copy: a file system or virtual machine snapshot, or a copy of the directory.
  • Another process committed offsets for the group ka2a.owner.<domain>.<endpoint>, for example a second host with the same endpoint, or an offset reset with kafka-consumer-groups.sh --reset-offsets.

The broker does not deliver the input before the committed offset again. The store cannot know if those requests ran. In quarantine, the node still records new input, but as held.

  1. Stop the owner.
  2. Make sure that only one host runs the endpoint. List the members of the group: kafka-consumer-groups.sh --bootstrap-server ... --describe --group ka2a.owner.<domain>.<endpoint> --members (with the admin credentials of your cluster). When the owner is stopped, the group must have no member. Stop every other member.
  3. Read the details: ka2a doctor --state-dir DIR and ka2a observe --state-dir DIR --list partitions. Note the committed and the accounted offsets.
  4. Find the state that is lost: the input between the accounted and the committed offset, and the requests that this endpoint sent before the loss. Reconcile them with the peers and with the systems that the handlers change.
  5. Record the decision: ka2a quarantine release --state-dir DIR --operator NAME --note "TEXT".
  6. Resolve each held item with the runbook Held work.
  7. Start the owner.

Never delete the consumer group or reset its offsets to the start to "get the input back". A group without a committed offset starts at the earliest retained record. The node would then run every request again that is still inside its signed expiry.

Lost or damaged state directory

Symptom: the state directory is gone (a lost disk or host), or the owner start fails with state_invalid and ka2a doctor reports a store.open problem, for example corrupt, unsafe_storage or state_missing.

Do not delete the files ka2a.db-wal or ka2a.db-shm. The WAL file can hold committed transactions.

  1. Stop the owner. Copy the whole state directory to a safe place before you change anything. The copy holds message bodies: keep it private.
  2. For unsafe_storage, fix the ownership and the modes. The service user must own the directory (mode 0700) and each file (mode 0600, one link). No parent directory may be writable by others. Then run ka2a doctor --state-dir DIR again. You do not need a restore.
  3. For a damaged database (corrupt) or a lost directory, check the disk and the file system first (dmesg, the SMART data, fsck on an unmounted file system).
  4. If you have a backup, restore it: ka2a restore --state-dir DIR --from FILE. If ka2a restore refuses the damaged directory, move the directory away and restore into the empty path. Then follow Quarantine after a restore.
  5. If you have no backup, create a new store with ka2a init and the same domain, endpoint, catalog and signing key. The first start enters quarantine with broker_ahead, because the consumer group has committed offsets. Follow Store quarantined: broker_ahead. The requests that the endpoint sent before the loss are lost: their responses arrive for unknown operations. Ask the peers which work is open.
  6. Find the cause before the next start, for example a full disk, a failed disk or a second owner.

Move an endpoint to a new host

Symptom: you replace the host of an endpoint, and the old state directory is readable.

Move the state directory. Do not create a new store on the new host: a new store loses the deduplication state and enters quarantine with broker_ahead.

  1. On the new host, install the same version of ka2a, the service user, the catalog, the signing key and the broker credentials of the endpoint. The broker principal must stay the same, so that the ACLs stay valid.
  2. On the old host, stop and disable the owner: systemctl disable --now ka2a.service.
  3. Copy the whole state directory while the owner is stopped, with every file: tar -C /var/lib/ka2a -cpf - billing | ssh NEW-HOST 'tar -C /var/lib/ka2a -xpf -'. Never copy the files of a running owner, and never copy ka2a.db alone: without store.epoch the store enters quarantine (epoch_witness_missing).
  4. On the new host, give the files to the service user (the user ID can differ): chown -R ka2a:ka2a /var/lib/ka2a/billing. Keep mode 0700 on the directory and 0600 on the files.
  5. Run ka2a doctor --state-dir /var/lib/ka2a/billing as the service user. Expect store.restore "No restore evidence" and store.quarantine "No quarantine".
  6. Start the owner on the new host.
  7. Remove the old state directory, or keep it only as an offline copy. Never start an owner on the old host again. The owner lock cannot see a second host. Two live copies of one store are not detected and can run a request twice.

Disk full

Symptom: Submit or ka2a send fails with storage_unavailable, ka2a_admission_retries_total rises, or ka2a backup fails with storage_full. df shows that the file system of the state directory is full. ka2a doctor can show store.storage as ok: the budget is not full.

Disk full and storage pressure are different conditions:

Disk fullStorage pressure
LimitThe free space of the file system.Limits.RetainedBytes, the retained-state budget of the store.
Errorstorage_unavailable (the store code storage_full).storage_pressure.
CauseOther files, backups, logs, or a budget that is larger than the disk.More retained state than the budget.

The node does not lose accounted work. A received record that the store cannot write stays unaccounted; the node tries it again, and the broker keeps it. A sender gets an error from Submit and nothing is queued.

  1. Find what uses the space: du -xsh /var/lib/ka2a/* /var/backups/ka2a/* and the other directories on the file system.
  2. Free space outside the state directory: move old backups to another disk, rotate the logs. Never delete a file in the state directory.
  3. The node continues by itself when space is free. You do not need a restart. Check that ka2a_admission_retries_total stops rising and that ka2a doctor reports no problem.
  4. Prevent the next event. The database file does not shrink when retention removes rows: SQLite reuses the free pages. Size the disk for the budget plus 50 %, plus the WAL (up to 64 MiB after a checkpoint), plus the backups on that disk. See Retention budget sizing. Alert on the free space of the file system, for example below 20 %.

Clock drift and expired records

Symptom: ka2a_input_rejected_total or ka2a_responses_rejected_total rises on a receiver. ka2a observe --state-dir DIR --list rejections shows the reasons future_time, expired or key_not_valid. On a sender, ka2a_publish_expired_total rises. A broker connection fails with tls_failed and the hint that a certificate is not valid yet or has expired.

Every record carries its signed creation time and expiry. ka2a uses the wall clock of the host for these rules:

  • A receiver rejects a record whose creation time is later than its own clock plus the allowed skew (future_time). The skew is Horizons.ClockSkew, default 5 minutes.
  • A receiver rejects a record at or after its signed expiry (expired). A sender does not publish an expired record (ka2a_publish_expired_total).
  • A signing key verifies a record only when the signed creation time is inside the window of the key (key_not_valid). The creation time comes from the clock of the sender.
  • Retention removes completed state by the clock of the local host. A replay after the purge is refused also when the clock moved back.
  • TLS checks the validity of the broker and client certificates with the local clock.

Requirement: synchronize the clock of every host with NTP (for example chrony or systemd-timesyncd). Keep the offset between any two hosts well below the skew of 5 minutes. Alert when the offset of a host is above 1 minute.

  1. Check the clock on the sender and on the receiver: timedatectl or chronyc tracking.
  2. Fix the time synchronization. Do not set Horizons.ClockSkew larger to hide a drift.
  3. A rejected request never reached a handler. The receiver keeps the rejection. On the sender, the operation waits for a response that does not come. Check the rejection on the receiver, then abandon the old operation on the sender (ka2a held abandon --state-dir DIR --operation <operation-id> --reason peer_gone) and send the request again with a new operation key.
  4. A rejected response leaves the operation of the sender without a result. Use the runbook Unknown outcomes and exit code 6. The handler ran on the peer: do not send a new request without a check.

Consumer restarts

Symptom: ka2a_consumer_restarts_total rises. The node logs "ka2a consumer failed; a new consumer starts". Or the consumer is stuck: the accounted offsets in ka2a observe --state-dir DIR --list partitions do not advance while the lag of the group grows.

The consumer stops after an internal inconsistency, for example a record for a partition that the group did not assign to it. The node then creates a new consumer with a backoff of 50 ms to 5 s. The group delivers again from the committed offset. The store deduplicates the records again (ka2a_request_duplicates_total rises). No work is lost, but a high rate costs throughput.

  1. Check the members of the group ka2a.owner.<domain>.<endpoint>: kafka-consumer-groups.sh --bootstrap-server ... --describe --group ka2a.owner.<domain>.<endpoint> --members. Expect exactly one member. Another member, for example a second host with the same endpoint or a tool that uses this group, causes rebalances. Stop it at once: two owners of one endpoint can run a request twice.
  2. Check the broker: rebalances and leader changes in the broker logs, and Status.Broker of the node.
  3. For a stuck consumer, check these causes:
    • ka2a_admission_retries_total rises: the store cannot write. See Storage pressure and Disk full.
    • Status.Broker.DenialScope is group_read or topic_read. See Missing broker ACL.
    • ka2a_checkpoint_failures_total rises: the broker refuses the offset commits. Check the broker and the group.
  4. If the restarts continue and the causes above do not apply, keep the logs of the node and of the brokers, then restart the owner. A restart is safe: unfinished work stays durable. Report the case with the logs (they hold no bodies).

All ka2a documents