ka2a documentation
Runbooks
Developer preview. Not yet production ready. This page is rendered from docs/runbooks.md of the ka2a repository at revision 96fb45e5e5f6693e779837998c3aa9581d6a37f2. It describes the behavior of that revision.
Contents
Each runbook has a symptom, a check and the steps. The error reference explains each error code and reason, and the metrics reference each metric. Commands that change state need the owner lock: stop the owner first. They exit with 5 while an owner runs.
Read commands never change the store or the broker: ka2a observe, ka2a operation show,
ka2a doctor, ka2a snapshot export, ka2a serve-ui, ka2a topics verify, ka2a keys public,
ka2a catalog validate, ka2a catalog digest and ka2a topics acl-plan.
Held work
Symptom: ka2a doctor shows a warning in store.held_work. The metrics show held items.
An Operation.Wait returns held.
ka2a holds work when it cannot know the outcome of a handler call: the handler returned a non-A2A error, panicked or timed out, or the process stopped while the handler ran. ka2a never runs held work again by itself.
- List the held dispatch rows:
ka2a observe --state-dir DIR --list dispatch. The columnINis the inbound sequence number. - Read one operation:
ka2a operation show --state-dir DIR <operation>. - Check the external effect of the handler in the system that the handler changes. Use the
operation ID (
--raw-idsshows it). - Decide:
- The earlier attempt had no effect, or the handler deduplicates the operation ID: run
ka2a held release --state-dir DIR --dispatch IN --reason checked_no_effect. The next start calls the handler again with the same operation identity. - The effect happened, or you cannot know: run
ka2a held abandon --state-dir DIR --dispatch IN --reason effect_unknown. The node sends no response. The requester sees no result.
- The earlier attempt had no effect, or the handler deduplicates the operation ID: run
- For a sent request that waits too long, run
ka2a held abandon --state-dir DIR --operation <operation-id> --reason peer_gone. A queued request is then never published again, andWaitreturnsabandoned. - Start the owner.
Never release a held row only to make the warning go away. A release can repeat an effect.
Quarantine after a restore
Symptom: the node logs a quarantine event with a reason token. ka2a doctor shows a
problem in store.quarantine. No handler runs.
Reasons: a restore with ka2a restore, a backup copy opened as a store, a missing or foreign
epoch witness, a database behind its witness, or a committed consumer offset beyond the
accounted progress (broker_ahead).
While the store is in quarantine, senders still admit work, and the node still publishes frozen records. Receivers deduplicate the republished bytes.
- Stop the owner.
- Read the state:
ka2a doctor --state-dir DIRandka2a observe --state-dir DIR. - Reconcile the work after the backup time with the peers and with the systems that the handlers change. Work after the backup can be lost (a sent request that the backup does not hold) or seen again (a received request that ran after the backup).
- Record the decision:
ka2a quarantine release --state-dir DIR --operator NAME --note "TEXT". The command starts a new store generation and records an incident with your name and note. - Resolve each held item with the runbook Held work. A quarantine release does not run, requeue or discard held work.
- Start the owner.
Broker outage
Symptom: Status.Broker.State is unreachable with a failure class, for example
connection_refused, timeout or disconnected. The log line is
ka2a broker unreachable class=<class>. ka2a run --ui-listen shows it in the UI.
During an outage the node admits work locally (stage locally_queued). It publishes the
work when a broker answers. A Wait returns unknown_outcome at its deadline; the work
continues.
- Check the class.
authentication_failedandtls_failedare not an outage: use Refused broker connection. - Check the brokers and the network from the host of the owner:
ka2a topics verify --brokers ... --catalog ... --domain D --endpoint Ewith the connection flags of the owner. - Do not restart the owner to fix an outage. The node reconnects by itself.
- Watch the retained-state budget. Admitted work stays in the store until a broker acknowledges it.
- After the outage, the state becomes
reachable. The node publishes the queued work.
A node that starts without a broker starts without broker contact, unless
Config.WaitForBroker is set. Its mailbox check is broker_unreachable.
Refused broker connection
Symptom: ka2a doctor or ka2a topics verify exits with 3. The check broker.connection
is a problem, and the result of topics verify is connection_failed. A running node shows
Status.Broker.FailureClass tls_failed or authentication_failed, and
Status.Broker.FailureHint (also in the run UI after "cause:") holds the fixed cause of the
table below, for example the revocation hint.
Cause in broker.connection | Class | Fix |
|---|---|---|
| the broker certificate does not chain to a trusted CA | tls_failed | Give the CA of the broker with --tls-ca (or Kafka.TLS.RootCAs). |
| the broker certificate is not valid for the server name | tls_failed | Use a broker address that the certificate names, or set --tls-server-name. |
| the broker certificate is not valid (for example expired or not yet valid) | tls_failed | Renew the broker certificate. Check the clock of the host. |
| the broker listener needs TLS, and the client connected without TLS | tls_failed | Add --tls. |
| the broker listener does not speak TLS | tls_failed | Use the TLS listener of the broker. |
| SASL authentication failed (user name, password or mechanism) | authentication_failed | Check --sasl-username, --sasl-password-file and --sasl-mechanism. Check the user on the broker. |
| the broker does not enable this SASL mechanism | authentication_failed | Use a mechanism that the listener enables. |
| the broker listener needs a client certificate (mutual TLS), and the client sent none that a CA trusted by the broker issued | tls_failed | Give --tls-cert and --tls-key. Use a certificate of a CA in the trust store of the broker listener. |
| the broker refused the client certificate (untrusted CA, expired, revoked or not valid for client authentication) | tls_failed | Run ka2a doctor with the same flags: broker.client_certificate shows the validity period. Renew an expired certificate. Check the trust store and the revocation list of the broker. |
| a certificate of the broker chain is revoked by the revocation list of the client (--tls-crl) | tls_failed | The broker presents a certificate that a list of --tls-crl revokes. Install a new broker certificate. If the list is wrong, publish a corrected list of the CA and send SIGHUP (or restart). Never remove --tls-crl to connect to a revoked broker. |
| the TLS handshake failed (for example no common TLS version or cipher, or a missing or refused client certificate) | tls_failed | Check the TLS versions of the listener (ka2a needs TLS 1.2 or 1.3), then the client certificate. |
The output never holds the password. Fix the configuration, then run ka2a doctor again
before you restart the owner. A running node with a refused connection keeps admitting work
locally. Its start log says no broker answered with the class.
Revocation list due
Symptom: ka2a doctor --tls-crl ... reports broker.revocation_lists as a warning ("the
earliest next update is ..., in less than 24 hours"), or
ka2a_revocation_lists_next_update_timestamp_seconds is less than a day ahead. A start or a
reload after that time fails with "--tls-crl: ... next update" and the node keeps the lists
it has.
- Get the current lists of each CA of
--tls-cafrom your CA (the next list must be published before the next update of the current one). - Write them to the path of
--tls-crl(PEM blocksX509 CRL, complete lists only). - Run
ka2a doctor --tls-ca CA --tls-crl NEW .... Expectbroker.revocation_listsokwith a later next update and nobroker.connectionproblem (a list that revokes the broker certificate givestls_failedwith the revocation hint). - Send
SIGHUPtoka2a run. Expect "broker credentials reloaded" with "The revocation lists are due for their next update at ..." on stderr, and the new time in the run UI and inka2a_revocation_lists_next_update_timestamp_seconds. A refused reload keeps the previous lists and the previous time; the line names the class. - A running node keeps checking with its lists after their next update; only the next start
or reload refuses them. Do not remove
--tls-crlto get past a stale list.
Refused or expiring client certificate
Symptom: ka2a run, send or topics verify fails with "--tls-cert: the client certificate
expired at ...", or ka2a doctor reports broker.client_certificate as a warning or a problem.
- Issue a new client certificate with the same subject, so the principal and its ACLs stay the same.
- Write the key with mode 0600, owned by the user that runs ka2a.
- Run
ka2a doctor --tls-cert NEW --tls-key NEW-KEY .... Expectbroker.client_certificateokand nobroker.connectionproblem. - Reload now, without a restart: write the new files at the paths of
--tls-certand--tls-key, then sendSIGHUPtoka2a run. Expect "broker credentials reloaded" with the new end of validity on stderr, and the new time inka2a_client_certificate_not_after_timestamp_seconds. - If the reload is refused, the node keeps the previous certificate. The line names the
class in parentheses, for example
tls_failedfor a certificate of a CA that a broker does not trust. The node checks every broker of--brokers, so one broker with another trust store refuses the reload; fix the trust store of that broker first. Fix the files and sendSIGHUPagain. When the previous certificate has expired already, restart the owner instead: unfinished work stays durable across the restart.
An expired certificate stops new broker connections only. Open connections stay until they close, so reload before the end of the validity, not after it.
A key file with group or other permissions, or a key that does not match the certificate, is refused at start. The message names the flag, never the key.
Rotated SASL password
Symptom: the broker administrator changes the password of the SASL user of an endpoint.
- Write the new password to the file of
--sasl-password-file(mode 0600). - Send
SIGHUPtoka2a run. Expect "broker credentials reloaded". - A refusal with
authentication_failedmeans that the broker does not accept the new password yet. Open connections stay authenticated with the old password, so the node keeps working. Fix the password and sendSIGHUPagain. - Requests whose produce failed the authentication in the meantime stay locally queued
(
ka2a_publish_unknown_totalcounts the attempts). The node sends them again after the reload, and the handler runs once; only the horizon of a request ends its attempts. - A refusal with
invalid_credentialsmeans that the file is empty, unsafe (group or other permissions) or unreadable. The node keeps the previous password.
Catalog on disk differs from the running digest
Symptom: ka2a_catalog_disk{state="differs"} is 1, the log says "ka2a catalog on disk
differs from the running digest", or ka2a doctor reports it in catalog.configuration.
- Check that the new catalog is the reviewed one:
ka2a catalog validate FILEandka2a catalog digest FILE. - Restart the owner to adopt it. The running node keeps the previous catalog until then.
With state="refused", a node refuses the file on disk. Do not restart: the start would
fail. Fix or restore the file first.
Missing broker ACL
Symptom: ka2a doctor or ka2a topics verify exits with 3 with the problem
broker.authorization and the result authorization_failed. A running node logs
"ka2a broker authorization failed" with class=authorization_failed and a scope, and its
status shows Broker.DenialScope. The broker stays reachable.
| Scope | Missing ACL | Effect |
|---|---|---|
group_read | Read on the group ka2a.owner.<domain>.<endpoint> | The node consumes nothing. Requests to it wait in its mailbox. |
topic_read | Read on the own mailbox topic | The node consumes nothing. |
topic_describe | Describe or DescribeConfigs on the own mailbox topic | The mailbox check stays unverified (no_permission). |
topic_write | Write on the mailbox topic of a peer | The request or the response fails with publish_rejected (publish_unauthorized). It is not retried. |
cluster | IdempotentWrite on the cluster (brokers before Kafka 3.0) | Publications fail. |
- Print the minimal set:
ka2a topics acl-plan --catalog ... --domain ... --endpoint ... --principal User:<name>. With mutual TLS, the principal comes from the client certificate andssl.principal.mapping.rules; with SASL it isUser:<user name>. - Apply the missing entries with an administrative principal. ka2a never changes ACLs.
- Run
ka2a topics verifyandka2a doctoragain. Expect nobroker.authorizationproblem. - The client retries a denied group or fetch with a backoff of at most 5 s, so the node
continues without a restart. Send a request that failed with
publish_rejectedagain; the rejected record is not retried.
Signing key expires soon
Symptom: ka2a catalog validate reports the warning key_expiring, or no_signing_key for an
endpoint. An owner whose key window ended does not start: ka2a doctor and ka2a run fail with
key_not_valid.
- Make the new key on the host of the owner:
ka2a keys generate --endpoint DOMAIN/ENDPOINT --key-id NEW-ID --out DIR/NEW-ID.key.json --not-before <switch time>. The directory must have mode 0700 (or 0750). - In a reviewed change of the catalog, add the printed entry. Set
not_afterof the old key after the switch time plus the clock skew, and set itsreplaced_bytoNEW-ID. - Run
ka2a catalog validate NEW-CATALOG. It must exit with 0. Without thereplaced_bylink it reportsoverlapping_validity. - Distribute the catalog. On each host, compare
ka2a catalog digestwith the reviewed value. Restart each owner. - At the switch time, restart the owner with
--signing-key DIR/NEW-ID.key.json. - Keep the old key in the catalog until its records leave the retention (see Key rotation).
When a key leaked, set revoked of the key instead and rotate at once. A revoked key never
verifies, also for retained records.
Catalog check fails
Symptom: ka2a catalog validate exits with 3, or a node refuses the catalog with
invalid_catalog or unsafe_file.
| Class | Fix |
|---|---|
unsafe_file | Make the file a regular file that root or the owner user owns, and that group and others cannot write (for example mode 0644). |
syntax, format, limit, invalid_field | Fix the named member. The catalog accepts no unknown member. |
duplicate_endpoint, duplicate_topic | List each endpoint once. Give each endpoint its own mailbox topic. |
duplicate_key | Give each key its own key ID and its own key pair. Make a new key with ka2a keys generate. |
unknown_endpoint | Add the endpoint to endpoints, or remove the key or the grant. |
duplicate_grant | Merge the operations of the source and destination pair into one grant. |
rotation | Let replaced_by name a later key of the same endpoint. Remove the cycle. |
overlapping_validity | Link the keys with replaced_by, end the old window before the new one starts, or set revoked. |
Fix the catalog in the reviewed source and validate it again before you distribute it.
Storage pressure
Symptom: Submit and ka2a send fail with storage_pressure. The receiver sends an
admission_paused event and admits nothing more from the partition. ka2a doctor shows
store.storage near the budget.
ka2a never removes unresolved state to make room.
- Read the retained bytes and the budget:
ka2a doctor --state-dir DIR. - Find what keeps the state:
ka2a observe --state-dir DIR. Look for held work, failed outbox records and old unanswered requests. - Resolve held work with the runbook Held work. Abandon sent requests that will never get a response.
- Wait for retention. Completed state leaves after its retention period.
- If the load needs more space, stop the owner and start it with a larger
Limits.RetainedBytes. Check the free disk space first. See Retention budget sizing.
identity_conflict
Symptom: Submit, ka2a send or a peer answer fails with identity_conflict.
The operation key (SubmitOptions.OperationKey, --operation-id) is already used by another
request: the body, the operation, the recovery class or the horizon differs. ka2a never sends the
second request.
- Do not reuse a key for a new request. Make a new key (a lowercase UUIDv4) for each new request.
- To repeat the original request, send the identical request with the same key. You get the original operation and its result, with no new bytes on the broker.
- In a context, a principal can send messages only to its own contexts. A message that names
the context of another principal also fails with
identity_conflict. Use a new context.
Unknown outcomes and exit code 6
Symptom: ka2a send, task get, tasks list or task cancel exits with 6. Wait returns
unknown_outcome.
Exit code 6 is not a failure. The request is admitted and durable. ka2a does not know the
outcome yet: no response came within --wait, or the exchange is held. A timeout is never a
remote failure.
- Do not send a new request with a new key. That can repeat the effect.
- Read the stages:
ka2a operation show --state-dir DIR <operation>. The stages arelocally_queued, broker acknowledged, remote recorded and the response. - Repeat the command with the same
--operation-idand the same message. It returns the original operation and waits again. A later start of the node publishes the request again if necessary and records the response. - If the exchange is held, use the runbook Held work.
- If the peer is gone for good, abandon the request:
ka2a held abandon --state-dir DIR --operation <operation-id> --reason peer_gone.
Store quarantined: broker_ahead
Symptom: the node logs "ka2a store quarantined: the broker committed input that the store
never accounted for". ka2a doctor shows store.quarantine with the reason broker_ahead
and a store.partitions problem: "the broker committed an offset beyond the accounted
progress". ka2a_quarantine_active is 1.
When a partition is assigned, the node compares the committed offset of its consumer group with the progress that the store accounted for. The node commits only accounted input. A committed offset beyond that progress means that the store lost input that it once accounted for. These are the usual causes:
- The state directory is new or empty (
ka2a initafter a loss, a new container without its volume), but the consumer group has committed offsets. - The state directory is an older copy: a file system or virtual machine snapshot, or a copy of the directory.
- Another process committed offsets for the group
ka2a.owner.<domain>.<endpoint>, for example a second host with the same endpoint, or an offset reset withkafka-consumer-groups.sh --reset-offsets.
The broker does not deliver the input before the committed offset again. The store cannot know if those requests ran. In quarantine, the node still records new input, but as held.
- Stop the owner.
- Make sure that only one host runs the endpoint. List the members of the group:
kafka-consumer-groups.sh --bootstrap-server ... --describe --group ka2a.owner.<domain>.<endpoint> --members(with the admin credentials of your cluster). When the owner is stopped, the group must have no member. Stop every other member. - Read the details:
ka2a doctor --state-dir DIRandka2a observe --state-dir DIR --list partitions. Note the committed and the accounted offsets. - Find the state that is lost: the input between the accounted and the committed offset, and the requests that this endpoint sent before the loss. Reconcile them with the peers and with the systems that the handlers change.
- Record the decision:
ka2a quarantine release --state-dir DIR --operator NAME --note "TEXT". - Resolve each held item with the runbook Held work.
- Start the owner.
Never delete the consumer group or reset its offsets to the start to "get the input back". A group without a committed offset starts at the earliest retained record. The node would then run every request again that is still inside its signed expiry.
Lost or damaged state directory
Symptom: the state directory is gone (a lost disk or host), or the owner start fails with
state_invalid and ka2a doctor reports a store.open problem, for example corrupt,
unsafe_storage or state_missing.
Do not delete the files ka2a.db-wal or ka2a.db-shm. The WAL file can hold committed
transactions.
- Stop the owner. Copy the whole state directory to a safe place before you change anything. The copy holds message bodies: keep it private.
- For
unsafe_storage, fix the ownership and the modes. The service user must own the directory (mode 0700) and each file (mode 0600, one link). No parent directory may be writable by others. Then runka2a doctor --state-dir DIRagain. You do not need a restore. - For a damaged database (
corrupt) or a lost directory, check the disk and the file system first (dmesg, the SMART data,fsckon an unmounted file system). - If you have a backup, restore it:
ka2a restore --state-dir DIR --from FILE. Ifka2a restorerefuses the damaged directory, move the directory away and restore into the empty path. Then follow Quarantine after a restore. - If you have no backup, create a new store with
ka2a initand the same domain, endpoint, catalog and signing key. The first start enters quarantine withbroker_ahead, because the consumer group has committed offsets. Follow Store quarantined: broker_ahead. The requests that the endpoint sent before the loss are lost: their responses arrive for unknown operations. Ask the peers which work is open. - Find the cause before the next start, for example a full disk, a failed disk or a second owner.
Move an endpoint to a new host
Symptom: you replace the host of an endpoint, and the old state directory is readable.
Move the state directory. Do not create a new store on the new host: a new store loses the
deduplication state and enters quarantine with broker_ahead.
- On the new host, install the same version of ka2a, the service user, the catalog, the signing key and the broker credentials of the endpoint. The broker principal must stay the same, so that the ACLs stay valid.
- On the old host, stop and disable the owner:
systemctl disable --now ka2a.service. - Copy the whole state directory while the owner is stopped, with every file:
tar -C /var/lib/ka2a -cpf - billing | ssh NEW-HOST 'tar -C /var/lib/ka2a -xpf -'. Never copy the files of a running owner, and never copyka2a.dbalone: withoutstore.epochthe store enters quarantine (epoch_witness_missing). - On the new host, give the files to the service user (the user ID can differ):
chown -R ka2a:ka2a /var/lib/ka2a/billing. Keep mode 0700 on the directory and 0600 on the files. - Run
ka2a doctor --state-dir /var/lib/ka2a/billingas the service user. Expectstore.restore"No restore evidence" andstore.quarantine"No quarantine". - Start the owner on the new host.
- Remove the old state directory, or keep it only as an offline copy. Never start an owner on the old host again. The owner lock cannot see a second host. Two live copies of one store are not detected and can run a request twice.
Disk full
Symptom: Submit or ka2a send fails with storage_unavailable,
ka2a_admission_retries_total rises, or ka2a backup fails with
storage_full. df shows that the file system of the state directory is full. ka2a doctor can show store.storage as ok: the budget is not full.
Disk full and storage pressure are different conditions:
| Disk full | Storage pressure | |
|---|---|---|
| Limit | The free space of the file system. | Limits.RetainedBytes, the retained-state budget of the store. |
| Error | storage_unavailable (the store code storage_full). | storage_pressure. |
| Cause | Other files, backups, logs, or a budget that is larger than the disk. | More retained state than the budget. |
The node does not lose accounted work. A received record that the store cannot write stays
unaccounted; the node tries it again, and the broker keeps it. A sender gets an error from
Submit and nothing is queued.
- Find what uses the space:
du -xsh /var/lib/ka2a/* /var/backups/ka2a/*and the other directories on the file system. - Free space outside the state directory: move old backups to another disk, rotate the logs. Never delete a file in the state directory.
- The node continues by itself when space is free. You do not need a restart. Check that
ka2a_admission_retries_totalstops rising and thatka2a doctorreports no problem. - Prevent the next event. The database file does not shrink when retention removes rows: SQLite reuses the free pages. Size the disk for the budget plus 50 %, plus the WAL (up to 64 MiB after a checkpoint), plus the backups on that disk. See Retention budget sizing. Alert on the free space of the file system, for example below 20 %.
Clock drift and expired records
Symptom: ka2a_input_rejected_total or ka2a_responses_rejected_total rises on a receiver.
ka2a observe --state-dir DIR --list rejections shows the reasons future_time, expired or
key_not_valid. On a sender, ka2a_publish_expired_total rises. A broker connection fails
with tls_failed and the hint that a certificate is not valid yet or has expired.
Every record carries its signed creation time and expiry. ka2a uses the wall clock of the host for these rules:
- A receiver rejects a record whose creation time is later than its own clock plus the
allowed skew (
future_time). The skew isHorizons.ClockSkew, default 5 minutes. - A receiver rejects a record at or after its signed expiry (
expired). A sender does not publish an expired record (ka2a_publish_expired_total). - A signing key verifies a record only when the signed creation time is inside the window of
the key (
key_not_valid). The creation time comes from the clock of the sender. - Retention removes completed state by the clock of the local host. A replay after the purge is refused also when the clock moved back.
- TLS checks the validity of the broker and client certificates with the local clock.
Requirement: synchronize the clock of every host with NTP (for example chrony or systemd-timesyncd). Keep the offset between any two hosts well below the skew of 5 minutes. Alert when the offset of a host is above 1 minute.
- Check the clock on the sender and on the receiver:
timedatectlorchronyc tracking. - Fix the time synchronization. Do not set
Horizons.ClockSkewlarger to hide a drift. - A rejected request never reached a handler. The receiver keeps the rejection. On the
sender, the operation waits for a response that does not come. Check the rejection on the receiver,
then abandon the old operation on the sender
(
ka2a held abandon --state-dir DIR --operation <operation-id> --reason peer_gone) and send the request again with a new operation key. - A rejected response leaves the operation of the sender without a result. Use the runbook Unknown outcomes and exit code 6. The handler ran on the peer: do not send a new request without a check.
Consumer restarts
Symptom: ka2a_consumer_restarts_total rises. The node logs "ka2a consumer failed; a new
consumer starts". Or the consumer is stuck: the accounted offsets in
ka2a observe --state-dir DIR --list partitions do not advance while the lag of the group
grows.
The consumer stops after an internal inconsistency, for example a record for a partition
that the group did not assign to it. The node then creates a new consumer with a backoff of
50 ms to 5 s. The group delivers again from the committed offset. The store deduplicates the
records again (ka2a_request_duplicates_total rises). No work is lost, but a high rate costs
throughput.
- Check the members of the group
ka2a.owner.<domain>.<endpoint>:kafka-consumer-groups.sh --bootstrap-server ... --describe --group ka2a.owner.<domain>.<endpoint> --members. Expect exactly one member. Another member, for example a second host with the same endpoint or a tool that uses this group, causes rebalances. Stop it at once: two owners of one endpoint can run a request twice. - Check the broker: rebalances and leader changes in the broker logs, and
Status.Brokerof the node. - For a stuck consumer, check these causes:
ka2a_admission_retries_totalrises: the store cannot write. See Storage pressure and Disk full.Status.Broker.DenialScopeisgroup_readortopic_read. See Missing broker ACL.ka2a_checkpoint_failures_totalrises: the broker refuses the offset commits. Check the broker and the group.
- If the restarts continue and the causes above do not apply, keep the logs of the node and of the brokers, then restart the owner. A restart is safe: unfinished work stays durable. Report the case with the logs (they hold no bodies).