ka2a documentation

Production deployment

Developer preview. Not yet production ready. This page is rendered from docs/production-deployment.md of the ka2a repository at revision 96fb45e5e5f6693e779837998c3aa9581d6a37f2. It describes the behavior of that revision.

Contents

This guide tells you how to deploy ka2a on a secured Kafka cluster. Read the runbooks before you go live. The configuration reference lists every flag and every field of ka2a.Config.

ka2a is still a developer preview. Before you plan a deployment, read the section "What was not executed" of the newest release-candidate report, and the section "Security limitations" of the report of 2026-09-30.

Topology

  • Each endpoint has one owner process. The owner holds the state directory of the endpoint with an exclusive lock. It is the only member of the consumer group of the endpoint.
  • Each endpoint has one mailbox topic. The trusted catalog gives its name.
  • Run the owner on the host that has the state directory on a local disk. Do not put a state directory on a network file system. The store refuses a file system that is not on its allowlist (unsafe_storage).
  • Two hosts must never run the same endpoint. The owner lock (flock) does not see a lock of another host.
  • Embed pkg/ka2a in your application to serve requests. ka2a run has no application handler: it answers SendMessage and CancelTask with the A2A error UnsupportedOperation.
 application process (pkg/ka2a)         application process (pkg/ka2a)
 endpoint billing                        endpoint orders
  state dir (SQLite, 0700)                state dir (SQLite, 0700)
         |  TLS + SASL                           |  TLS + SASL
         v                                       v
 +----------------------- Kafka cluster ----------------------+
 |  mailbox topic of billing        mailbox topic of orders   |
 |  replication factor 3, min.insync.replicas 2, no unclean   |
 +------------------------------------------------------------+

The qualified profile

Use the durability profile qualified for production. It is the default of ka2a.Config and of the --profile flag.

SettingQualifiedDevelopment
Brokers3 or more1
Replication factor of the mailbox31
min.insync.replicas21
unclean.leader.election.enablefalsefalse
Broker TLSRequiredOptional (AllowPlaintext)
Broker-loss durabilityYesNo
  • The node verifies its mailbox topic at the start of Run. A mismatch of a setting stops Run with invalid_config. A setting that the broker does not report stays unverified.
  • The runtime never creates or changes a topic. Create each mailbox once with ka2a topics provision. Check it with ka2a topics verify or ka2a doctor.
  • The producer uses acks=all and idempotence. The consumer reads committed records only and never commits automatically. ka2a sets these client settings. You cannot change them.

Broker TLS and SASL

Networked brokers need TLS, authentication and ACLs. ka2a supports these SASL mechanisms: SCRAM-SHA-512, SCRAM-SHA-256 and PLAIN. Use SCRAM-SHA-512. PLAIN sends the password inside the TLS connection. ka2a refuses PLAIN without TLS. Only the administrative commands accept --allow-plaintext-credentials, for a local test broker.

Command line:

ka2a topics provision --brokers kafka1:9093,kafka2:9093,kafka3:9093 \
  --catalog /etc/ka2a/catalog.json --domain example --endpoint billing \
  --tls --tls-ca /etc/ka2a/kafka-ca.pem \
  --sasl-mechanism SCRAM-SHA-512 --sasl-username ka2a-billing \
  --sasl-password-file /etc/ka2a/kafka-password
  • --tls turns on TLS. --tls-ca or --tls-server-name also turn it on.
  • --tls-ca is a PEM file with the CA certificates. Without it, ka2a uses the system pool.
  • --tls-server-name sets the name that the broker certificate must hold. Without it, ka2a uses the host of each broker address.
  • --sasl-password-file must have mode 0600. ka2a never accepts a password as a flag value. It removes trailing line breaks.

pkg/ka2a:

cfg.Kafka.TLS = &tls.Config{MinVersion: tls.VersionTLS12, RootCAs: pool}
cfg.Kafka.SASL = &ka2a.SASL{Mechanism: "SCRAM-SHA-512", Username: user, Password: password}

Mutual TLS

A broker listener with ssl.client.auth=required needs a client certificate. Give it with --tls-cert and --tls-key (or Kafka.TLS.Certificates in pkg/ka2a). SASL is then optional: without SASL, the broker maps the subject of the certificate to the principal with ssl.principal.mapping.rules, for example RULE:^CN=([^,]+)$/$1/,DEFAULT gives User:<CN>.

ka2a run --brokers kafka1:9093,kafka2:9093,kafka3:9093 \
  --catalog /etc/ka2a/catalog.json --signing-key /etc/ka2a/billing.key.json \
  --tls-ca /etc/ka2a/kafka-ca.pem \
  --tls-cert /etc/ka2a/billing-client.pem --tls-key /etc/ka2a/billing-client-key.pem
  • --tls-cert is a PEM file: the leaf certificate first, then the intermediate CAs. The certificate must allow TLS client authentication.
  • --tls-key is the unencrypted PEM key of the leaf. It must be a regular file of the user that runs ka2a, with mode 0600. ka2a refuses a key that does not match the certificate. The key never appears in an output or a log.
  • run, send, the task commands and the topic commands refuse a certificate that is expired or not yet valid. ka2a doctor reports broker.client_certificate: a problem outside the validity period and a warning in the last 30 days. Renew the certificate before the warning becomes a problem, and reload it without a restart (see "Credential rotation without a restart").

Revoked broker certificates

--tls-crl FILE gives the broker commands the certificate revocation lists (CRL) of the broker chain. Each handshake then refuses a broker whose leaf or intermediate certificate a list revokes: tls_failed with the hint "a certificate of the broker chain is revoked by the revocation list of the client (--tls-crl)". ka2a reads only the local file. It never fetches a CRL distribution point and never asks an OCSP responder, so distribute the lists like the CA bundle.

  • The file holds PEM blocks X509 CRL (or one DER list), at most 64 lists and 8 MiB.
  • --tls-crl needs --tls-ca. A certificate of that bundle must sign each list; for a list of an intermediate CA, put the intermediate in the bundle. A list of the root that revokes that intermediate still refuses its brokers: ka2a examines every certificate of the chain, also a CA of the bundle.
  • Give complete lists. ka2a refuses a delta list, an indirect list (entries of other CAs), a list of attribute certificates and a list or an entry with a critical extension that it does not read. Partitioned lists (an issuing distribution point that only names the point) and several lists of one CA are fine; a certificate that any of them revokes is refused.
  • Each list must be current when ka2a reads it: after its this update and before its next update. ka2a refuses a stale list at the start and at a reload, with a line that names --tls-crl. A running node keeps its lists after their next update, so publish the next list before that time and send SIGHUP to ka2a run.
  • Watch the next update. ka2a doctor --tls-crl reports broker.revocation_lists: ok with the earliest next update of the lists, or a warning when it is less than 24 hours away (lists live for days, so the window of the client certificate, 30 days, does not fit). A running ka2a run shows the earliest next update of the lists in use in the run UI and in ka2a_revocation_lists_next_update_timestamp_seconds, and the reload line names it; an accepted reload updates it, a refused one keeps it. Alert on ka2a_revocation_lists_next_update_timestamp_seconds - time() < 86400 for a value above 0.
  • ka2a run reads the file again at SIGHUP. The broker check of the reload uses the new lists, so a list that revokes the certificate of a reachable broker is refused before the node uses it.
  • In pkg/ka2a, build the same check with ka2a.ParseBrokerRevocationLists(caPEM, crl, now). It applies the rules above (the same code as --tls-crl) and refuses bad input with ErrInvalidConfig and a fixed cause. Set its VerifyConnection as Config.Kafka.TLS.VerifyConnection. A refused handshake is ErrBrokerCertificateRevoked, shown as tls_failed with the revocation hint. NextUpdate() gives the earliest next update: warn before it, read the current lists, parse them again and call Node.ReloadCredentials with the new function. Give NextUpdate() as Config.Kafka.RevocationListsNextUpdate and BrokerCredentials.RevocationListsNextUpdate, so Status.Credentials and the metric show it (the node does not read it from the function). Node.ReloadCredentials may replace the function but refuses a configuration without it (incompatible_credentials), like one that turns off the verification.
  • A revoked client certificate is the task of the broker: give the broker its own CRL or remove the CA from its trust store.

ACLs

Give the principal of an endpoint only the permissions that it needs. ka2a topics acl-plan prints the minimal set from the catalog and the kafka-acls.sh commands. It never contacts the broker and never applies an ACL. Review the commands, then let an administrative principal apply them.

ka2a topics acl-plan --catalog /etc/ka2a/catalog.json --domain example --endpoint billing \
  --principal User:billing

With --output json, the command prints the document ka2a.topics.acl-plan/1. For the catalog of "Catalog format", the plan of billing is (the members commands and notes are left out here):

{
  "format": "ka2a.topics.acl-plan/1",
  "principal": "User:billing",
  "endpoint": "example/billing",
  "mailbox_topic": "ka2a.v1.mailbox.f835940d72fabe417b968e3591020099dc8102b39c6f7f28aabe7689afc9d194",
  "consumer_group": "ka2a.owner.example.billing",
  "acls": [
    {
      "resource_type": "TOPIC",
      "resource_name": "ka2a.v1.mailbox.f835940d72fabe417b968e3591020099dc8102b39c6f7f28aabe7689afc9d194",
      "pattern_type": "LITERAL",
      "operation": "Read",
      "permission": "ALLOW",
      "host": "*",
      "purpose": "consume the own mailbox (Read implies Describe)"
    },
    {
      "resource_type": "TOPIC",
      "resource_name": "ka2a.v1.mailbox.f835940d72fabe417b968e3591020099dc8102b39c6f7f28aabe7689afc9d194",
      "pattern_type": "LITERAL",
      "operation": "DescribeConfigs",
      "permission": "ALLOW",
      "host": "*",
      "purpose": "check the mailbox against the declared profile"
    },
    {
      "resource_type": "GROUP",
      "resource_name": "ka2a.owner.example.billing",
      "pattern_type": "LITERAL",
      "operation": "Read",
      "permission": "ALLOW",
      "host": "*",
      "purpose": "join the consumer group and commit offsets (Read implies Describe)"
    },
    {
      "resource_type": "TOPIC",
      "resource_name": "ka2a.v1.mailbox.3ac8c0a334732d3a44d5491b15cc2e992983d949a54f61a0a973d25288dcc9cb",
      "pattern_type": "LITERAL",
      "operation": "Write",
      "permission": "ALLOW",
      "host": "*",
      "purpose": "send responses to example/orders (Write implies Describe)"
    }
  ]
}
ResourceOperationWhy
Topic: own mailboxRead (implies Describe)Consume the own mailbox.
Topic: own mailboxDescribeConfigsThe mailbox check. Without it, the configuration checks stay unverified.
Group: ka2a.owner.<domain>.<endpoint>Read (implies Describe)Join the group and commit offsets.
Topic: mailbox of each peer that the endpoint sends requests to, or that sends requests to itWrite (implies Describe)Requests go to the mailbox of the peer; responses go to the mailbox of the requester.
Topic: each mailboxCreate, DescribeConfigsOnly the administrative principal that runs ka2a topics provision.
ClusterIdempotentWriteOnly on brokers older than Kafka 3.0.

The real-broker tests apply exactly the printed set to two endpoints and exchange a signed request (testdata/compose/kafka-mtls-acl.yml). When a permission is missing, the node status and log show authorization_failed with the scope of the missing ACL, and doctor and topics verify show the problem broker.authorization. The runbook explains the fix.

Credential rotation without a restart

A running owner reloads its broker credentials without a restart: the CA bundle (--tls-ca), the revocation lists (--tls-crl), the server name, the client certificate and key (--tls-cert, --tls-key) and the SASL password file (--sasl-password-file). The catalog and the signing key are not broker credentials; a change of them needs a restart (see "Catalog changes need a restart").

  1. Replace the files at their paths. Keep the key and the password file private (mode 0600). Write the certificate and its key before the next step: ka2a reads both at the reload.
  2. Send SIGHUP to the ka2a run process (kill -HUP <pid>). In pkg/ka2a, call Node.ReloadCredentials(ctx, ka2a.BrokerCredentials{TLS: ..., SASL: ...}).
  3. Read the line on stderr, or the status. ka2a run prints "broker credentials reloaded; new broker connections use them" with the end of the validity of the client certificate, or "broker credentials not reloaded (CLASS)" with the cause. Every refusal line names its class in place of CLASS, for example invalid_credentials for an empty password file or tls_failed.

The node checks the new material before it uses it:

  • TLS and SASL stay on or off, and the SASL mechanism stays. A tls.Config of Node.ReloadCredentials must not turn off the verification of the broker certificate (InsecureSkipVerify) or lower MinVersion. A change of them is refused with incompatible_credentials; restart the owner for such a change. Node.ReloadCredentials can change the SASL user name; ka2a run reads only the files again and keeps --sasl-username until a restart. A new principal needs the ACLs of ka2a topics acl-plan first.
  • The command reads the files with the rules of the start: a private key file that matches the certificate, a certificate that allows client authentication and is valid now. The node refuses an expired certificate or an empty password with invalid_credentials.
  • The node sends one metadata request for its own mailbox with the new material to each broker of --brokers (Config.Kafka.Brokers), in parallel (at most 8 at a time), each through a separate connection, all within the request timeout. Each broker must accept the TLS handshake and the SASL authentication, and must allow Describe on the mailbox. A refusal of any broker gives tls_failed, authentication_failed or authorization_failed with a fixed hint, also when another broker accepted the material: a broker with another trust store or user database would refuse the new connections after the swap. A broker that does not answer does not refuse the reload when another broker accepted it. When no broker answers, the class is that of the outage (for example connection_refused) of the first broker, and the reload is refused too: the node never swaps to material that no broker accepted. Only the bootstrap list is checked: list every broker in --brokers (or a name per broker) so that the check reaches each trust store; a broker that only the cluster metadata names is seen at the next connection to it.
  • After the checks, the node swaps the material as a whole. Each new broker connection of the producer, the consumer and the admin client uses it. An open connection keeps the material of its handshake until it closes. Kafka does not authenticate an open TLS connection again, and a SASL connection authenticates again only when the broker sets connections.max.reauth.ms. So a broker that removes the old certificate or password from its trust store does not close open connections. Restart the owner when the old material must stop at once, for example after a leak.
  • A refusal changes nothing: the node keeps the previous material and its work.

Accepted, queued and in-flight work does not change in either case. No record is published again, and no consumer group rebalance starts. The status (Status.Credentials), the metrics and the run UI show the time and the result of the last reload and the end of the validity of the client certificate:

MetricMeaning
ka2a_credentials_last_reload{result}1 for the result of the last reload: none, reloaded or refused.
ka2a_credentials_last_reload_timestamp_secondsTime of the last reload, or 0.
ka2a_credentials_reloads_total{result}Reloads by result: reloaded or refused.
ka2a_client_certificate_not_after_timestamp_secondsEnd of the validity of the client certificate in use, or 0. Alert well before it.
ka2a_revocation_lists_next_update_timestamp_secondsEarliest next update of the revocation lists of the broker chain in use (--tls-crl), or 0 without lists. Publish the next lists and send SIGHUP before it.
ka2a_catalog_disk{state}The catalog file compared with the running catalog: same, differs or refused.

Rotate a broker CA in two reloads: first add the new CA to the bundle and reload; then let the broker change its certificate; then remove the old CA and reload again. A reload with only the new CA before the broker changes its certificate is refused with tls_failed. This procedure ran on a real broker (TestIntegrationBrokerCARotation, k0-decisions section 31.2). Reload every node that talks to the broker in each step, also the nodes that only answer requests.

Set connections.max.reauth.ms on the broker so that open SASL connections move to a new password (k0-decisions section 31.1):

  • Same user, new password: set the new password on the broker, then reload the nodes at once. SCRAM keeps one password per user, so the change removes the old password. At the next re-authentication, each open connection authenticates with the new password. A re-authentication between the change and the reload fails, and the client connects again after the reload.
  • New user: create the user and its ACLs (ka2a topics acl-plan), reload the nodes, wait for one connections.max.reauth.ms, then delete the old user. Kafka refuses a re-authentication that changes the principal. So each open connection fails once and the client connects again with the new user. No exchange fails.
  • A produce that the broker refuses because the authentication failed is no rejection: the request stays locally queued, and the node sends it again after the reload (ka2a_publish_unknown_total counts the attempt; k0-decisions section 33.5). It expires at its horizon like work during an outage. Keep the time between the change on the broker and the reload short, or use a new user, which never has such a window. A denied ACL (authorization_failed) stays a rejection.

Check the connection

Check the connection with ka2a doctor or ka2a topics verify. When TLS or SASL refuses the connection, the check broker.connection is a problem, the command exits with 3, and the detail names the cause. The runbook lists each cause. The output never holds the password or the client key.

Signing keys

Each endpoint has one Ed25519 signing key in a private file (format ka2a.signing-key.v1, mode 0600). Make the key with ka2a keys generate:

install -d -m 0700 /etc/ka2a/keys
ka2a keys generate --endpoint example/billing --key-id billing-2026-10 \
  --out /etc/ka2a/keys/billing-2026-10.key.json \
  --not-before 2026-10-01T00:00:00Z --not-after 2027-10-01T00:00:00Z
  • The command makes the key from the system random source and writes the file with mode
    1. It never replaces an existing file, also not through a symbolic link.
  • The directory must be owned by the user who runs the command. Group and others must not write it, and others must not read it. Use mode 0700, or 0750 for a service group. The command refuses another directory with unsafe_file.
  • --endpoint is DOMAIN/ENDPOINT. --not-before defaults to the current time. Without --not-after the key has no end.
  • The command prints the catalog key entry on stdout. Only the public key is in the output:
{
  "key_id": "billing-2026-10",
  "algorithm": "ed25519",
  "public_key": "qu0s8m-HwO406XRi3sf7YKKrINtzpJ1gEsJkJXTkWDo",
  "domain": "example",
  "endpoint": "billing",
  "not_before": "2026-10-01T00:00:00Z",
  "not_after": "2027-10-01T00:00:00Z"
}

The public_key is the 32-byte Ed25519 public key in base64url without padding. Your key has a different value. Add the entry to the keys list of the catalog in a reviewed change (see "Catalog management"). To print the entry of an existing key file again, run ka2a keys public /etc/ka2a/keys/billing-2026-10.key.json --endpoint example/billing --not-before ... [--not-after ...]. It reads the file with the rules of a node: a regular file that you own with mode 0600.

Run the commands on the host of the owner, so the private key never leaves it. Never commit a key file. Never send it in a log or a ticket. --output json prints the documents ka2a.keys.generate/1 and ka2a.keys.public/1, with the entry in the member catalog_key.

Key rotation with validity windows

A catalog key has not_before, an optional not_after, revoked and replaced_by. A record verifies only when its signed creation time is inside the window of its key. A retained record that the old key signed inside its window still verifies after the rotation.

  1. Make the new key file with ka2a keys generate --not-before <switch time>. Do not use it yet.
  2. Add the printed entry to the catalog. Set not_after of the old key to a time after the switch. Leave a margin of at least the clock skew (default 5 minutes). Set replaced_by of the old key to the new key ID. Without this link, ka2a catalog validate reports the two keys as overlapping_validity.
  3. Run ka2a catalog validate on the new catalog. Distribute it to every endpoint. Restart each owner, so it loads the catalog. Until the restart, the status, the metrics (ka2a_catalog_disk{state="differs"}) and ka2a doctor show that the catalog on disk differs from the running digest.
  4. At the switch time, restart the owner of the endpoint with the new key file.
  5. Keep the old key in the catalog until its records leave the retention. The longest signed lifetime is the response horizon (default 8 days) plus the clock skew.
  6. Remove the old key, or set revoked when the key leaked. A revoked key never verifies, also for retained records.

Catalog management

The trusted catalog (format ka2a.catalog.v1) is the only source of peers, mailbox topics, keys and grants. A discovery result never grants a permission.

  • Keep the catalog under version control. Review each change by two people. ka2a has no command that edits the catalog: the catalog is the root of trust, and a rotation changes two keys in one reviewed change (k0-decisions section 27.5).
  • Check each change with ka2a catalog validate FILE before you distribute it. The command reports each finding with a severity, a class and a JSON path, and exits with 3 when it finds a problem. Problems are the defects that make a node refuse the catalog (for example duplicate_key, unknown_endpoint, rotation, unsafe_file), and two keys of one endpoint with overlapping windows and no replaced_by link (overlapping_validity). Warnings name keys that end within --expiry-days days (default 30) without a replacement that is valid at their end (key_expiring), and endpoints without a key that is valid now (no_signing_key). --strict counts warnings as problems. Run it in the review pipeline of the catalog.
  • ka2a catalog digest FILE prints the digest that a node records, sha256:<hex> of the exact bytes. Compare it on each host after the distribution.
  • Give the same catalog to every endpoint. A grant names the source, the destination and the operations (SendMessage, GetTask, ListTasks, CancelTask).
  • ka2a init --catalog records the digest of the catalog as a configuration version. ka2a doctor warns when the catalog differs from the recorded version (catalog.configuration). Run ka2a init again with the new catalog after you review the change.
  • A node loads the catalog at Open. Restart the owner after a catalog change.

Catalog format

A catalog is one JSON file of at most 1 MiB. This catalog lets orders send the four operations to billing. Each endpoint has one key, and the key of billing is the entry that ka2a keys generate printed above:

{
  "format": "ka2a.catalog.v1",
  "mailbox": {
    "prefix": "ka2a",
    "rule": "sha256-v1"
  },
  "endpoints": [
    {
      "domain": "example",
      "endpoint": "billing"
    },
    {
      "domain": "example",
      "endpoint": "orders"
    }
  ],
  "keys": [
    {
      "key_id": "billing-2026-10",
      "algorithm": "ed25519",
      "public_key": "qu0s8m-HwO406XRi3sf7YKKrINtzpJ1gEsJkJXTkWDo",
      "domain": "example",
      "endpoint": "billing",
      "not_before": "2026-10-01T00:00:00Z",
      "not_after": "2027-10-01T00:00:00Z"
    },
    {
      "key_id": "orders-2026-10",
      "algorithm": "ed25519",
      "public_key": "_1xL8cugh1KnTLaF-VAkfgQWiDZ0BBLJeHamMYlBODY",
      "domain": "example",
      "endpoint": "orders",
      "not_before": "2026-10-01T00:00:00Z",
      "not_after": "2027-10-01T00:00:00Z"
    }
  ],
  "grants": [
    {
      "source": {"domain": "example", "endpoint": "orders"},
      "destination": {"domain": "example", "endpoint": "billing"},
      "operations": ["SendMessage", "GetTask", "ListTasks", "CancelTask"]
    }
  ]
}
MemberRule
formatAlways ka2a.catalog.v1.
mailboxrule is always sha256-v1. prefix starts each derived mailbox topic name: <prefix>.v1.mailbox.<SHA-256 of "<domain>\n<endpoint>" in hex>.
endpoints1 to 1024 endpoints. domain and endpoint match [a-z0-9][a-z0-9._-]{0,63}. The optional mailbox_topic replaces the derived topic name. The optional agent_card (name, version, description, sha256) is descriptive only and grants nothing.
keysAt most 4096 entries, as ka2a keys generate prints them. Optional: not_after, revoked and replaced_by (see "Key rotation with validity windows").
grantsAt most 16384 grants. Each grant names a source, a destination and its operations. A response needs no grant.

The catalog accepts no unknown member. The node reads the file only when it is a regular file (not a symbolic link) that root or the owner user owns, and that group and others cannot write. ka2a catalog validate reports each of these rules. The test TestDocumentationJSONExamples (internal/cli, make verify) checks every JSON example of this documentation with the product.

Catalog changes need a restart

A running node never adopts a changed catalog or signing key (k0-decisions section 30.4). The catalog is the root of trust of every record in flight: the grants and keys that admitted a request also verify its response, the store records one configuration version at each start, and the mailbox topic decides the consumer subscription. A restart gives one clear boundary, and it is crash-safe: queued work stays durable, and held work stays held.

The node compares the catalog file with the running digest at each retention interval (one minute), at each credential reload and at Node.CheckCatalog. The status (Status.Catalog), the metric ka2a_catalog_disk, the run UI and one log line show a file that differs (with both digests) or that a node would refuse (refused, with the class; a restart would fail). ka2a doctor reports catalog.configuration as a warning: "The catalog on disk differs from the running digest" while an owner holds the store.

State directory ownership

  • ka2a init creates the state directory with mode 0700. Its parent must exist. Create the parent as the service user.
  • Run every owner command as the user that owns the state directory: run, send, backup, restore, quarantine release, held release and held abandon.
  • Only one process can own the directory. Other commands that need the owner lock exit with 5 while an owner runs. Read commands (observe, doctor, snapshot export, serve-ui) open the store read-only and never change it.
  • The SQLite files hold message bodies. Treat the directory and every backup as secret.

File system of the state directory

  • The owner open accepts these local file systems: on Linux ext2/3/4, xfs, btrfs, zfs, f2fs, bcachefs, tmpfs and overlay; on macOS apfs and hfs. It refuses every other type (for example NFS, CIFS or FUSE) and a read-only mount with unsafe_storage.
  • tmpfs and overlay are accepted for development and tests, but they are not durable. tmpfs loses the state at a reboot. overlay is the writable layer of a container; it is lost when the container is replaced. Then the node loses its accepted work and its deduplication state, and the new store enters quarantine with broker_ahead (see Run ka2a in a container).
  • ka2a doctor names the type in store.storage and reports store.file_system: a warning for tmpfs, overlay and a file system that the owner open refuses. Status.StoreFileSystem of the library shows the type of a running node, and a snapshot shows it in storage.file_system.
  • In production, put the state directory on a persistent local disk or volume. A bind mount or a volume shows the type of its backing file system, not overlay.

Backup and restore

ka2a backup is an offline backup. It needs the owner lock, so the owner must stop first. A backup refuses (exit 5) while a node owns the state directory. There is no online backup in this revision: an embedding application has no backup call either.

  1. Stop the owner. ka2a run drains within --drain-timeout (default 30s).
  2. Run ka2a backup --state-dir DIR --out FILE. The file must not exist. The directory of FILE must be owned by your user and must not be writable by group or others. Its parent directories must be owned by root or by your user. The command increments the live epoch after the copy.
  3. Record the backup ID and the sha256 that the command prints (--output json gives ka2a.backup/1).
  4. Start the owner again.
  5. Store the backup with the same protection as the state directory. See Data protection.

To restore:

  1. Stop the owner.
  2. Run ka2a restore --state-dir DIR --from FILE. The command keeps the replaced database under a new name (ka2a.db.pre-restore-<time>) and writes a recovery marker.
  3. Start the owner. The store enters quarantine: no dispatch starts.
  4. Follow the runbook Quarantine after a restore.

A restore never gives exactly-once recovery. Work after the backup can be lost or repeated. The quarantine makes this visible.

Downtime and recovery point

  • Downtime: each backup stops the endpoint for the drain, the copy and the start. The copy time grows with the size of the database. Measure it on a copy of your store. While the owner is stopped, requests to the endpoint wait in its mailbox topic. An embedding application cannot submit work while its node is closed.
  • Recovery point: a restore returns the state of the backup time. The broker does not give back the input after that time: the consumer group committed it, so the restored store enters quarantine (broker_ahead or a restore reason). Requests that the endpoint sent after the backup are not in the restored store. The recovery point objective is the backup interval.
  • Trade-off: a shorter interval gives a smaller loss after a restore and more downtime. A backup does not protect against the loss of a broker: the qualified profile does that (replication factor 3, min.insync.replicas 2). A backup protects against the loss or damage of the state directory.
  • A restore is the last option. First try to keep the existing state directory. See Lost or damaged state directory.

A backup schedule

Run the backup in a quiet period. This example uses a systemd timer and a script that runs as root. It stops the unit of Run ka2a as a service, writes the backup as the user ka2a and starts the unit again, also after a failed backup.

/usr/local/sbin/ka2a-backup-billing (owner root, mode 0755):

#!/bin/sh
set -eu
dir=/var/backups/ka2a/billing          # owner ka2a, mode 0700
out="$dir/billing-$(date -u +%Y%m%dT%H%M%SZ).db"
systemctl stop ka2a.service
trap 'systemctl start ka2a.service' EXIT
runuser -u ka2a -- /usr/local/bin/ka2a backup --state-dir /var/lib/ka2a/billing \
  --out "$out" --output json > "$out.json"
# Keep 14 days of backups. Remove the older files with their reports.
find "$dir" -name 'billing-*.db*' -mtime +14 -delete

/etc/systemd/system/ka2a-backup-billing.service:

[Unit]
Description=Offline backup of the ka2a endpoint example/billing

[Service]
Type=oneshot
ExecStart=/usr/local/sbin/ka2a-backup-billing

/etc/systemd/system/ka2a-backup-billing.timer:

[Unit]
Description=Daily offline backup of the ka2a endpoint example/billing

[Timer]
OnCalendar=*-*-* 03:17:00 UTC
Persistent=true

[Install]
WantedBy=timers.target

Enable it with systemctl enable --now ka2a-backup-billing.timer. With cron, run the same script from the crontab of root, for example 17 3 * * * /usr/local/sbin/ka2a-backup-billing.

Verify a backup with a restore drill

A backup that you never restored is not verified. Do a restore drill regularly, for example once a month, and after each upgrade of ka2a.

  1. Compare the checksum: sha256sum FILE must equal the sha256 of the backup report.
  2. Restore into a scratch directory, never into the live state directory: ka2a restore --state-dir /var/tmp/ka2a-drill/billing --from FILE. Create the parent directory first with mode 0700.
  3. Run ka2a doctor --state-dir /var/tmp/ka2a-drill/billing. Expect store.open ok with the schema version, store.identity with the endpoint of the backup, and store.restore as a problem with the reasons restore_marker and restored_backup. Exit code 3 is the expected result here.
  4. Run ka2a observe --state-dir /var/tmp/ka2a-drill/billing and compare the counts with your expectation.
  5. Never start an owner (ka2a run or pkg/ka2a) on the drill directory with the broker settings of the endpoint. It would join the consumer group of the live endpoint.
  6. Remove the drill directory. It holds message bodies.

Retention and encryption of backups

  • Keep backups for at least the longest retention of your stores (default 9 days) plus the time that you need to find a damage. Keep the backup of the last schema version before an upgrade until you do not need a rollback (see Upgrade notes).
  • A backup holds message bodies in plaintext. Encrypt it before it leaves the host, for example with age or gpg to a recipient key whose private key is not on the host. Remove the plaintext file after a successful encryption and a checksum comparison.
  • Keep the backups on another disk or host than the state directory.
  • Delete expired backups with a method that fits your storage. A removed file on a copy-on-write file system or in object storage with versions can stay readable.

Retention budget sizing

Limits.RetainedBytes is the retained-state budget of a store (default 1 GiB). When an admission would exceed the budget, it fails with storage_pressure. The store never removes unresolved state to make room. Retention removes completed state after its retention period.

Size the budget from the bytes that each operation keeps and from the retention periods:

budget >= operations per second x bytes per operation x retention seconds x 1.5
  • An operation keeps its request and response records, the frozen bytes and the task state. Measure the bytes per operation on your payloads: run a load for some minutes and read the retained bytes in ka2a observe.
  • The default retention of results is 8 days (Horizons.ResultRetention). The dedup retentions are 8 and 9 days.
  • The 30-minute soak of the qualification (20 operations per second, payloads from 64 B to 16 KiB) grew each store by about 10 MiB per minute. At that rate the default budget is full after about 100 minutes, before the first retention removes anything. That soak is not a capacity figure; measure your own rate.
  • Keep 50 % free disk space beyond the budget for the SQLite WAL and the backups.
  • Watch the metrics and ka2a doctor (store.storage). Act before the budget is full; see Storage pressure.

Health and readiness

ka2a has no HTTP health or readiness endpoint. ka2a run --metrics-listen serves only /metrics on a loopback address. An embedding application reads Node.Status or serves Node.MetricsHandler itself. A scrape returns 503 with the error code when the node cannot read its store.

Questionpkg/ka2aMetric
Does the node work? (liveness)Run has not returned. Status.Running is true.ka2a_node_running 1
Can the node take work? (readiness)WaitReady returned nil. Status.Ready is true and Status.Closed is false.ka2a_node_ready 1 and ka2a_node_closed 0
Does the broker answer?Status.Broker.State is reachable.ka2a_broker_state
Does a handler start?Status.Quarantine.Active is false.ka2a_quarantine_active 0

Use these rules:

  • Liveness is the process. ka2a run exits when Run fails, and a supervisor restarts it (see Run ka2a as a service). In an embedding application, treat a non-nil error from Run as fatal for the node: call Close, then exit or Open a new node. Do not add a probe that restarts a node that still runs.
  • Readiness is Status.Ready. Ready means that restart recovery ran and the background loops run, so Submit admits work. It does not mean that a broker answered, unless Config.WaitForBroker is set.
  • Do not restart on these conditions. A restart does not fix them, and it adds a drain and a start:
    • Status.Broker.State is unreachable: the node reconnects by itself. See Broker outage.
    • Status.Quarantine.Active is true: an operator must decide. See Quarantine after a restore.
    • storage_pressure, held work, ka2a_catalog_disk{state="differs"} or an unverified mailbox check.
  • Alert with the rules of examples/alerts.yml. The metrics reference explains each metric and its threshold. This table names the runbook of each rule:
Rule in alerts.ymlRunbook
KA2ANodeDown, KA2ANodeNotReadyLost or damaged state directory, Disk full
KA2ABrokerUnreachable, KA2AOutboxStalledBroker outage
KA2AMailboxCheckFailedRefused broker connection, Missing broker ACL
KA2AHeldWorkHeld work
KA2AQuarantineActiveQuarantine after a restore, Store quarantined: broker_ahead
KA2ARetainedBytesNearBudgetStorage pressure
KA2AClientCertificateExpiresSoon, KA2ACredentialReloadRefusedRefused or expiring client certificate, Rotated SASL password
KA2ARevocationListsDueRevocation list due
KA2ACatalogDiffersCatalog on disk differs from the running digest
KA2AConsumerRestartsConsumer restarts
KA2APublishRejectedMissing broker ACL
  • Also alert on two conditions that the rules do not cover: the free space of the file system of the state directory (Disk full), and the clock offset of the host (Clock drift and expired records). The error reference explains each error code, rejection reason and quarantine reason.

The metrics listener accepts only loopback addresses. Scrape it with an agent on the same host, for example a Prometheus agent or an OpenTelemetry collector.

Run ka2a as a service

docs/examples/ka2a.service is an example systemd unit for one endpoint with ka2a run. It does these things:

  • It runs as the dedicated system user ka2a, without capabilities, with NoNewPrivileges=yes, ProtectSystem=strict and a system call filter.
  • StateDirectory=ka2a/billing with StateDirectoryMode=0700 is the only writable path.
  • TimeoutStopSec=60s is longer than --drain-timeout (30s) plus the SQLite busy timeout (5s). A shorter stop timeout kills the drain. The data stays safe, but the next start holds the handler calls that were running.
  • ExecReload sends SIGHUP. systemctl reload ka2a reloads the broker credentials and checks the catalog file.
  • Restart=on-failure restarts after a fatal error. RestartPreventExitStatus=2 4 5 stops the restarts after a usage error and when another process owns the state directory.
  • It starts after time-sync.target. See Clock drift and expired records.

systemd-analyze verify accepts the unit (systemd 255, with the binary path changed to an existing file), and systemd-analyze security rates its exposure 1.4 ("OK"). The unit did not run under systemd in a test. Check it on your distribution before you go live.

Do not use DynamicUser=yes. ka2a requires that the effective user owns the state files and the secret files.

Run ka2a in a container

docs/examples/Containerfile builds a static CGO_ENABLED=0 binary and copies it into gcr.io/distroless/static-debian12:nonroot. The project does not publish an image, and no test built this file.

  • Keep the state directory on a persistent volume of a local file system (ext4 or xfs). Do not keep it in the writable layer of the container. The store accepts the overlay file system, but that layer is lost with the container. A new container then starts with an empty store, and the store enters quarantine with broker_ahead. ka2a doctor warns (store.file_system) when the state directory is on overlay.
  • Run one container per endpoint. Never mount one volume in two containers, also not on two hosts: the owner lock does not work across hosts.
  • Set the stop timeout of the container runtime (for example docker run --stop-timeout 60) above --drain-timeout plus 5s.
  • The metrics listener accepts only loopback addresses, so a probe from outside the container cannot reach it. Use the exit of the process as the liveness signal.

Data protection

ka2a signs the records. It does not encrypt the message bodies. The bodies are in plaintext in these places:

PlaceHow longControl
Kafka mailbox topicsretention.ms of the topic. The node refuses a topic whose retention is shorter than the longest record horizon (default 192h).Keep the retention at the minimum (ka2a topics provision sets 192h). Do not set retention.ms=-1. Use broker TLS. Encrypt the broker disks. Limit the topic ACLs (ka2a topics acl-plan).
SQLite store in the state directoryUntil retention removes the completed state (default 8 or 9 days). Held and unresolved work stays until an operator resolves it.Mode 0700 and 0600 (the store refuses other modes). Encrypt the disk, for example with LUKS.
Free pages and the WAL of the storeThe owner sets secure_delete=ON: SQLite overwrites deleted rows with zeros. Old pages in the WAL stay until SQLite overwrites the WAL.Encrypt the disk. A store that a version before secure_delete wrote can have old free pages: a backup and a restore write a new file without them.
BackupsUntil you delete them.See Retention and encryption of backups.
ka2a.db.pre-restore-<time> fileska2a restore keeps the replaced database.Delete it after the quarantine release when you do not need it.
Snapshots (ka2a snapshot export)Until you delete them.A snapshot holds no bodies by default. --raw-ids adds raw identifiers.

Logs, metrics, the UI and the default snapshots hold no bodies (KA2A-R1-017).

There is no call that removes one message before its retention ends. Node.Purge applies the retention once; it never removes state before its retention period or unresolved state. If you must remove a body at once, you must remove the store and the topic records. This loses the deduplication state of the endpoint. Do not put data in a message that you must be able to erase before the retention ends.

Go-live checklist

Do each step before the first production start of an endpoint. The threat model explains why.

  1. Read the section "What was not executed" of the newest release-candidate report.
  2. Build the binary from a reviewed commit with make release-artifacts. Run make release-check and make vulncheck. See Releasing.
  3. Use the durability profile qualified: 3 brokers, replication factor 3, min.insync.replicas 2, unclean.leader.election.enable=false.
  4. Use broker TLS. Use SCRAM or mutual TLS. Never use --allow-plaintext or --allow-plaintext-credentials.
  5. Give each endpoint its own broker principal with only the ACLs of ka2a topics acl-plan. Run ka2a topics verify with the credentials of the endpoint.
  6. With --tls-crl, plan the publication of the revocation lists before their next update.
  7. Make each signing key on the host of its owner (ka2a keys generate). Keep it with mode 0600, owned by the service user. Give each key a not_after and plan its rotation.
  8. Review the catalog. Distribute it as a file owned by root, mode 0644. Compare ka2a catalog digest on each host with the reviewed value.
  9. Put the state directory on a local disk of the allowlist (ext4 or xfs), on an encrypted volume, owned by the service user, mode 0700. One host per endpoint.
  10. Synchronize the clock of every host with NTP. Alert when the offset is above 1 minute.
  11. Size Limits.RetainedBytes and the disk (see Retention budget sizing). Keep 50 % free disk space beyond the budget.
  12. Install the service with a stop timeout above the drain timeout plus 5s (Run ka2a as a service).
  13. Scrape the metrics on the host and add the alerts of Health and readiness.
  14. Schedule backups, encrypt them and do one restore drill (Backup and restore).
  15. Run ka2a doctor with the flags of the owner. Expect no problem.
  16. Give each handler the recovery class HoldOnAmbiguity, unless the handler deduplicates the operation ID. Name the people who resolve held work and quarantines.
  17. Read the runbooks and the security policy.

All ka2a documents