ka2a documentation
Production deployment
Developer preview. Not yet production ready. This page is rendered from docs/production-deployment.md of the ka2a repository at revision 96fb45e5e5f6693e779837998c3aa9581d6a37f2. It describes the behavior of that revision.
Contents
This guide tells you how to deploy ka2a on a secured Kafka cluster. Read the
runbooks before you go live. The
configuration reference lists every flag and every field of
ka2a.Config.
ka2a is still a developer preview. Before you plan a deployment, read the section "What was not executed" of the newest release-candidate report, and the section "Security limitations" of the report of 2026-09-30.
Topology
- Each endpoint has one owner process. The owner holds the state directory of the endpoint with an exclusive lock. It is the only member of the consumer group of the endpoint.
- Each endpoint has one mailbox topic. The trusted catalog gives its name.
- Run the owner on the host that has the state directory on a local disk. Do not put a state
directory on a network file system. The store refuses a file system that is not on its
allowlist (
unsafe_storage). - Two hosts must never run the same endpoint. The owner lock (
flock) does not see a lock of another host. - Embed
pkg/ka2ain your application to serve requests.ka2a runhas no application handler: it answersSendMessageandCancelTaskwith the A2A errorUnsupportedOperation.
application process (pkg/ka2a) application process (pkg/ka2a)
endpoint billing endpoint orders
state dir (SQLite, 0700) state dir (SQLite, 0700)
| TLS + SASL | TLS + SASL
v v
+----------------------- Kafka cluster ----------------------+
| mailbox topic of billing mailbox topic of orders |
| replication factor 3, min.insync.replicas 2, no unclean |
+------------------------------------------------------------+
The qualified profile
Use the durability profile qualified for production. It is the default of ka2a.Config
and of the --profile flag.
| Setting | Qualified | Development |
|---|---|---|
| Brokers | 3 or more | 1 |
| Replication factor of the mailbox | 3 | 1 |
min.insync.replicas | 2 | 1 |
unclean.leader.election.enable | false | false |
| Broker TLS | Required | Optional (AllowPlaintext) |
| Broker-loss durability | Yes | No |
- The node verifies its mailbox topic at the start of
Run. A mismatch of a setting stopsRunwithinvalid_config. A setting that the broker does not report stays unverified. - The runtime never creates or changes a topic. Create each mailbox once with
ka2a topics provision. Check it withka2a topics verifyorka2a doctor. - The producer uses
acks=alland idempotence. The consumer reads committed records only and never commits automatically. ka2a sets these client settings. You cannot change them.
Broker TLS and SASL
Networked brokers need TLS, authentication and ACLs. ka2a supports these SASL mechanisms:
SCRAM-SHA-512, SCRAM-SHA-256 and PLAIN. Use SCRAM-SHA-512. PLAIN sends the
password inside the TLS connection. ka2a refuses PLAIN without TLS. Only the
administrative commands accept --allow-plaintext-credentials, for a local test broker.
Command line:
ka2a topics provision --brokers kafka1:9093,kafka2:9093,kafka3:9093 \
--catalog /etc/ka2a/catalog.json --domain example --endpoint billing \
--tls --tls-ca /etc/ka2a/kafka-ca.pem \
--sasl-mechanism SCRAM-SHA-512 --sasl-username ka2a-billing \
--sasl-password-file /etc/ka2a/kafka-password
--tlsturns on TLS.--tls-caor--tls-server-namealso turn it on.--tls-cais a PEM file with the CA certificates. Without it, ka2a uses the system pool.--tls-server-namesets the name that the broker certificate must hold. Without it, ka2a uses the host of each broker address.--sasl-password-filemust have mode 0600. ka2a never accepts a password as a flag value. It removes trailing line breaks.
pkg/ka2a:
cfg.Kafka.TLS = &tls.Config{MinVersion: tls.VersionTLS12, RootCAs: pool}
cfg.Kafka.SASL = &ka2a.SASL{Mechanism: "SCRAM-SHA-512", Username: user, Password: password}
Mutual TLS
A broker listener with ssl.client.auth=required needs a client certificate. Give it with
--tls-cert and --tls-key (or Kafka.TLS.Certificates in pkg/ka2a). SASL is then
optional: without SASL, the broker maps the subject of the certificate to the principal with
ssl.principal.mapping.rules, for example RULE:^CN=([^,]+)$/$1/,DEFAULT gives User:<CN>.
ka2a run --brokers kafka1:9093,kafka2:9093,kafka3:9093 \
--catalog /etc/ka2a/catalog.json --signing-key /etc/ka2a/billing.key.json \
--tls-ca /etc/ka2a/kafka-ca.pem \
--tls-cert /etc/ka2a/billing-client.pem --tls-key /etc/ka2a/billing-client-key.pem
--tls-certis a PEM file: the leaf certificate first, then the intermediate CAs. The certificate must allow TLS client authentication.--tls-keyis the unencrypted PEM key of the leaf. It must be a regular file of the user that runs ka2a, with mode 0600. ka2a refuses a key that does not match the certificate. The key never appears in an output or a log.run,send, the task commands and the topic commands refuse a certificate that is expired or not yet valid.ka2a doctorreportsbroker.client_certificate: a problem outside the validity period and a warning in the last 30 days. Renew the certificate before the warning becomes a problem, and reload it without a restart (see "Credential rotation without a restart").
Revoked broker certificates
--tls-crl FILE gives the broker commands the certificate revocation lists (CRL) of the broker
chain. Each handshake then refuses a broker whose leaf or intermediate certificate a list
revokes: tls_failed with the hint "a certificate of the broker chain is revoked by the
revocation list of the client (--tls-crl)". ka2a reads only the local file. It never fetches
a CRL distribution point and never asks an OCSP responder, so distribute the lists like the
CA bundle.
- The file holds PEM blocks
X509 CRL(or one DER list), at most 64 lists and 8 MiB. --tls-crlneeds--tls-ca. A certificate of that bundle must sign each list; for a list of an intermediate CA, put the intermediate in the bundle. A list of the root that revokes that intermediate still refuses its brokers: ka2a examines every certificate of the chain, also a CA of the bundle.- Give complete lists. ka2a refuses a delta list, an indirect list (entries of other CAs), a list of attribute certificates and a list or an entry with a critical extension that it does not read. Partitioned lists (an issuing distribution point that only names the point) and several lists of one CA are fine; a certificate that any of them revokes is refused.
- Each list must be current when ka2a reads it: after its this update and before its next
update. ka2a refuses a stale list at the start and at a reload, with a line that names
--tls-crl. A running node keeps its lists after their next update, so publish the next list before that time and sendSIGHUPtoka2a run. - Watch the next update.
ka2a doctor --tls-crlreportsbroker.revocation_lists:okwith the earliest next update of the lists, or a warning when it is less than 24 hours away (lists live for days, so the window of the client certificate, 30 days, does not fit). A runningka2a runshows the earliest next update of the lists in use in the run UI and inka2a_revocation_lists_next_update_timestamp_seconds, and the reload line names it; an accepted reload updates it, a refused one keeps it. Alert onka2a_revocation_lists_next_update_timestamp_seconds - time() < 86400for a value above 0. ka2a runreads the file again atSIGHUP. The broker check of the reload uses the new lists, so a list that revokes the certificate of a reachable broker is refused before the node uses it.- In
pkg/ka2a, build the same check withka2a.ParseBrokerRevocationLists(caPEM, crl, now). It applies the rules above (the same code as--tls-crl) and refuses bad input withErrInvalidConfigand a fixed cause. Set itsVerifyConnectionasConfig.Kafka.TLS.VerifyConnection. A refused handshake isErrBrokerCertificateRevoked, shown astls_failedwith the revocation hint.NextUpdate()gives the earliest next update: warn before it, read the current lists, parse them again and callNode.ReloadCredentialswith the new function. GiveNextUpdate()asConfig.Kafka.RevocationListsNextUpdateandBrokerCredentials.RevocationListsNextUpdate, soStatus.Credentialsand the metric show it (the node does not read it from the function).Node.ReloadCredentialsmay replace the function but refuses a configuration without it (incompatible_credentials), like one that turns off the verification. - A revoked client certificate is the task of the broker: give the broker its own CRL or remove the CA from its trust store.
ACLs
Give the principal of an endpoint only the permissions that it needs.
ka2a topics acl-plan prints the minimal set from the catalog and the kafka-acls.sh
commands. It never contacts the broker and never applies an ACL. Review the commands, then let
an administrative principal apply them.
ka2a topics acl-plan --catalog /etc/ka2a/catalog.json --domain example --endpoint billing \
--principal User:billing
With --output json, the command prints the document ka2a.topics.acl-plan/1. For the
catalog of "Catalog format", the plan of billing is (the members commands and notes are
left out here):
{
"format": "ka2a.topics.acl-plan/1",
"principal": "User:billing",
"endpoint": "example/billing",
"mailbox_topic": "ka2a.v1.mailbox.f835940d72fabe417b968e3591020099dc8102b39c6f7f28aabe7689afc9d194",
"consumer_group": "ka2a.owner.example.billing",
"acls": [
{
"resource_type": "TOPIC",
"resource_name": "ka2a.v1.mailbox.f835940d72fabe417b968e3591020099dc8102b39c6f7f28aabe7689afc9d194",
"pattern_type": "LITERAL",
"operation": "Read",
"permission": "ALLOW",
"host": "*",
"purpose": "consume the own mailbox (Read implies Describe)"
},
{
"resource_type": "TOPIC",
"resource_name": "ka2a.v1.mailbox.f835940d72fabe417b968e3591020099dc8102b39c6f7f28aabe7689afc9d194",
"pattern_type": "LITERAL",
"operation": "DescribeConfigs",
"permission": "ALLOW",
"host": "*",
"purpose": "check the mailbox against the declared profile"
},
{
"resource_type": "GROUP",
"resource_name": "ka2a.owner.example.billing",
"pattern_type": "LITERAL",
"operation": "Read",
"permission": "ALLOW",
"host": "*",
"purpose": "join the consumer group and commit offsets (Read implies Describe)"
},
{
"resource_type": "TOPIC",
"resource_name": "ka2a.v1.mailbox.3ac8c0a334732d3a44d5491b15cc2e992983d949a54f61a0a973d25288dcc9cb",
"pattern_type": "LITERAL",
"operation": "Write",
"permission": "ALLOW",
"host": "*",
"purpose": "send responses to example/orders (Write implies Describe)"
}
]
}
| Resource | Operation | Why |
|---|---|---|
| Topic: own mailbox | Read (implies Describe) | Consume the own mailbox. |
| Topic: own mailbox | DescribeConfigs | The mailbox check. Without it, the configuration checks stay unverified. |
Group: ka2a.owner.<domain>.<endpoint> | Read (implies Describe) | Join the group and commit offsets. |
| Topic: mailbox of each peer that the endpoint sends requests to, or that sends requests to it | Write (implies Describe) | Requests go to the mailbox of the peer; responses go to the mailbox of the requester. |
| Topic: each mailbox | Create, DescribeConfigs | Only the administrative principal that runs ka2a topics provision. |
| Cluster | IdempotentWrite | Only on brokers older than Kafka 3.0. |
The real-broker tests apply exactly the printed set to two endpoints and exchange a signed
request (testdata/compose/kafka-mtls-acl.yml). When a permission is missing, the node status
and log show authorization_failed with the scope of the missing ACL, and doctor and
topics verify show the problem broker.authorization. The
runbook explains the fix.
Credential rotation without a restart
A running owner reloads its broker credentials without a restart: the CA bundle
(--tls-ca), the revocation lists (--tls-crl), the server name, the client certificate and
key (--tls-cert, --tls-key) and the SASL password file (--sasl-password-file). The catalog and the signing key are not
broker credentials; a change of them needs a restart (see "Catalog changes need a restart").
- Replace the files at their paths. Keep the key and the password file private (mode 0600). Write the certificate and its key before the next step: ka2a reads both at the reload.
- Send
SIGHUPto theka2a runprocess (kill -HUP <pid>). Inpkg/ka2a, callNode.ReloadCredentials(ctx, ka2a.BrokerCredentials{TLS: ..., SASL: ...}). - Read the line on stderr, or the status.
ka2a runprints "broker credentials reloaded; new broker connections use them" with the end of the validity of the client certificate, or "broker credentials not reloaded (CLASS)" with the cause. Every refusal line names its class in place of CLASS, for exampleinvalid_credentialsfor an empty password file ortls_failed.
The node checks the new material before it uses it:
- TLS and SASL stay on or off, and the SASL mechanism stays. A
tls.ConfigofNode.ReloadCredentialsmust not turn off the verification of the broker certificate (InsecureSkipVerify) or lowerMinVersion. A change of them is refused withincompatible_credentials; restart the owner for such a change.Node.ReloadCredentialscan change the SASL user name;ka2a runreads only the files again and keeps--sasl-usernameuntil a restart. A new principal needs the ACLs ofka2a topics acl-planfirst. - The command reads the files with the rules of the start: a private key file that matches
the certificate, a certificate that allows client authentication and is valid now. The
node refuses an expired certificate or an empty password with
invalid_credentials. - The node sends one metadata request for its own mailbox with the new material to each
broker of
--brokers(Config.Kafka.Brokers), in parallel (at most 8 at a time), each through a separate connection, all within the request timeout. Each broker must accept the TLS handshake and the SASL authentication, and must allow Describe on the mailbox. A refusal of any broker givestls_failed,authentication_failedorauthorization_failedwith a fixed hint, also when another broker accepted the material: a broker with another trust store or user database would refuse the new connections after the swap. A broker that does not answer does not refuse the reload when another broker accepted it. When no broker answers, the class is that of the outage (for exampleconnection_refused) of the first broker, and the reload is refused too: the node never swaps to material that no broker accepted. Only the bootstrap list is checked: list every broker in--brokers(or a name per broker) so that the check reaches each trust store; a broker that only the cluster metadata names is seen at the next connection to it. - After the checks, the node swaps the material as a whole. Each new broker connection of
the producer, the consumer and the admin client uses it. An open connection keeps the
material of its handshake until it closes. Kafka does not authenticate an open TLS
connection again, and a SASL connection authenticates again only when the broker sets
connections.max.reauth.ms. So a broker that removes the old certificate or password from its trust store does not close open connections. Restart the owner when the old material must stop at once, for example after a leak. - A refusal changes nothing: the node keeps the previous material and its work.
Accepted, queued and in-flight work does not change in either case. No record is published
again, and no consumer group rebalance starts. The status (Status.Credentials), the
metrics and the run UI show the time and the result of the last reload and the end of the
validity of the client certificate:
| Metric | Meaning |
|---|---|
ka2a_credentials_last_reload{result} | 1 for the result of the last reload: none, reloaded or refused. |
ka2a_credentials_last_reload_timestamp_seconds | Time of the last reload, or 0. |
ka2a_credentials_reloads_total{result} | Reloads by result: reloaded or refused. |
ka2a_client_certificate_not_after_timestamp_seconds | End of the validity of the client certificate in use, or 0. Alert well before it. |
ka2a_revocation_lists_next_update_timestamp_seconds | Earliest next update of the revocation lists of the broker chain in use (--tls-crl), or 0 without lists. Publish the next lists and send SIGHUP before it. |
ka2a_catalog_disk{state} | The catalog file compared with the running catalog: same, differs or refused. |
Rotate a broker CA in two reloads: first add the new CA to the bundle and reload; then let
the broker change its certificate; then remove the old CA and reload again. A reload with
only the new CA before the broker changes its certificate is refused with tls_failed.
This procedure ran on a real broker (TestIntegrationBrokerCARotation, k0-decisions
section 31.2). Reload every node that talks to the broker in each step, also the nodes that
only answer requests.
Set connections.max.reauth.ms on the broker so that open SASL connections move to a new
password (k0-decisions section 31.1):
- Same user, new password: set the new password on the broker, then reload the nodes at once. SCRAM keeps one password per user, so the change removes the old password. At the next re-authentication, each open connection authenticates with the new password. A re-authentication between the change and the reload fails, and the client connects again after the reload.
- New user: create the user and its ACLs (
ka2a topics acl-plan), reload the nodes, wait for oneconnections.max.reauth.ms, then delete the old user. Kafka refuses a re-authentication that changes the principal. So each open connection fails once and the client connects again with the new user. No exchange fails. - A produce that the broker refuses because the authentication failed is no rejection: the
request stays locally queued, and the node sends it again after the reload
(
ka2a_publish_unknown_totalcounts the attempt; k0-decisions section 33.5). It expires at its horizon like work during an outage. Keep the time between the change on the broker and the reload short, or use a new user, which never has such a window. A denied ACL (authorization_failed) stays a rejection.
Check the connection
Check the connection with ka2a doctor or ka2a topics verify. When TLS or SASL refuses the
connection, the check broker.connection is a problem, the command exits with 3, and the
detail names the cause. The runbook lists each cause.
The output never holds the password or the client key.
Signing keys
Each endpoint has one Ed25519 signing key in a private file (format ka2a.signing-key.v1,
mode 0600). Make the key with ka2a keys generate:
install -d -m 0700 /etc/ka2a/keys
ka2a keys generate --endpoint example/billing --key-id billing-2026-10 \
--out /etc/ka2a/keys/billing-2026-10.key.json \
--not-before 2026-10-01T00:00:00Z --not-after 2027-10-01T00:00:00Z
- The command makes the key from the system random source and writes the file with mode
- It never replaces an existing file, also not through a symbolic link.
- The directory must be owned by the user who runs the command. Group and others must not
write it, and others must not read it. Use mode 0700, or 0750 for a service group. The
command refuses another directory with
unsafe_file. --endpointisDOMAIN/ENDPOINT.--not-beforedefaults to the current time. Without--not-afterthe key has no end.- The command prints the catalog key entry on stdout. Only the public key is in the output:
{
"key_id": "billing-2026-10",
"algorithm": "ed25519",
"public_key": "qu0s8m-HwO406XRi3sf7YKKrINtzpJ1gEsJkJXTkWDo",
"domain": "example",
"endpoint": "billing",
"not_before": "2026-10-01T00:00:00Z",
"not_after": "2027-10-01T00:00:00Z"
}
The public_key is the 32-byte Ed25519 public key in base64url without padding. Your key
has a different value. Add the entry to the keys list of the catalog in a reviewed change
(see "Catalog management"). To print the entry of an existing key file again, run
ka2a keys public /etc/ka2a/keys/billing-2026-10.key.json --endpoint example/billing --not-before ... [--not-after ...]. It reads the file with the rules of a node: a regular file
that you own with mode 0600.
Run the commands on the host of the owner, so the private key never leaves it. Never commit a
key file. Never send it in a log or a ticket. --output json prints the documents
ka2a.keys.generate/1 and ka2a.keys.public/1, with the entry in the member catalog_key.
Key rotation with validity windows
A catalog key has not_before, an optional not_after, revoked and replaced_by. A record
verifies only when its signed creation time is inside the window of its key. A retained
record that the old key signed inside its window still verifies after the rotation.
- Make the new key file with
ka2a keys generate --not-before <switch time>. Do not use it yet. - Add the printed entry to the catalog. Set
not_afterof the old key to a time after the switch. Leave a margin of at least the clock skew (default 5 minutes). Setreplaced_byof the old key to the new key ID. Without this link,ka2a catalog validatereports the two keys asoverlapping_validity. - Run
ka2a catalog validateon the new catalog. Distribute it to every endpoint. Restart each owner, so it loads the catalog. Until the restart, the status, the metrics (ka2a_catalog_disk{state="differs"}) andka2a doctorshow that the catalog on disk differs from the running digest. - At the switch time, restart the owner of the endpoint with the new key file.
- Keep the old key in the catalog until its records leave the retention. The longest signed lifetime is the response horizon (default 8 days) plus the clock skew.
- Remove the old key, or set
revokedwhen the key leaked. A revoked key never verifies, also for retained records.
Catalog management
The trusted catalog (format ka2a.catalog.v1) is the only source of peers, mailbox topics,
keys and grants. A discovery result never grants a permission.
- Keep the catalog under version control. Review each change by two people. ka2a has no command that edits the catalog: the catalog is the root of trust, and a rotation changes two keys in one reviewed change (k0-decisions section 27.5).
- Check each change with
ka2a catalog validate FILEbefore you distribute it. The command reports each finding with a severity, a class and a JSON path, and exits with 3 when it finds a problem. Problems are the defects that make a node refuse the catalog (for exampleduplicate_key,unknown_endpoint,rotation,unsafe_file), and two keys of one endpoint with overlapping windows and noreplaced_bylink (overlapping_validity). Warnings name keys that end within--expiry-daysdays (default 30) without a replacement that is valid at their end (key_expiring), and endpoints without a key that is valid now (no_signing_key).--strictcounts warnings as problems. Run it in the review pipeline of the catalog. ka2a catalog digest FILEprints the digest that a node records,sha256:<hex>of the exact bytes. Compare it on each host after the distribution.- Give the same catalog to every endpoint. A grant names the source, the destination and the
operations (
SendMessage,GetTask,ListTasks,CancelTask). ka2a init --catalogrecords the digest of the catalog as a configuration version.ka2a doctorwarns when the catalog differs from the recorded version (catalog.configuration). Runka2a initagain with the new catalog after you review the change.- A node loads the catalog at
Open. Restart the owner after a catalog change.
Catalog format
A catalog is one JSON file of at most 1 MiB. This catalog lets orders send the four
operations to billing. Each endpoint has one key, and the key of billing is the entry that
ka2a keys generate printed above:
{
"format": "ka2a.catalog.v1",
"mailbox": {
"prefix": "ka2a",
"rule": "sha256-v1"
},
"endpoints": [
{
"domain": "example",
"endpoint": "billing"
},
{
"domain": "example",
"endpoint": "orders"
}
],
"keys": [
{
"key_id": "billing-2026-10",
"algorithm": "ed25519",
"public_key": "qu0s8m-HwO406XRi3sf7YKKrINtzpJ1gEsJkJXTkWDo",
"domain": "example",
"endpoint": "billing",
"not_before": "2026-10-01T00:00:00Z",
"not_after": "2027-10-01T00:00:00Z"
},
{
"key_id": "orders-2026-10",
"algorithm": "ed25519",
"public_key": "_1xL8cugh1KnTLaF-VAkfgQWiDZ0BBLJeHamMYlBODY",
"domain": "example",
"endpoint": "orders",
"not_before": "2026-10-01T00:00:00Z",
"not_after": "2027-10-01T00:00:00Z"
}
],
"grants": [
{
"source": {"domain": "example", "endpoint": "orders"},
"destination": {"domain": "example", "endpoint": "billing"},
"operations": ["SendMessage", "GetTask", "ListTasks", "CancelTask"]
}
]
}
| Member | Rule |
|---|---|
format | Always ka2a.catalog.v1. |
mailbox | rule is always sha256-v1. prefix starts each derived mailbox topic name: <prefix>.v1.mailbox.<SHA-256 of "<domain>\n<endpoint>" in hex>. |
endpoints | 1 to 1024 endpoints. domain and endpoint match [a-z0-9][a-z0-9._-]{0,63}. The optional mailbox_topic replaces the derived topic name. The optional agent_card (name, version, description, sha256) is descriptive only and grants nothing. |
keys | At most 4096 entries, as ka2a keys generate prints them. Optional: not_after, revoked and replaced_by (see "Key rotation with validity windows"). |
grants | At most 16384 grants. Each grant names a source, a destination and its operations. A response needs no grant. |
The catalog accepts no unknown member. The node reads the file only when it is a regular file
(not a symbolic link) that root or the owner user owns, and that group and others cannot write.
ka2a catalog validate reports each of these rules. The test TestDocumentationJSONExamples
(internal/cli, make verify) checks every JSON example of this documentation with the product.
Catalog changes need a restart
A running node never adopts a changed catalog or signing key (k0-decisions section 30.4). The catalog is the root of trust of every record in flight: the grants and keys that admitted a request also verify its response, the store records one configuration version at each start, and the mailbox topic decides the consumer subscription. A restart gives one clear boundary, and it is crash-safe: queued work stays durable, and held work stays held.
The node compares the catalog file with the running digest at each retention interval (one
minute), at each credential reload and at Node.CheckCatalog. The status
(Status.Catalog), the metric ka2a_catalog_disk, the run UI and one log line show a file
that differs (with both digests) or that a node would refuse (refused, with the class;
a restart would fail). ka2a doctor reports catalog.configuration as a warning: "The
catalog on disk differs from the running digest" while an owner holds the store.
State directory ownership
ka2a initcreates the state directory with mode 0700. Its parent must exist. Create the parent as the service user.- Run every owner command as the user that owns the state directory:
run,send,backup,restore,quarantine release,held releaseandheld abandon. - Only one process can own the directory. Other commands that need the owner lock exit
with 5 while an owner runs. Read commands (
observe,doctor,snapshot export,serve-ui) open the store read-only and never change it. - The SQLite files hold message bodies. Treat the directory and every backup as secret.
File system of the state directory
- The owner open accepts these local file systems: on Linux ext2/3/4, xfs, btrfs, zfs, f2fs,
bcachefs, tmpfs and overlay; on macOS apfs and hfs. It refuses every other type (for
example NFS, CIFS or FUSE) and a read-only mount with
unsafe_storage. - tmpfs and overlay are accepted for development and tests, but they are not durable. tmpfs
loses the state at a reboot. overlay is the writable layer of a container; it is lost when
the container is replaced. Then the node loses its accepted work and its deduplication
state, and the new store enters quarantine with
broker_ahead(see Run ka2a in a container). ka2a doctornames the type instore.storageand reportsstore.file_system: a warning for tmpfs, overlay and a file system that the owner open refuses.Status.StoreFileSystemof the library shows the type of a running node, and a snapshot shows it instorage.file_system.- In production, put the state directory on a persistent local disk or volume. A bind mount or a volume shows the type of its backing file system, not overlay.
Backup and restore
ka2a backup is an offline backup. It needs the owner lock, so the owner must stop first. A
backup refuses (exit 5) while a node owns the state directory. There is no online backup
in this revision: an embedding application has no backup call either.
- Stop the owner.
ka2a rundrains within--drain-timeout(default 30s). - Run
ka2a backup --state-dir DIR --out FILE. The file must not exist. The directory ofFILEmust be owned by your user and must not be writable by group or others. Its parent directories must be owned by root or by your user. The command increments the live epoch after the copy. - Record the backup ID and the
sha256that the command prints (--output jsongiveska2a.backup/1). - Start the owner again.
- Store the backup with the same protection as the state directory. See Data protection.
To restore:
- Stop the owner.
- Run
ka2a restore --state-dir DIR --from FILE. The command keeps the replaced database under a new name (ka2a.db.pre-restore-<time>) and writes a recovery marker. - Start the owner. The store enters quarantine: no dispatch starts.
- Follow the runbook Quarantine after a restore.
A restore never gives exactly-once recovery. Work after the backup can be lost or repeated. The quarantine makes this visible.
Downtime and recovery point
- Downtime: each backup stops the endpoint for the drain, the copy and the start. The copy time grows with the size of the database. Measure it on a copy of your store. While the owner is stopped, requests to the endpoint wait in its mailbox topic. An embedding application cannot submit work while its node is closed.
- Recovery point: a restore returns the state of the backup time. The broker does not give
back the input after that time: the consumer group committed it, so the restored store
enters quarantine (
broker_aheador a restore reason). Requests that the endpoint sent after the backup are not in the restored store. The recovery point objective is the backup interval. - Trade-off: a shorter interval gives a smaller loss after a restore and more downtime. A
backup does not protect against the loss of a broker: the qualified profile does that
(replication factor 3,
min.insync.replicas2). A backup protects against the loss or damage of the state directory. - A restore is the last option. First try to keep the existing state directory. See Lost or damaged state directory.
A backup schedule
Run the backup in a quiet period. This example uses a systemd timer and a script that runs
as root. It stops the unit of Run ka2a as a service, writes the
backup as the user ka2a and starts the unit again, also after a failed backup.
/usr/local/sbin/ka2a-backup-billing (owner root, mode 0755):
#!/bin/sh
set -eu
dir=/var/backups/ka2a/billing # owner ka2a, mode 0700
out="$dir/billing-$(date -u +%Y%m%dT%H%M%SZ).db"
systemctl stop ka2a.service
trap 'systemctl start ka2a.service' EXIT
runuser -u ka2a -- /usr/local/bin/ka2a backup --state-dir /var/lib/ka2a/billing \
--out "$out" --output json > "$out.json"
# Keep 14 days of backups. Remove the older files with their reports.
find "$dir" -name 'billing-*.db*' -mtime +14 -delete
/etc/systemd/system/ka2a-backup-billing.service:
[Unit]
Description=Offline backup of the ka2a endpoint example/billing
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/ka2a-backup-billing
/etc/systemd/system/ka2a-backup-billing.timer:
[Unit]
Description=Daily offline backup of the ka2a endpoint example/billing
[Timer]
OnCalendar=*-*-* 03:17:00 UTC
Persistent=true
[Install]
WantedBy=timers.target
Enable it with systemctl enable --now ka2a-backup-billing.timer. With cron, run the same
script from the crontab of root, for example 17 3 * * * /usr/local/sbin/ka2a-backup-billing.
Verify a backup with a restore drill
A backup that you never restored is not verified. Do a restore drill regularly, for example once a month, and after each upgrade of ka2a.
- Compare the checksum:
sha256sum FILEmust equal thesha256of the backup report. - Restore into a scratch directory, never into the live state directory:
ka2a restore --state-dir /var/tmp/ka2a-drill/billing --from FILE. Create the parent directory first with mode 0700. - Run
ka2a doctor --state-dir /var/tmp/ka2a-drill/billing. Expectstore.openok with the schema version,store.identitywith the endpoint of the backup, andstore.restoreas a problem with the reasonsrestore_markerandrestored_backup. Exit code 3 is the expected result here. - Run
ka2a observe --state-dir /var/tmp/ka2a-drill/billingand compare the counts with your expectation. - Never start an owner (
ka2a runorpkg/ka2a) on the drill directory with the broker settings of the endpoint. It would join the consumer group of the live endpoint. - Remove the drill directory. It holds message bodies.
Retention and encryption of backups
- Keep backups for at least the longest retention of your stores (default 9 days) plus the time that you need to find a damage. Keep the backup of the last schema version before an upgrade until you do not need a rollback (see Upgrade notes).
- A backup holds message bodies in plaintext. Encrypt it before it leaves the host, for
example with
ageorgpgto a recipient key whose private key is not on the host. Remove the plaintext file after a successful encryption and a checksum comparison. - Keep the backups on another disk or host than the state directory.
- Delete expired backups with a method that fits your storage. A removed file on a copy-on-write file system or in object storage with versions can stay readable.
Retention budget sizing
Limits.RetainedBytes is the retained-state budget of a store (default 1 GiB). When an
admission would exceed the budget, it fails with storage_pressure. The store never removes
unresolved state to make room. Retention removes completed state after its retention period.
Size the budget from the bytes that each operation keeps and from the retention periods:
budget >= operations per second x bytes per operation x retention seconds x 1.5
- An operation keeps its request and response records, the frozen bytes and the task state.
Measure the bytes per operation on your payloads: run a load for some minutes and read the
retained bytes in
ka2a observe. - The default retention of results is 8 days (
Horizons.ResultRetention). The dedup retentions are 8 and 9 days. - The 30-minute soak of the qualification (20 operations per second, payloads from 64 B to 16 KiB) grew each store by about 10 MiB per minute. At that rate the default budget is full after about 100 minutes, before the first retention removes anything. That soak is not a capacity figure; measure your own rate.
- Keep 50 % free disk space beyond the budget for the SQLite WAL and the backups.
- Watch the metrics and
ka2a doctor(store.storage). Act before the budget is full; see Storage pressure.
Health and readiness
ka2a has no HTTP health or readiness endpoint. ka2a run --metrics-listen serves only
/metrics on a loopback address. An embedding application reads Node.Status or serves
Node.MetricsHandler itself. A scrape returns 503 with the error code when the node cannot
read its store.
| Question | pkg/ka2a | Metric |
|---|---|---|
| Does the node work? (liveness) | Run has not returned. Status.Running is true. | ka2a_node_running 1 |
| Can the node take work? (readiness) | WaitReady returned nil. Status.Ready is true and Status.Closed is false. | ka2a_node_ready 1 and ka2a_node_closed 0 |
| Does the broker answer? | Status.Broker.State is reachable. | ka2a_broker_state |
| Does a handler start? | Status.Quarantine.Active is false. | ka2a_quarantine_active 0 |
Use these rules:
- Liveness is the process.
ka2a runexits whenRunfails, and a supervisor restarts it (see Run ka2a as a service). In an embedding application, treat a non-nil error fromRunas fatal for the node: callClose, then exit orOpena new node. Do not add a probe that restarts a node that still runs. - Readiness is
Status.Ready. Ready means that restart recovery ran and the background loops run, soSubmitadmits work. It does not mean that a broker answered, unlessConfig.WaitForBrokeris set. - Do not restart on these conditions. A restart does not fix them, and it adds a drain and a
start:
Status.Broker.Stateisunreachable: the node reconnects by itself. See Broker outage.Status.Quarantine.Activeis true: an operator must decide. See Quarantine after a restore.storage_pressure, held work,ka2a_catalog_disk{state="differs"}or an unverified mailbox check.
- Alert with the rules of examples/alerts.yml. The metrics reference explains each metric and its threshold. This table names the runbook of each rule:
| Rule in alerts.yml | Runbook |
|---|---|
KA2ANodeDown, KA2ANodeNotReady | Lost or damaged state directory, Disk full |
KA2ABrokerUnreachable, KA2AOutboxStalled | Broker outage |
KA2AMailboxCheckFailed | Refused broker connection, Missing broker ACL |
KA2AHeldWork | Held work |
KA2AQuarantineActive | Quarantine after a restore, Store quarantined: broker_ahead |
KA2ARetainedBytesNearBudget | Storage pressure |
KA2AClientCertificateExpiresSoon, KA2ACredentialReloadRefused | Refused or expiring client certificate, Rotated SASL password |
KA2ARevocationListsDue | Revocation list due |
KA2ACatalogDiffers | Catalog on disk differs from the running digest |
KA2AConsumerRestarts | Consumer restarts |
KA2APublishRejected | Missing broker ACL |
- Also alert on two conditions that the rules do not cover: the free space of the file system of the state directory (Disk full), and the clock offset of the host (Clock drift and expired records). The error reference explains each error code, rejection reason and quarantine reason.
The metrics listener accepts only loopback addresses. Scrape it with an agent on the same host, for example a Prometheus agent or an OpenTelemetry collector.
Run ka2a as a service
docs/examples/ka2a.service is an example systemd unit for one
endpoint with ka2a run. It does these things:
- It runs as the dedicated system user
ka2a, without capabilities, withNoNewPrivileges=yes,ProtectSystem=strictand a system call filter. StateDirectory=ka2a/billingwithStateDirectoryMode=0700is the only writable path.TimeoutStopSec=60sis longer than--drain-timeout(30s) plus the SQLite busy timeout (5s). A shorter stop timeout kills the drain. The data stays safe, but the next start holds the handler calls that were running.ExecReloadsendsSIGHUP.systemctl reload ka2areloads the broker credentials and checks the catalog file.Restart=on-failurerestarts after a fatal error.RestartPreventExitStatus=2 4 5stops the restarts after a usage error and when another process owns the state directory.- It starts after
time-sync.target. See Clock drift and expired records.
systemd-analyze verify accepts the unit (systemd 255, with the binary path changed to an
existing file), and systemd-analyze security rates its exposure 1.4 ("OK"). The unit did not
run under systemd in a test. Check it on your distribution before you go live.
Do not use DynamicUser=yes. ka2a requires that the effective user owns the state files
and the secret files.
Run ka2a in a container
docs/examples/Containerfile builds a static CGO_ENABLED=0 binary
and copies it into gcr.io/distroless/static-debian12:nonroot. The project does not publish
an image, and no test built this file.
- Keep the state directory on a persistent volume of a local file system (ext4 or xfs). Do
not keep it in the writable layer of the container. The store accepts the
overlayfile system, but that layer is lost with the container. A new container then starts with an empty store, and the store enters quarantine withbroker_ahead.ka2a doctorwarns (store.file_system) when the state directory is on overlay. - Run one container per endpoint. Never mount one volume in two containers, also not on two hosts: the owner lock does not work across hosts.
- Set the stop timeout of the container runtime (for example
docker run --stop-timeout 60) above--drain-timeoutplus 5s. - The metrics listener accepts only loopback addresses, so a probe from outside the container cannot reach it. Use the exit of the process as the liveness signal.
Data protection
ka2a signs the records. It does not encrypt the message bodies. The bodies are in plaintext in these places:
| Place | How long | Control |
|---|---|---|
| Kafka mailbox topics | retention.ms of the topic. The node refuses a topic whose retention is shorter than the longest record horizon (default 192h). | Keep the retention at the minimum (ka2a topics provision sets 192h). Do not set retention.ms=-1. Use broker TLS. Encrypt the broker disks. Limit the topic ACLs (ka2a topics acl-plan). |
| SQLite store in the state directory | Until retention removes the completed state (default 8 or 9 days). Held and unresolved work stays until an operator resolves it. | Mode 0700 and 0600 (the store refuses other modes). Encrypt the disk, for example with LUKS. |
| Free pages and the WAL of the store | The owner sets secure_delete=ON: SQLite overwrites deleted rows with zeros. Old pages in the WAL stay until SQLite overwrites the WAL. | Encrypt the disk. A store that a version before secure_delete wrote can have old free pages: a backup and a restore write a new file without them. |
| Backups | Until you delete them. | See Retention and encryption of backups. |
ka2a.db.pre-restore-<time> files | ka2a restore keeps the replaced database. | Delete it after the quarantine release when you do not need it. |
Snapshots (ka2a snapshot export) | Until you delete them. | A snapshot holds no bodies by default. --raw-ids adds raw identifiers. |
Logs, metrics, the UI and the default snapshots hold no bodies (KA2A-R1-017).
There is no call that removes one message before its retention ends. Node.Purge applies
the retention once; it never removes state before its retention period or unresolved state.
If you must remove a body at once, you must remove the store and the topic records. This
loses the deduplication state of the endpoint. Do not put data in a message that you must be
able to erase before the retention ends.
Go-live checklist
Do each step before the first production start of an endpoint. The threat model explains why.
- Read the section "What was not executed" of the newest release-candidate report.
- Build the binary from a reviewed commit with
make release-artifacts. Runmake release-checkandmake vulncheck. See Releasing. - Use the durability profile
qualified: 3 brokers, replication factor 3,min.insync.replicas2,unclean.leader.election.enable=false. - Use broker TLS. Use SCRAM or mutual TLS. Never use
--allow-plaintextor--allow-plaintext-credentials. - Give each endpoint its own broker principal with only the ACLs of
ka2a topics acl-plan. Runka2a topics verifywith the credentials of the endpoint. - With
--tls-crl, plan the publication of the revocation lists before their next update. - Make each signing key on the host of its owner (
ka2a keys generate). Keep it with mode 0600, owned by the service user. Give each key anot_afterand plan its rotation. - Review the catalog. Distribute it as a file owned by root, mode 0644. Compare
ka2a catalog digeston each host with the reviewed value. - Put the state directory on a local disk of the allowlist (ext4 or xfs), on an encrypted volume, owned by the service user, mode 0700. One host per endpoint.
- Synchronize the clock of every host with NTP. Alert when the offset is above 1 minute.
- Size
Limits.RetainedBytesand the disk (see Retention budget sizing). Keep 50 % free disk space beyond the budget. - Install the service with a stop timeout above the drain timeout plus 5s (Run ka2a as a service).
- Scrape the metrics on the host and add the alerts of Health and readiness.
- Schedule backups, encrypt them and do one restore drill (Backup and restore).
- Run
ka2a doctorwith the flags of the owner. Expect no problem. - Give each handler the recovery class
HoldOnAmbiguity, unless the handler deduplicates the operation ID. Name the people who resolve held work and quarantines. - Read the runbooks and the security policy.