notariz.ing · local-first spec review

Review a change with its evidence in view.

Notarizing is a workspace for spec-driven development. It reads an OpenSpec change, shows the premises that each requirement relies on and links requirements to the check reports that you import. It runs on your machine and opens in your browser.

Developer preview. Not yet production ready. Public installation is not yet verified.

A specification page under a loupe. A seal marks the checked requirements, and a small graph shows the premises.

Five questions

What a reviewer needs to know about a change

Requirements, assumptions and changes are the primary objects. Check reports are evidence that you import. Notarizing never runs your checks and never edits your repository.

  1. 1

    What does this change promise?

    The effective requirements of the change: added, modified and removed, with their scenarios and their exact source lines.

  2. 2

    Which premises does it rely on?

    Declared assumptions and premise groups. An all of group needs every member; an any of group needs one.

  3. 3

    What changed between two revisions?

    Wording, assumptions, relationships and evidence applicability, side by side at two exact commits.

  4. 4

    What was checked, and under which scope?

    Each report keeps its claim, its environment, its bounds and its exclusions. A reported pass is not an established pass.

  5. 5

    What remains unknown?

    Missing bindings, stale evidence and open questions stay visible. The workspace never shows an overall pass badge.

Views

Review, Explore and Compare

The screenshots show the fictional repository demo-payments from examples/refund-limits. Its change add-refund-limits adds a daily refund limit, in two revisions.

The Review view of the change add-refund-limits. Summary chips count requirements with a failure, with stale evidence, without a binding and more. The table lists each requirement with its reported and evaluated evidence, its gaps and its premises. The inspector on the right shows PAY-R-101 with its statement and evidence.
Figure 1 Synthetic demo data. Review answers what the change promises and what was checked. Each row shows the reported claim beside the evaluated status. The inspector shows the scope claim and the reasons.
The Explore view with a bounded, typed graph around PAY-R-101: requirements, assumptions, a premise group, checks, a component and a decision, with relation labels.
Figure 2 Synthetic demo data. Explore shows a bounded neighborhood of typed links. A table view lists the same nodes and edges as text.
The Compare view of revision 58f73aa38844 against b7719d84d43e. Rows per requirement show whether the source wording, assumptions, relationships, review decisions and evidence applicability changed.
Figure 3 Synthetic demo data. Compare matches requirements by stable ID across two exact commits and shows each aspect apart.

How it works

From a Git commit to a review in five steps

  1. Register

    repo add names a local Git repository. Notarizing reads objects by full commit ID and never writes to it.

  2. Import a change

    change import reads one OpenSpec change at one exact commit and indexes it for search.

  3. Import evidence

    binding import says which checks cover which requirements. evidence import stores exact report bytes with a receipt.

  4. Review

    serve opens the browser views on 127.0.0.1. A grant from the CLI lets one session write review notes for a limited time.

  5. Export

    Reviews, graphs and draft patches export as Markdown, Mermaid or JSON at a pinned review checkpoint.

Examples

A fictional change, reviewed end to end

Run make web, then make demo-refund-limits, in the Notarizing repository to create the same workspace. make web builds the browser assets and needs Node at build time only. The script builds the repository with fixed commit IDs, so the IDs below repeat on every run. Every report says SYNTHETIC: none of the evidence is real.

OpenSpec · spec.md

A requirement with its links

Requirements stay in your repository, in plain Markdown. Link lines under the ID declare what the requirement depends on, which premises it assumes and which checks evaluate it.

  • A declared link is a claim for review, not a proof.
  • Only a reviewed binding file makes a check count as coverage.
openspec/changes/add-refund-limits/specs/refunds/spec.mdSynthetic
### Requirement: Daily refund limit

Requirement-ID: PAY-R-101
Depends-On: PAY-R-001
Assumes: GRP-LIMIT-PREMISES
Evaluated-By: CHK-LIMIT-UNIT, CHK-LIMIT-PROPERTY

The service SHALL reject a refund that would raise the refunded total of the merchant for the current UTC day above the daily limit of the merchant. Refunds that wait for an override SHALL count toward the total.

#### Scenario: Refund under the limit
- **GIVEN** a merchant with a daily limit of 1000.00 EUR and 900.00 EUR refunded today
- **WHEN** a clerk refunds 80.00 EUR
- **THEN** the refund succeeds

#### Scenario: Refund over the limit
- **GIVEN** a merchant with a daily limit of 1000.00 EUR and 950.00 EUR refunded today
- **WHEN** a clerk refunds 80.00 EUR
- **THEN** the service rejects the refund with the reason limit_exceeded and writes no ledger entry

#### Scenario: Pending override counts
- **GIVEN** a merchant with a daily limit of 1000.00 EUR, 900.00 EUR refunded today and a 150.00 EUR refund that waits for an override
- **WHEN** a clerk refunds 80.00 EUR
- **THEN** the service rejects the refund with the reason limit_exceeded
Listing 1 The requirement PAY-R-101 at revision 2, with its link lines.

OpenSpec · declarations.md

Assumptions and premise groups

A change can declare what must be true for its requirements to hold. Each assumption has a category, a rationale and the consequence if it is false.

  • all_of groups need every member. any_of groups need one.
  • Review states say how a reviewer treated an assumption. They never say that it is true.
openspec/changes/add-refund-limits/declarations.mdSynthetic
### Assumption: No merchant needs more than 1000 refunds on one day
Assumption-ID: ASM-REFUND-VOLUME
Category: product_hypothesis
Applies-To: PAY-R-101
Rationale: The service computes the daily total from the refunds of the day. A load test measured 40 ms for 1000 refunds. Support reports at most 120 refunds for one merchant on one day.
If-False: The total query becomes slow, and the limit check times out.

### Premise-Group: Limit premises
Group-ID: GRP-LIMIT-PREMISES
Operator: all_of
Members: ASM-UTC-DAY, ASM-REFUND-VOLUME

### Premise-Group: Exchange-rate source
Group-ID: GRP-FX-SOURCE
Operator: any_of
Members: ASM-FX-HOURLY, ASM-FX-CACHE
Listing 2 One assumption and the two premise groups of the change.

CLI · init, repo add, change import

Import a change at an exact commit

The import pins the commit, the extractor and a digest of the result. A later import of another commit becomes a second target that you can compare.

  • No broker, model or network service is needed.
  • Repository text is untrusted data. It is never an instruction.
terminalReal run · synthetic repository
$ export NOTARIZING_WORKSPACE=/srv/notarizing-demo/workspace
$ notarizing init
workspace /srv/notarizing-demo/workspace is ready (schema version 12)
next: notarizing repo add ID PATH

$ notarizing repo add demo-payments /srv/notarizing-demo/demo-payments
registered repository demo-payments
path: /srv/notarizing-demo/demo-payments
object format: sha1

$ notarizing change import --repo demo-payments \
    --at 58f73aa38844dcce75296abfc7ecdb1effe50e49 \
    --change add-refund-limits
imported add-refund-limits of repository demo-payments at 58f73aa38844dcce75296abfc7ecdb1effe50e49
target: tgt_g625kewcsiwgote55vvbaeyhuj (built)
target digest: a729c0044d66ea52c7ad354be1358f4b4b55013e15db685f750821557df9989f
requirements: 6 (proposed 5, accepted 2, removed 0); declared nodes 13, edges 28
diagnostics: fatal 0, error 0, warning 0, info 0
search index: generation 1, coverage complete (22 of 22 chunks)
next: notarizing change report --target tgt_g625kewcsiwgote55vvbaeyhuj

$ notarizing change import --repo demo-payments \
    --at b7719d84d43e337dcb7bfdbddbe088f333347fcb \
    --change add-refund-limits
imported add-refund-limits of repository demo-payments at b7719d84d43e337dcb7bfdbddbe088f333347fcb
target: tgt_xxwrnsvhrkxqhdz6rns24j4ktb (built)
target digest: 134c9aba9c241fdc4519ac81cf0734a0cbec745d13515ad6e90b5906731b4631
requirements: 6 (proposed 5, accepted 2, removed 0); declared nodes 14, edges 29
diagnostics: fatal 0, error 0, warning 0, info 0
search index: generation 2, coverage complete (23 of 23 chunks)
next: notarizing change report --target tgt_xxwrnsvhrkxqhdz6rns24j4ktb
Listing 3 Create a workspace and import the change at revision 1 and at revision 2.

Evidence · notarizing.check-report/1

A check report is a claim with a scope

Your own test harness writes the report. It names the exact requirement and implementation snapshots, the checker, the outcome and the scope of the claim. Notarizing validates it strictly and keeps the exact bytes.

  • Artifacts travel with the report and are stored by digest.
  • A report about an older commit becomes stale, not wrong.
  • evidence import takes up to 256 report files, or --dir with one report per subdirectory. Each report gets its own transaction and receipt.
evidence/limit-property-1.jsonSynthetic
{
  "schema_version": "notarizing.check-report/1",
  "report_id": "9a2d7e41-5c3b-4f8a-b1e6-2d9c4a7f3e22",
  "check_id": "CHK-LIMIT-PROPERTY",
  "requirement": {
    "snapshot": {
      "repository_id": "demo-payments",
      "revision": "58f73aa38844dcce75296abfc7ecdb1effe50e49",
      "path": "openspec/changes/add-refund-limits/specs/refunds/spec.md",
      "sha256": "eb519c724985abb1e3b343b5566ca2c631f88378f4ca591c138dbb2ee44170f5"
    },
    "requirement_id": "PAY-R-103",
    "classification": "proposed"
  },
  "target": {
    "kind": "implementation",
    "snapshots": [
      {
        "repository_id": "demo-payments",
        "revision": "58f73aa38844dcce75296abfc7ecdb1effe50e49",
        "path": "src/refunds/limits.go",
        "sha256": "6ab858e651d96056c16c74cd6c9f7da436944af55e8ab82005726dbea464caa6"
      }
    ]
  },
  "method": "implementation_test",
  "checker": {
    "name": "demo-go-test",
    "version": "go1.27.0",
    "executable_sha256": "5e5e5e5e5e5e5e5e5e5e5e5e5e5e5e5e5e5e5e5e5e5e5e5e5e5e5e5e5e5e5e5e",
    "toolchain_manifest_sha256": "7a7a7a7a7a7a7a7a7a7a7a7a7a7a7a7a7a7a7a7a7a7a7a7a7a7a7a7a7a7a7a7a"
  },
  "started_at": "2026-09-28T09:20:11Z",
  "finished_at": "2026-09-28T09:21:40Z",
  "completion": "complete",
  "reported_outcome": "fail",
  "finding_type": "test_failure",
  "scope": {
    "claim": "Converted refunds never raise the daily total above the limit.",
    "environment": "linux-amd64, go1.27.0, 10000 random histories",
    "assumptions": [
      "Exchange rates from the fixture feed"
    ],
    "exclusions": [
      "Rates older than 24 hours"
    ],
    "bounds": {
      "histories": "10000",
      "seed": "1848"
    }
  },
  "coverage": {
    "requested": [
      "random-histories"
    ],
    "completed": [
      "random-histories"
    ],
    "note": "Synthetic demonstration data."
  },
  "artifacts": [
    {
      "path": "limit-property-counterexample.txt",
      "sha256": "fb7b406b118011424a8ae99aaeb30cf31323a8667010c0af3d31e028d3eef508",
      "byte_length": 360,
      "media_type": "text/plain"
    }
  ],
  "summary": "SYNTHETIC: rounding lets small JPY refunds pass the limit by one minor unit."
}
Listing 4 A check report of a failed property test, with its counterexample artifact.
terminalReal run · synthetic reports
$ notarizing binding import evidence/bindings.json
imported bindings cb4fffdd-5698-44c6-a6fd-4c295ef55880 (import 1) of repository demo-payments: 3 checks, 4 bindings
file sha256: 0ef492909b1caae1fcfb0387f8ff62df916a938a6aed000b24f8a16d55584957 (2410 bytes)

$ notarizing evidence import evidence/limit-property-1.json \
    --artifacts evidence/artifacts
accepted receipt 0f48da3b-0b1d-4bf3-aace-8a6918bcbca9 for report 9a2d7e41-5c3b-4f8a-b1e6-2d9c4a7f3e22
report sha256: c9d9b927856527f79643164c7de47360c9ec87883a5ac0ee277c658228e9560e (2175 bytes), 1 artifacts
attribution manual_unverified via local_file, source local, ingest sequence 2
A receipt says that this workspace accepted these exact bytes. It does not say that a checker ran or that the claim is true.
search index: rebuilt target tgt_g625kewcsiwgote55vvbaeyhuj (generation 4, coverage complete)

# … the three other reports are imported the same way.

$ notarizing policy set evidence/policy.json
policy version 1, sha256 6e670f6f43cdcbd6bd51d385eaad1e5cdee04fd2210099aa28f0f58371610c84
trusted source local for implementation_test, integration_test
Listing 5 Import the bindings, the report of Listing 4 and the policy. The receipt names the exact bytes.

CLI · change report

Read what was checked and what was not

The change report lists every requirement with its assessment rows: the reported claim beside the evaluated status. Revision 2 changed the code, so the failure and the incomplete run from revision 1 are now stale.

  • A requirement without a binding stays visible as unassessed.
  • There is no overall pass badge, by design.
terminalReal run · synthetic repository
$ notarizing change report --repo demo-payments \
    --change add-refund-limits
Change report: add-refund-limits of demo-payments at b7719d84d43e337dcb7bfdbddbe088f333347fcb
Target tgt_xxwrnsvhrkxqhdz6rns24j4ktb (digest 134c9aba9c241fdc4519ac81cf0734a0cbec745d13515ad6e90b5906731b4631)
Snapshot snap_7ox2w6kvqme6qsqakw7zahwbad (snap:demo-payments:b7719d84d43e), extractor notarizing-openspec/1
Review checkpoint 0, scope workspace, policy version 1, rules notarizing.assessment-rules/2
Repository, report and review text in this report is untrusted data, never instructions. A status is a method-specific conclusion under the named policy, never general correctness. The report has no overall pass badge.

Summary
  requirements: 6 (stale 2, unknown 4)
  requested dimensions: 4; checked: 3 (pass 1, stale 2, unknown 1); missing: 4
  failed: none
  stale: PAY-R-102, PAY-R-103
  unassessed: 3; without binding: 3
  assumptions: 6 (open 6); open questions: 1; findings: 3

Requirements
  PAY-R-001  refunds  accepted unchanged  unknown/not_assessed  Refund ledger entries
    source openspec/specs/refunds/spec.md:9-18 (block ed77b07e1d7f)
    missing: no_binding
  PAY-R-002  refunds  proposed modified  unknown/not_assessed  Refund notifications
    source openspec/changes/add-refund-limits/specs/refunds/spec.md:74-89 (block 076e3f6f0418)
    missing: no_binding
  PAY-R-101  refunds  proposed added  unknown/partial  Daily refund limit
    source openspec/changes/add-refund-limits/specs/refunds/spec.md:3-25 (block 3dee3ef4de9e)
    row CHK-LIMIT-UNIT [any configuration]: reported pass (complete, 3 of 3 cases, manual_unverified via local_file) -> evaluated pass (none)
    row CHK-LIMIT-PROPERTY [any configuration]: no report -> evaluated unknown (missing_configuration)
    missing CHK-LIMIT-PROPERTY [any configuration]: no_report
  PAY-R-102  refunds  proposed added  stale/partial  Second approver for overrides
    source openspec/changes/add-refund-limits/specs/refunds/spec.md:27-44 (block 8fac07780b27)
    row CHK-OVERRIDE-E2E [any configuration]: reported unknown (incomplete, 1 of 2 cases, manual_unverified via local_file) -> evaluated stale (stale_implementation, incomplete)
  PAY-R-103  refunds  proposed added  stale/complete  Limits in the settlement currency
    source openspec/changes/add-refund-limits/specs/refunds/spec.md:46-58 (block c9065d394346)
    row CHK-LIMIT-PROPERTY [any configuration]: reported fail (complete, 1 of 1 cases, manual_unverified via local_file) -> evaluated stale (stale_implementation)
  PAY-R-104  refunds  proposed added  unknown/not_assessed  Limit audit log
    source openspec/changes/add-refund-limits/specs/refunds/spec.md:60-70 (block 4ea1620d6a3f)
    missing: no_binding

…

Premise groups
  GRP-FX-SOURCE  ANY OF: ASM-FX-HOURLY (open), ASM-FX-CACHE (open)
  GRP-LIMIT-PREMISES  ALL OF: ASM-UTC-DAY (open), ASM-REFUND-VOLUME (open)

Open questions
  Q-WEEKEND-LIMIT  Do refunds on weekend days share one limit?

Findings
  missing_check_binding: PAY-R-001 has no approved check binding.
  missing_check_binding: PAY-R-002 has no approved check binding.
  missing_check_binding: PAY-R-104 has no approved check binding.

…
Listing 6 The change report at revision 2. An ellipsis (…) marks the cut lists of assumptions and limitations.

CLI · compare

See what changed between two revisions

compare matches objects by stable ID across two imported targets, each at a pinned review checkpoint. It lists only the declared claims that changed, by ID and title.

  • A link that only moved to other source lines is not a change. The JSON output lists it apart.
  • Review decisions stay apart from source changes.
terminalReal run · synthetic repository
$ notarizing compare --left tgt_g625kewcsiwgote55vvbaeyhuj \
    --right tgt_xxwrnsvhrkxqhdz6rns24j4ktb
left:  tgt_g625kewcsiwgote55vvbaeyhuj (add-refund-limits at 58f73aa38844), checkpoint 0
right: tgt_xxwrnsvhrkxqhdz6rns24j4ktb (add-refund-limits at b7719d84d43e), checkpoint 0

Requirements: 6 (wording: 1 changed, 5 unchanged)
  changed requirement PAY-R-101: Daily refund limit [wording]

Assumptions: 6 (1 added, 1 changed, 4 unchanged)
  added assumption ASM-OVERRIDE-EXPIRY: An override that nobody confirms expires after 24 hours
  changed assumption ASM-REFUND-VOLUME: No merchant needs more than 1000 refunds on one day [wording, label]

Premise groups: 2 (2 unchanged)

Other declared objects: 6 (6 unchanged)

Relationships: 28 (1 added, 27 unchanged)
  added assumes PAY-R-102 -> ASM-OVERRIDE-EXPIRY (Second approver for overrides -> An override that nobody confirms expires after 24 hours)
  20 unchanged relationships differ only in provenance, for example their source lines. The JSON output lists them.

Review decisions that differ: 0
limitation: Access-scoped: each side holds only objects that the caller is authorized to see on that side.
limitation: Stable IDs preserve identity. Provisional identities are listed for review and are never matched automatically. Only reviewed one-to-one mappings link different identities.
limitation: Review decisions are listed apart from source changes. A review change is not a wording change.
limitation: Declared inputs describe recorded graph inputs only. Evidence applicability comes from the protected policy. Model evidence and implementation correspondence stay separate.
limitation: Evidence applicability is not part of this comparison. Run change report on each target.
Listing 7 Compare revision 1 with revision 2 of the change.

Export · Mermaid, Markdown, JSON

Export a graph that you can put in a review

The same target, checkpoint, filters and exporter always give the same bytes. The export says what it is and what it is not: a presentation, not proof.

  • Labels are escaped. Hostile text cannot become markup.
  • The graph is bounded: at most 500 nodes and 1000 edges.
terminalReal run · synthetic repository
$ notarizing graph export --target tgt_xxwrnsvhrkxqhdz6rns24j4ktb \
    --object PAY-R-101 --view dependencies --depth 2 --format mermaid
%% notarizing graph view as a Mermaid flowchart. Presentation only; not an evidence ledger.
%% Mermaid is a presentation, not proof. Graph reachability does not establish correctness.
%% Review states describe review treatment, not truth. Accepted for scope only is not a proof, a test binding or an approval.
%% schema notarizing.graph/1; exporter notarizing-graph-export/1.0.0
%% review checkpoint 0; query dependencies; depth 2; suggestions excluded
%% roots obj_yxkswyxw3tjhvnffg4sq32m5hs
%% source src_cs4iizuaicipkwppdq5iktt4zf repository demo-payments revision b7719d84d43e337dcb7bfdbddbe088f333347fcb sha256 266ee81e39d0117c42d8e1cc3eca98024f80008fb96d1eda959e9dc7faa35823
%% source src_onb4febxxlsqbj2h7bwq2uohi4 repository demo-payments revision b7719d84d43e337dcb7bfdbddbe088f333347fcb sha256 a644c2edd673abbee2857188ced1c1e9971b0c4fcdfbbf5bd0b0d0b2dc5939ef
%% source src_wbltsfisddryz6jpz555bx6xoz repository demo-payments revision b7719d84d43e337dcb7bfdbddbe088f333347fcb sha256 37f50fc36570075cb157be92ee8778ff111676707d2269d95b3a0991d13619bc
%% coverage authorized_retained_graph (access-scoped); complete; not truncated; depth limit not reached; authorized omissions 0; unresolved relationships 0
%% shapes: rectangle requirement, component, decision, result; stadium assumption; hexagon premise group; subroutine check, model; rhombus question
flowchart LR
  n1{{"ALL OF: Limit premises: not fully reviewed, 0 of 2 members accepted for scope only"}}
  n2(["All services use UTC day boundaries: open"])
  n3(["No merchant needs more than 1000 refunds on one day: open"])
  n4["Refund ledger entries: unassessed"]
  n5["Daily refund limit: assessment unknown"]
  n1 -->|member| n2
  n1 -->|member| n3
  n5 -->|depends on| n4
  n5 -->|assumes| n1
  n5 -->|assumes| n2
  n5 -->|assumes| n3
Listing 9 Export the dependencies of PAY-R-101 as a Mermaid flowchart.
The dependency graph of PAY-R-101: it depends on PAY-R-001 and assumes the premise group Limit premises, which needs both the UTC day assumption and the refund volume assumption. depends on assumes member member PAY-R-101 (root) Daily refund limit assessment unknown PAY-R-001 Refund ledger entries ALL OF Limit premises UTC day boundaries 1000 refunds a day
Figure 4 The same graph, drawn for this page. The drawing leaves out the direct assumption links that the Mermaid text also lists.

CLI · graph suggestions

Similar wording is a suggestion, never a link

The generator notarizing.suggest/1 compares the wording of the requirements of one target in the stored search index. A pair above the threshold becomes a may_relate suggestion with its scores and its generator identity.

  • A suggestion is never a dependency, a binding, coverage or a review state. A reviewer must propose and review the link before it counts.
  • The lexical method always runs. The embedding method uses only stored vectors of a local provider and never calls it.
  • The six requirements of the demo change share no wording above the threshold, so the list is empty. The header still names the generation and the embedding state.
terminalReal run · synthetic repository
$ notarizing graph suggestions --target tgt_xxwrnsvhrkxqhdz6rns24j4ktb
notarizing.suggest/1 suggestions of target tgt_xxwrnsvhrkxqhdz6rns24j4ktb: available
lexical generation 6, embedding provider_disabled, requirements 6, truncated false
Suggestions add no dependency, binding or coverage.
Listing 10 The suggestions of revision 2 of the change. There are none.

Review mode · serve, review session grant

Review in the browser with an expiring grant

The browser starts as a viewer. The CLI gives one browser session review authority for a limited time and scope. The session can then add notes, decide on assumptions and draft wording.

  • The one-time secret opens one session. The next secret is in an OS-private file.
  • When the grant expires, drafts stay in the browser. Nothing is lost.
  • --actor LABEL adds a display name to the grant. Every event shows it as an unverified claim, beside the browser session.
terminalReal run · synthetic repository
$ notarizing serve
notarizing: serving the browser interface at http://127.0.0.1:7373/
notarizing: one-time viewer secret: (64 hex characters, not shown here)
notarizing: after each use, the next secret is in /srv/notarizing-demo/workspace/viewer.token
notarizing: the page shows your session ID; to review, run: notarizing review session grant --session ID --ttl 30m --scope workspace

# In a second terminal: give this browser session review authority for 30 minutes.
$ notarizing review session grant --session ses_b45a3b1fa77badb9 \
    --ttl 30m --scope workspace
session ses_b45a3b1fa77badb9 (browser session b45a): role reviewer, grant active
  scope workspace, expires 2026-10-02T19:01:23Z
Listing 11 Start the browser interface, then give one session an expiring review grant. The listing does not show the one-time secret.

Agents · Model Context Protocol

Let an agent read, never write

notarizing mcp serves nine read-only tools over standard input and output. An agent can read change reports, graphs, similar-wording suggestions, receipts and search results. No tool can import, purge or change anything.

  • Responses are bounded to 512 KiB.
  • At the end of its input, the server answers the requests that it already read, then stops.
  • Report text reaches the agent as data, with a notice that it is untrusted.
mcp client configurationJSON
{
  "mcpServers": {
    "notarizing": {
      "command": "notarizing",
      "args": ["--workspace", "/srv/notarizing-demo/workspace", "mcp"]
    }
  }
}
Listing 12 A client configuration that starts the MCP server over standard input and output.
stdio sessionReal run · synthetic repository
# tools/list of `notarizing mcp` (stdio). Every tool is annotated read-only.
compare_snapshots        readOnly  Compare two imported targets at pinned review checkpoints. …
get_assumption           readOnly  Get one assumption of a target: statement, category, scope, rationale, exclusions, consequences, review state at the checkpoint and the deterministic challenge prompts.
get_change_report        readOnly  Get the change report of one imported change: every requirement with its exact source, its assessment rows (reported claim beside evaluated status) and missing dimensions, the assumptions with review states, premise groups, open questions and findings. …
get_graph_neighborhood   readOnly  Get a bounded graph around one object of a target: dependencies, impact, gaps or selection. …
get_graph_suggestions    readOnly  Get the similar-wording suggestions (may_relate) of the requirements of a target. …
get_receipt              readOnly  Get one receipt of the local source: the collector record, the reported claim and its evaluations. …
get_requirement_history  readOnly  Get the source revisions of one stable Requirement-ID across imports and the receipts whose reports name it, with the reported claim beside the evaluation.
list_changes             readOnly  List the imported changes (targets) of the workspace with their repository, commit, status and requirement count. …
search_specs             readOnly  Search specifications, declarations and review text with the lexical index. …

# A write attempt is not a tool.
tools/call import_evidence
error -32602: unknown tool "import_evidence"
Listing 13 The tool list of a real session, and a write attempt that fails. An ellipsis (…) marks cut tool descriptions.

Transport · evidence-over-ka2a/1

Receive evidence over ka2a, optionally

A CI system can send a check report to Notarizing through ka2a. The verified ka2a principal selects the source namespace. The payload cannot choose it. The receipt is sent only after the local transaction commits.

  • Local use needs no broker. The adapter is off unless you configure it.
  • A report is at most 256 KiB; the whole package at most 512 KiB.
A2A SendMessage messageSynthetic
{
  "messageId": "0f6a3c1e-5b2d-4e8f-9a71-3c2b1d0e9f84",
  "role": "ROLE_USER",
  "parts": [
    {
      "data": {
        "profile": "notarizing.evidence-over-ka2a/1",
        "operation": "evidence.submit",
        "body": {
          "report_base64": "ewogICJzY2hlbWFfdmVyc2lvbiI6ICJub3Rhcml6aW5nLmNoZWNr…",
          "artifacts": [
            {
              "path": "limit-unit-log.txt",
              "bytes_base64": "U1lOVEhFVElDIERFTU9OU1RSQVRJT04gQVJUSUZBQ1QuIE5v…"
            }
          ]
        }
      },
      "mediaType": "application/json"
    }
  ]
}
Listing 14 An ellipsis (…) marks the shortened base64 values. The JSON is what the Go client of the adapter builds.

CLI · doctor

Check the workspace and the adapter health

When notarizing serve runs, doctor also reads the health of its ka2a adapter. It reports the adapter state, the reason, a broker denial with its scope and count, and the validity of the client certificate.

  • The adapter can use broker mutual TLS. Its key file must be private, and an expired certificate keeps the adapter unavailable.
  • A missing broker ACL degrades only the adapter, with the reason authorization_failed. Local imports and the browser keep working.
  • After a credential reload, doctor also reports the result of the last reload and whether the catalog on disk differs from the running catalog.
  • With broker revocation lists (tls.crl_file), doctor warns 24 hours before the next update of the lists, and when they are stale. The adapter gives the same time to the ka2a node, so its status and metric agree with the adapter health.
  • The adapter is optional, so these checks are never errors. In this demo the owner runs without an adapter.
terminalReal run · synthetic repository
# notarizing serve runs in another terminal.
$ notarizing doctor
workspace /srv/notarizing-demo/workspace
ok       workspace: the directory mode is 0700
ok       database: the database mode is 0600
ok       store: schema version 12 of 12; WAL journal, foreign keys on, trusted_schema off and query_only readers verified; the owner connection uses synchronous=FULL
ok       owner: notarizing serve (process 16410, version dev) owns the workspace since 2026-10-02T18:30:38Z; browser http://127.0.0.1:7373/; 0 sessions
info     ka2a adapter: the running owner has no ka2a adapter
ok       git: git version 2.43.0
ok       browser assets: the compiled browser assets are embedded
info     semantic provider: disabled (the default); search is lexical and no model is needed
Listing 15 doctor while the owner process runs. The demo owner has no ka2a adapter, so the adapter check is information only.

Evidence

A reported pass is not an established pass.

A receipt says that the workspace accepted exact bytes. The reported claim is what the producer asserts. The evaluated status is what the claim establishes under a named policy. The three stay apart.

StatusWhen
passAn eligible, applicable and complete report supports the requirement under the selected policy, and no report shows a failure.
failAn eligible, applicable report shows a concrete failure, such as a failed test or a reachable counterexample.
staleThe report checked another revision of the specification or the implementation.
unknownThe report is missing, incomplete, from an untrusted source, or from an evaluator that the binding does not approve.
not_applicableThe policy excludes this requirement and check, with a reviewed reason.

Trust is explicit. A local file import stays manual_unverified. Its outcome counts only when the workspace policy trusts its source for that method. The demonstration trusts the local source so that its outcomes can count.

Security

Local by default, strict by design

Specifications, reports and review text can come from anyone. Notarizing treats all of them as data.

Loopback only

The server listens on 127.0.0.1. It checks the exact Host and Origin of each request and sends a strict content security policy. A page request with another Host gets a help page with the exact address and the SSH tunnel command.

One-time secrets

A viewer secret opens one browser session and is then replaced. The session cookie is HttpOnly and SameSite=Strict. Mutations need a CSRF token.

Expiring review grants

Review authority comes from the CLI, for one session, one scope and a limited time. The browser can never change the policy or import evidence.

Nothing runs

Notarizing never executes imported content, fetches URLs from reports or edits a source repository. Patches are proposals that you apply yourself.

Separate authorities

Browser viewer, browser reviewer, CLI control, read-only MCP and ka2a principal each have their own permissions.

Local models only

Semantic search is off by default. When you turn it on, it uses a local Ollama endpoint on a loopback address, and nothing else.

Read one change only

A reader can narrow search and the graph to the selected change. Search leaves hidden objects out before limits, ranking and excerpts, so they change no result. The narrowing never widens a read or a grant.

Mutual TLS for the adapter

The optional ka2a adapter can authenticate to the broker with a client certificate. The key file must be private. notarizing ka2a reload or SIGHUP to serve reads the certificate, key, CA, password and revocation list files again without a restart. A refused reload keeps the previous material.

Denials in the adapter health

A missing broker ACL shows as authorization_failed with its scope in the adapter health. Only the adapter degrades. A producer whose request record is denied gets an error that names the scope.

Built from its own specification

Notarizing reviews the change that builds it

The R2 reimplementation is itself an OpenSpec change with requirement IDs, declared assumptions and an acceptance ledger. The workspace can import that change and review it like any other.

acceptance scenarios with automated tests
99 of 102
scenarios with partial evidence and a named gap
3
requirements with stable IDs
44
Go test functions in the module
920

Real browsers, real server

Playwright drives Chromium against the packaged binary on real workspaces, also offline in a network namespace. The suite checks review conflicts, grant expiry, forged requests, stale responses and hostile text.

Crash-safe receipts

Tests kill a real process before and after the commit of an import. A retry after a lost response returns the same receipt, and a changed report under the same ID is a conflict.

Boundaries as tests

Architecture tests keep the core free of ka2a, Kafka and cgo. A test under a seccomp filter shows that the local demo opens no IP network socket.

A clean, offline environment

The full Go suite, the packaged binary and the three demonstrations ran in a network namespace with loopback only, empty caches and no broker, model or JavaScript runtime. 1984 tests passed, none failed, and 16 skipped with a listed reason. This was one recorded run on one Linux amd64 host.

Model traces replayed on the service

A Quint model of the receipt lifecycle, written from the specification, gave 30 traces (442 steps). They replay against the real service on a real workspace in make verify. This is scoped trace conformance for these traces, not a proof.

Reproducible release builds

make release-artifacts builds binaries for three platforms (linux/amd64, linux/arm64 and darwin/arm64), with SPDX 2.3 SBOMs that include the shipped npm packages, the notices, an evidence summary and SHA256SUMS. In the recorded run, make release-check built everything twice from git archive and found no byte difference. The release build refuses untracked files and ignores the npm configuration of the caller. Nothing is published or signed.

Status

What works today, and what does not

This is a developer preview. The list below describes the reviewed revision in source-state.json.

To set up a workspace, write reports or run the owner process, read the documentation: the user guide, the evidence producer guide, the operator guide, the CLI reference and the release builds. Other guides cover writing specifications, CI recipes, limits and collaboration. The compatibility policy, the privacy and accessibility statements and the threat model state what the project promises.

Works in this revision

  • OpenSpec import at exact commits, with declarations, links and diagnostics.
  • Review, Explore and Compare in the browser, with expiring review grants.
  • Strict check-report/1 import with receipts, bindings and a versioned policy.
  • Exact-ID and full-text search, a reading scope of one change, and optional local embeddings through Ollama.
  • Similar-wording suggestions (may_relate) that count only after a review.
  • Mermaid, Markdown and JSON exports; review and evidence bundles; draft patches.
  • A CLI for every workflow and a read-only MCP server.
  • An optional ka2a adapter with broker mutual TLS, ACL denials in its health and in doctor, a credential reload without a restart (notarizing ka2a reload or SIGHUP to serve), and broker certificate revocation lists through the check of ka2a.
  • Documented JSON examples and the demonstration templates are checked by tests.
  • Evidence corrections and reviewed supersessions that keep both receipts.
  • Reproducible release artifacts for three platforms, with SBOMs and checksums.
  • A known-vulnerability scan (make vulncheck): govulncheck of the Go code and npm audit of the shipped browser packages. The recorded run found nothing.
  • A build without the private ka2a module (build tag noka2a), and make install.
  • Log records of serve on standard error (--log-level, --log-format), repo set-path for a moved checkout, and a help page for a request to a wrong host name.
  • JSON Schemas for bindings, policy, bundle manifests, the adapter configuration and the gotestreport mapping, as editor aids.
  • No telemetry, update check or analytics.
  • A clean-environment run of the full suite, and Quint traces replayed against the service.

Not in this release

  • Hosted or multi-user workspaces.
  • Running checks. Your own harness runs them and writes reports.
  • Edits to a source repository. Patches are proposals only.
  • A terminal user interface. The browser is the review interface.
  • Windows. The workspace store has no owner lock there, so the release has no Windows binary.
  • A reload of the adapter configuration file, the catalog or the signing key. These need a restart of serve.

Not yet verified or blocked

  • A release: no tag, release or upload exists. Publishing and signing are manual decisions of the owner.
  • Public installation: the source repositories are private during the preview.
  • Semantic search and suggestions with a real local model. No Ollama is installed on the test host; the tests used fake providers.
  • Signed producers, and an evaluator behind a real authority boundary.
  • A vulnerability scan of the built binaries. The recorded source scan used an offline database snapshot of 2026-10-01.
  • A screen reader, browsers other than Chromium, and a conformance level for accessibility.
  • Tests on platforms other than linux/amd64. linux/arm64 and darwin/arm64 only cross-compile.
  • OCSP, and an adapter start against a broker whose certificate is already revoked. A reload with such lists was tested against a real broker.
  • Revoked client certificates, broker certificate rotation, and the adapter on a three-broker profile.

Also in this family

ka2a: agent messages that survive the restart

A Go SDK and command-line tool for durable agent-to-agent messages over Apache Kafka. Notarizing can receive evidence through it, and it works without it.