Rebuilding from the spec: a scaffold

Table of Contents

1. Status of this page

Scaffold. The method below is in use on one project. Two of eight planned rebuilds exist. This page records what has been measured, marks what has been relayed from elsewhere, and names the gaps as gaps. It is not a synthesis, and no cross-project agreement is claimed.

2. The method

One hub. The hub holds the specification, the oracle, and the acceptance gates. Around it sit N bare spokes. Each spoke rebuilds the same system independently, in its own language, carrying no governance of its own. One shared oracle grades every spoke against the same properties.

Roles are coordinator, worker, and gate. The coordinator schedules and reconciles. Workers build. Gates decide whether a build counts. An append-only review ledger records every case the text got wrong and the build found.

Where two spokes disagree, the disagreement is evidence about the specification. The implementations are the instrument. A clause that two competent readers resolve two ways is an ambiguous clause, whatever its author intended.

The spec governs its own derivatives, which is what makes the disagreement readable:

This document governs. README.md, CLAUDE.md, AGENTS.md, berths.tsv, docker-compose.yml, and toxiproxy.json are derived from it. When a derived file disagrees with this spec, the derived file is the bug.

2.1. A rule without a runnable oracle is not a rule

The method fits in one line of the governing spec, at the head of its rules section:

Each rule names its oracle. A rule without a runnable oracle is not a rule.

That is a rule of authorship. Prose that cannot be checked is demoted to commentary at the moment it is written, before any build reads it. The browser-harness work in this repository reached the same place from the other side, as a failure observed in practice: a gate that cannot execute proves nothing. Authorship discipline and gate discipline are the same constraint seen from two ends.

The consequence in the source project is that each invariant carries its check. R1 (one service, one directory, one process, one port) is checked by berths.tsv having exactly one row per service directory with each port appearing once. R2 (services share no code) is checked by a grep for cross-directory imports returning nothing. R6 (observers read and never call) is checked by the observer source containing no POST, PUT or DELETE. R8 (sensitive input is not logged) is checked by a request carrying a PAN producing no log line that contains it.

Each check is mechanical, cheap, and falsifiable on a working tree. None of them requires reading the prose to evaluate.

2.2. The harness is a contract, not a model

The spec states its own scope in a way that keeps the oracle central:

It is a test harness, not a discrete-event simulation: timing is wall-clock, transport is real HTTP and FIFOs, and the value is in the contracts and the oracle, not in the model fidelity.

Fidelity is explicitly out of scope, which means a rebuild is graded on contract conformance alone. That is what lets eight implementations in eight languages be compared without arguing about behaviour the spec never claimed.

3. Why rebuild more than once

One implementation cannot falsify its own specification. The author reads the text, writes the code, and the gaps close silently in the author's head. The result is a system that works and a text that only appears to say why.

The second implementation is the measurement. It turns "the spec is clear" from an opinion into an observation with a number attached: the count of clauses that two builds read differently. N rebuilds give N choose 2 pairwise comparisons over the same oracle.

TODO: state the pairwise count for the live instance once a second spoke has passed the full gate set. Two spokes exist; the comparison has not been run.

4. The oracle and the gates

The oracle is a single program. It reads every event stream and exits nonzero on the first violated property. The properties are ordering and accounting claims over the event log: stage transitions per order arrive in order with non-decreasing timestamps; a release is followed by a packed stage within a bound derived from the declared delay ranges; allocated never exceeds on-hand per SKU; each sourced order names exactly one warehouse; each recorded transition exists in the machine's published table; sink counts equal what a scripted generator sent.

One detail in the spec's oracle section matters more than the properties themselves. Each property is negative-tested with a deliberately broken stream. The oracle has to be shown failing before its passes mean anything. That is the authorship-side answer to the first gating finding below.

Gates come in two kinds. The documents gate checks that org exports, that every link resolves, that the derived berth table matches the spec table, and that derived files carry their derived-file header. The code gate runs the harness up, healthy, and down on two operating systems, runs each rule oracle that was itself negative-tested, runs the property oracle against both a recorded fixture and a live run, and records the result in the commit's git note along with the cell it ran on.

5. The ledger

The review ledger is append-only. One entry per case the text got wrong and the build found. Entries are never edited, so the record of what the specification used to say survives its own repair.

R7 is where the ledger meets the rules. Nothing is merged for convenience; a new capability that could be a separate process is a separate process. Its oracle is a review-ledger entry required for any exception. An exception to a rule is therefore never an edit to the rule. The ledger is how the spec records being wrong without quietly becoming right.

Per the source project, such ledgers mostly contain four shapes:

  1. A case no clause named. The text enumerates the situations it expects; the build reaches one that is absent from the enumeration.
  2. A field required on one side of an edge and absent on the other. Producer and consumer were specified in different sections, and the sections do not agree.
  3. A capability added in one section while the section enumerating absences still denies it. The specification contradicts itself across an edit.
  4. A check that reports success because its own mechanism is broken. See the gating findings below.

The copy of the governing spec read for this note is v0.1.0, dated 2026-09-12, and its ledger reads "Entries so far: none", with a reconcile issue open against v0.2. The four shapes above are therefore relayed from the source project rather than counted from that document.

TODO: publish entry counts by shape once the ledger has entries that can be quoted in public.

6. Gating findings measured here

A gate is the thing that says a build counts. Four findings about gates were measured in this repository during browser-harness work. They are generic, they are already folded into the shared package, and they are stated here because they are the part of this work with local evidence behind it.

6.1. A gate that cannot execute is indistinguishable from a gate that ran and failed

A spec harness here shipped with two fatal errors. All fifteen targets printed "failed" and exited nonzero. That output is identical to fifteen real findings. The fix is to assert a positive count of units examined. An exit code alone cannot separate "found a bug" from "never started".

6.2. A standing known-bad fixture must assert its own badness

A canary page documented as gated off had been un-gated. Calibrating the harness against it did not fail. It went vacuous: the check passed because there was nothing left to catch. A known-bad fixture has to verify that it is still bad before anything calibrates against it.

6.3. Append-only measurement artifacts produce silent wrong deltas

Trace files that accumulate across runs mixed old and new states inside a before/after comparison. The delta was computed over the union of two runs and reported without error. Measurement artifacts need a run identity, or a clean slate per run.

6.4. Vacuous guards inflate green

A property guarded on a condition that never admits passes forever. Its output is identical to that of a property that genuinely holds. Every guarded property needs a second assertion: that the guard admitted at least once.

7. Eventing: one contract, two transports

The eventing work applies the same idea one level down. One contract is stated once. Two transports are driven from it. The dashboard is then a differential test of the contract, and any divergence is a property of the text rather than of either broker.

Greyscale screenshot of a self-contained eventing dashboard. Two panels, MQTT on port 11883 and NATS, both marked running, each reporting 1010 messages received under a grind x5 load, with sample message bodies showing external_ref, template, order_id and vars.total_cents. A contract panel below lists the MQTT topic pattern wharfinger slash mail slash request slash external_ref at qos 0 retained false, the NATS subject mail.request.external_ref, the shared payload schema, and a source line citing the spec section it came from.

Figure 1: Local eventing dashboard. MQTT on port 11883 and NATS side by side, both running, each having received 1010 messages under a grind x5 load. The contract panel below states one topic pattern per transport, one payload schema, and the spec section the contract is drawn from verbatim.

The panels: MQTT on port 11883, NATS beside it, both marked running, 1010 messages each under a grind x5 load. Message bodies carry external_ref, template (payment_captured or order_confirmed), order_id and vars.total_cents.

The contract panel states the topic patterns: wharfinger/mail/request/<external_ref> for MQTT at qos 0, retained false; mail.request.<external_ref> for NATS. One payload schema serves both. A source line cites the spec section the contract is drawn from verbatim, so the dashboard and the text can be diffed by eye.

TODO: reconcile that source line. The copy of the governing spec read for this note is v0.1.0 and contains no mail contract: its contracts cover the seven services in the tarball (pickstage, wmsim, omsim, solidus-sm, taxsim, risksim, otelsink) plus tally, which is unbuilt. Either the dashboard cites a later spec version, or the contract is ahead of the document. The method says the derived file is the bug; establishing which file is derived here is the open item.

TODO: record what the differential has actually caught. The harness runs; no divergence has been written up.

8. Status and what is missing

This is the section that keeps everything above it defensible.

The live instance is a private repo simulating a commerce and fulfilment estate: seven services in the current tarball, up to eight language rebuilds planned, a live reproduction, MQTT and NATS eventing experiments, and a contract-evolution demonstration in which an enum rename propagates across implementations.

Its own coordinator reports:

area state
system structure complete
one full order flow complete
orders plus email incomplete; no order has produced its mail through a real driver
NATS integration not started
language rebuilds 2 of 8 real

Read the second row against the third. One order flow completes. The flow that ends in a delivered message has never completed through a real driver, which means the eventing contract above is exercised by the harness and has not yet been exercised by the system it specifies.

The spec's own build steps put the first real-system substitution last: replace the Solidus state-machine simulator with a real Solidus instance behind the same berth and diff its state-change table against the simulator's records for one scripted order. That step has not been reached.

TODO: what counts as "a real driver" needs a written definition before the orders-plus-email row can be closed.

9. Evidence classes

This is the blocker on publishing anything stronger than a scaffold. Two kinds of claim are in play and they carry different weight.

Measured here. The four gating findings above were observed in this repository, on this machine, at a known version. The harness that printed fifteen false findings was run and read. The canary page was checked and found un-gated. The quotations from the governing spec were read out of the document itself.

Relayed, unverified. Corroborating findings from two other projects arrived as self-reported messages from other sessions. They were never independently reproduced here. The coordinator's status table above is also a report rather than an observation. These are plausible and they agree with the measured half, which is exactly why they are dangerous: agreement between a measurement and a report is one measurement.

The distinction has to survive into any eventual synthesis, per claim, in the text. A page that blurs the two would be weaker evidence than either half standing alone. The site's verdict vocabulary already carries the right terms: reproduced for the first class, attributed for the second.

TODO: label every claim on this page with its class in a property drawer once the claims stop moving.

10. Open questions

  • How many spokes before the marginal spoke stops finding new ambiguity. Two is enough to measure disagreement. The shape of the curve after that is unknown.
  • Whether language diversity across spokes matters, or whether two builds in the same language by different workers find the same clauses.
  • Whether the oracle can be specified in the same document it grades. An oracle written from the same text may inherit the same blind spots.
  • Whether ledger entries should be graded by severity. Today every entry is one entry.
  • Whether the contract-evolution demonstration (an enum rename propagated across implementations) generalises to changes that are not renames.
  • Whether a rule's oracle can itself be ambiguous, and what checks the checker. Negative-testing each property is the current answer; it bottoms out somewhere.
  • When the shared package is stable enough to lock. It is still growing; publishing now would fix a snapshot taken mid-convergence.

11. Refutation conditions

The method described here is wrong, or at least oversold, if any of these is observed:

  1. Independent spokes stop disagreeing while the specification still contains clauses that a fresh reader resolves differently. The rebuilds would then be measuring shared context rather than the text.
  2. Ledger entries turn out to be dominated by implementation defects rather than specification defects, under a classification written before the entries were read.
  3. A second spoke costs more than the ambiguities it surfaces are worth, on a cost measure agreed in advance.
  4. The shared oracle is found to encode one reading of the specification, so that spokes converge on the oracle's interpretation instead of on the text.
  5. The four gating findings fail to reproduce outside this repository, which would make them local facts rather than generic ones.
  6. The eventing differential across two transports never diverges on any contract change, which would make the second transport decoration.
  7. A rule that names a runnable oracle proves no less ambiguous in practice than one that does not. The rules section's central claim would then be false.

12. TODO before this stops being a scaffold

  • [ ] Orders plus email completes through a real driver; define "real driver".
  • [ ] NATS integration started, and the differential run against a contract change.
  • [ ] A third spoke, so pairwise disagreement has more than one pair.
  • [ ] Reconcile the dashboard's source line with the spec version it cites.
  • [ ] Ledger entries by shape, with counts, quotable in public.
  • [ ] Per-claim evidence-class drawers (reproduced or attributed).
  • [ ] Independent reproduction of the relayed findings, or explicit demotion.
  • [ ] A decision on whether the shared package is stable enough to cite.