On this page

The orders · 09 Aug repair now has an owned interval, stable write and effect identities, a bounded late-data path and one conditional publish for candidate repair-07.

Now someone asks a fair question in review: how do we know any of that survives a real retry?

“The tasks are green” does not answer it. Neither does “I retried the DAG in staging and it worked.” Those are useful observations, but replayability is a claim about state transitions under failure. It needs a known start, a controlled disturbance, independent observations and named invariants.

The contracts established in the previous four parts become assertions here. The goal is to make them fail when the implementation stops honoring them.

A green task is not the system state

A scheduler knows about an attempt. The state we care about is spread across other boundaries:

  • the table or manifest reference readers resolve;
  • the immutable candidate and owned output digest;
  • the receipt held by an external side-effect consumer;
  • the partitions and models outside the declared repair closure; and
  • the durable recovery state that permits retry, quarantine or cleanup.

A task can be red after a successful commit response is lost, or green while an effect is applied twice. Task state is evidence about execution, not an oracle for the protocol.

For each test, I want four explicit parts:

  1. Setup — the run contract, code and schema revisions, starting references and existing effects.
  2. Stimulus — the exact action sequence and the named point where failure is introduced.
  3. Observation — fresh reads from the catalog, output store, effect receiver and repair boundary.
  4. Verdict — the invariant each observation supports or contradicts.

Missing any one part makes the result hard to reproduce, weakly scoped or dependent on the code under test reporting its own success.

Write the invariant before the test

“Replay the pipeline and compare row counts” sounds practical, but it leaves too much unnamed. Which slice should converge? Can another writer’s data change? Is one extra effect acceptable? May an unreferenced candidate remain after a known failure?

For this running example, I would start with six reviewable invariants:

  1. One visible result. Readers resolve either the previous revision or the complete certified candidate, never half of each.
  2. Replay convergence. Submitting the same run contract again creates neither a second logical result nor a duplicate business effect.
  3. Conflict preservation. A concurrent writer remains visible; the stale repair cannot silently overwrite it.
  4. Bounded repair. The declared dependency closure may change, while unrelated outputs retain their identities or digests.
  5. Uncertain means retained. An unknown commit result cannot authorize another blind commit or deletion.
  6. Traceable recovery. Retry, rebase, quarantine and reconciliation decisions link back to the same run, candidate and fault trial.

The subject matters. “No duplicates” is vague; “one receipt for this effect ID after two deliveries” is testable. “Unrelated data is safe” becomes an assertion only when the unchanged outputs are named.

Inject failure where ownership changes

I do not start by turning on random chaos across the whole environment. That usually produces an interesting incident and a weak explanation. I start with deterministic failpoints at seams where one component may have acted but the next component does not yet know it:

  • after candidate files are durable, before the visible commit;
  • after the catalog commit, before its acknowledgement reaches the worker;
  • after an external effect is applied, before its receipt is recorded by the sender; and
  • after an outcome becomes eligible for cleanup, before deletion begins.

Use a small test seam or controllable network proxy, but name the semantic boundary. AFTER_COMMIT_BEFORE_ACK is more useful than “random timeout after 18 seconds.” The matrix starts from one baseline publication and challenges it with five deterministic trials.

Replay proof matrixInject failure where ownership changes
5 controlled trials
TrialVisible stateLogical effectRecoveryBlast radius
01Exact replay
Same resultOne reference
Applied onceStable effect ID
ConvergedNo new intent
No changeOutside contract
02Pre-commit stop
Previous resultPointer unchanged
Not visibleCandidate only
Safe retryKnown failure
Owned tempRetained by policy
03Lost response
repair-07Commit landed
Applied onceNo blind repeat
ReconciledIdentity lookup
No cleanupUntil resolved
04Concurrent writer
N + 1 keptNo overwrite
Not repeatedStale intent held
RevalidateAssumptions first
Conflict scopedTouched data only
05Late repair
Repaired resultCertified publish
One correctionStable repair ID
ClosedManifest recorded
Declared closureOthers unchanged
05.1 / FAILURE COVERAGE

A scenario counts only when it names both the injected boundary and the state that must remain true. The baseline establishes the reference; these trials challenge it.

In the highlighted trial, readers already see repair-07 while the worker is uncertain. Evidence must show one candidate occurrence, one external receipt, a reconciled state and no premature cleanup—not merely another green attempt.

Allowed intermediate states belong in the contract

Not every leftover file is a leak. After a pre-commit stop, an unreferenced candidate may remain until retention permits cleanup while the visible reference stays unchanged. An unknown outcome protects the same files until reconciliation. Tests should assert the allowed state for each outcome, not demand immediate emptiness.

Observe through the boundary, not through the task

The useful test reads state through interfaces that a real consumer or operator would trust. For the running example, observe() resolves the catalog reference and recent history, reads the effect receipt from its destination, computes digests for the owned and unrelated outputs, and checks the durable recovery record.

One compact trial is enough to show the shape:

before = observe(contract)

run(contract, fail=AFTER_COMMIT_BEFORE_ACK)
retry(contract)

after = observe(contract)
assert after.visible_candidate == "repair-07"
assert after.commit_count("repair-07") == 1
assert after.effect_count(contract.effect_key) == 1
assert after.unrelated_digests == before.unrelated_digests

observe() must not return the worker’s cached outcome. Table history, the effect receiver and unrelated-output digests are separate sources. A total row count is a weak proxy for identity; a count scoped to one effect key or candidate directly tests the contract.

Generate sequences after the model is clear

Stabilize the deterministic trials before adding stateful tests. Once the model is explicit, a state machine can combine produce, validate, publish, lose acknowledgement, reconcile, retry, conflict and repair, checking invariants after each action. Its value is exploring missed orderings and shrinking a failure into a small permanent regression case—not randomness itself.

Prove that the proof can fail

A test suite should prove that its oracle is connected. In an isolated fixture, disable one guard:

  • replace the stable effect key with an attempt-local ID;
  • turn an unknown commit outcome into an immediate retry; or
  • let a repair rewrite an output outside its dependency closure.

The matching invariant should expose two receipts, two candidate occurrences or a changed unrelated digest. Intercept the boundary in the harness; do not add hidden test switches to production code.

Keep a proof bundle, not just a log archive

Keep a compact proof bundle containing:

  • the run contract, code and schema revision;
  • initial references and relevant input identities;
  • the trial name, failpoint and generated seed when one exists;
  • the action history and independently observed states;
  • the verdict for each applicable invariant; and
  • retained candidates or receipts needed to investigate the failure.
Replay evidenceA green task is one observation, not the verdict
Reproducible
Supported claimorders · 09 Aug is replayable within the tested boundaryEvidence complete
  1. 01 · Setup
    Contract
    orders · 09 Aug
    Base
    snapshot N
    Code
    immutable SHA
    Schema
    revision 12
  2. 02 · Stimulus
    Trial
    lost response
    Failpoint
    commit → ack
    Action
    retry contract
    Mode
    deterministic
  3. 03 · Observations
    Reference
    repair-07
    Commit
    1 occurrence
    Effect
    1 receipt
    Outside
    unchanged
  4. 04 · Verdict
    Visibility
    holds
    Replay
    converged
    Unknown
    reconciled
    Cleanup
    guarded
Not covered by this bundleMulti-region catalog outageSeparate integration trial required
05.2 / PROOF BUNDLE

Evidence is reviewable when another engineer can reconstruct the starting state, repeat the fault and obtain the same verdict. The claim remains limited to the boundaries the bundle records.

The “not covered” line is part of the proof. A lost-response trial says little about a multi-region catalog outage; naming the gap prevents a local result from becoming a universal claim.

Put each proof at the right cadence

Not every scenario belongs in every pull request.

  • CI: deterministic contract trials using stable identities, controlled failpoints and a small isolated state model. These should be fast enough to block a change.
  • Integration: real object storage, catalog and receiver boundaries, with connection faults and concurrent writers in an isolated environment.
  • Scheduled drills: longer generated sequences, retention behavior and recovery from failures that are too expensive for each commit.
  • Production exercises: only the risks that cannot be represented safely elsewhere, with an explicit blast radius and stop condition.

At every cadence, record the revision, supported invariant and age of the evidence.

What I want to hear in a review

Before accepting “this pipeline is replayable,” I want concise answers to these questions:

  1. Which run and output boundary does the claim cover?
  2. What independent oracle observes each external boundary?
  3. Which failures are injected, and what intermediate states are allowed?
  4. Does replay converge in data and side effects without overwriting a concurrent writer?
  5. Do unknown outcomes block retry and cleanup until reconciliation?
  6. Does repair leave outputs outside its declared closure unchanged?
  7. Can each oracle catch its broken guard, and which relevant risk remains untested?

That closes the series: owned runs, convergent effects, bounded correction and conditional publication must survive controlled failure and leave reproducible evidence.

A replayable pipeline is not one that never gets stuck. It is one whose state remains explainable when work repeats, whose recovery actions are constrained by what is actually known, and whose guarantees can be challenged without relying on luck.

Further reading