On this page

The late order from part 3 has reached an approved repair. The job rebuilds the 9 August slice, runs schema, key, reconciliation and scope checks, then reports a clean green result.

Before the repair publishes, another writer advances the table from snapshot N to N + 1. The repair still tries to commit. Its request reaches the catalog, but the response never reaches the worker.

The validation task passed. The workflow still cannot answer three basic questions:

  1. Which exact candidate did those checks certify?
  2. Did readers move to that candidate?
  3. Is it safe to retry or delete the candidate files?

This article makes that boundary operational: one candidate identity, evidence bound to it, and one conditional change to the state readers resolve.

“Validated” needs a subject

A validation report without a subject is little more than a timestamp. “The staging table passed at 14:20” is not enough if another process can overwrite that table at 14:21. Neither is “the latest output passed” when latest can name something different by the time publication starts.

For repair-07, I want the report to bind at least:

  • the base table revision used to plan the repair;
  • an immutable candidate or content digest;
  • the owned input and output scope;
  • the code, schema and contract revision; and
  • the individual checks and their evidence.

The candidate can be a staged snapshot, audit branch head or versioned manifest. What matters is that one identity means the same content throughout validation and publication.

If a process rewrites the candidate in place, it creates a new subject. If rebasing the repair on N + 1 changes its files or assumptions, it also creates a new subject. Evidence from repair-07 cannot be copied onto repair-08 just because both came from the same workflow.

Candidate protocolValidation certifies one candidate, not “the table”
Identity bound
Immutable candidaterepair-07
Base
snapshot N
Scope
orders · 09 Aug
Bound evidencereport-07
  • Schema
  • Keys
  • Reconcile
  • Scope
Subject · repair-07
Publish gateCompare + swap
Expect
snapshot N
Set
main → repair-07
Exception routes
  1. Evidence failsCandidate is not certified
    QuarantineKeep candidate + report
  2. Base movedCommit assumption changed
    Rebuild + validateCreate a new candidate ID
04.1 / CANDIDATE STATE

Evidence remains valid only while its candidate and commit assumptions remain unchanged. Failed checks quarantine the candidate; a changed base creates a new candidate and new evidence.

Checks certify repair-07, not “orders” in general. Failed evidence is quarantined; passed evidence reaches the publish gate with its expected base. After a conflict, refresh the table and retry only if the operation’s assumptions still hold. Any changed subject or assumption needs new evidence.

Files existing is not publication

Candidate files normally exist before a metadata commit and may be reusable across safe retries. Their presence is therefore a poor publication signal.

Readers need one authoritative reference. In a snapshot table, that is usually a catalog or table metadata reference that resolves to one current snapshot. In a manifest-based dataset, it may be a versioned manifest pointer. The protocol works only when readers agree to follow that reference instead of listing a directory and guessing which files belong together.

Publication is the conditional update of that reference:

compare current reference with expected base N
if equal: set reference to repair-07
otherwise: reject and refresh

The compare and update must be one atomic catalog operation. Checking N, waiting, and then writing a pointer in a second operation leaves a race between the two steps.

Before the swap readers resolve the old snapshot; afterward they resolve the new one. Producing the candidate may take minutes without exposing a half-published state.

A conflict changes the assumptions

Suppose another writer commits N + 1 after repair-07 is validated. The repair should refresh the table and ask whether its intended change is still valid.

For a pure append, it may be safe to attach the same new files to the newer snapshot. For a repair that replaces 9 August, the protocol must confirm that no concurrent change touched the replaced data or violated another declared assumption. If it did, the repair needs a new candidate based on the new visible state.

Evidence follows the same logic. Candidate-only checks may survive an unchanged candidate. Checks against the base table, replacement scope or visible totals do not survive a changed base automatically.

I prefer recording those dependencies in the report instead of treating validation as one boolean:

Evidence Bound to Reuse after base changes?
Schema compatibility Candidate + table schema Only if both identities are unchanged
Key uniqueness Candidate content Yes, if the candidate content is unchanged
Reconciliation total Candidate + source revision Only if both inputs are unchanged
Replacement scope Candidate + base snapshot Re-evaluate against the new base

A conservative starting rule is to rerun all checks after a rebase. Selective reuse requires more identity, not less.

A timeout is not a failed commit

Now return to the lost response. There are at least two possible worlds:

  • the atomic swap failed and repair-07 never became visible; or
  • the swap succeeded, readers can see repair-07, and only the acknowledgement was lost.

Blindly retrying assumes the first world. Deleting candidate files also assumes the first world. Either action can be wrong.

Commit knowledgeReconcile the pointer, not the timeout
One visible state
  1. Reader visibility
    Certified candidaterepair-07Built from snapshot N
    Atomic catalog commitCompare + swapExpect N · set main
    Reader referencemain → repair-07One visible snapshot
  2. Worker knowledge
    Commit responseLostOutcome is unknown
    Status checkReconcileCurrent + recent history
    Candidate foundMark successDo not commit again
04.2 / VISIBLE REFERENCE

A timeout changes what the worker knows, not necessarily what readers can see. Resolve the candidate against the current reference and its history before retrying or deleting files.

The visual shows the second world: readers already resolve repair-07, while the worker still knows only that the response was lost. Its uncertainty is about the response path, not reader atomicity.

The next action is reconciliation, not another commit. Use the stable publish or candidate identity to check the current reference and, where the table format supports it, recent reference history. If repair-07 is present, treat the publication as successful. If the system can prove it never committed, the workflow may retry under the normal concurrency rules. If status remains unknown, retain the artifacts and escalate; uncertainty is not cleanup authority.

This distinction is easy to lose when a scheduler has only success and failure task states. The data protocol needs a third durable state even if the orchestration UI represents it as a failed task: commit outcome unresolved.

Name the atomic unit

An atomic commit is always atomic over some boundary. A single Iceberg table snapshot does not make three other table snapshots visible in the same transaction. A database transaction does not include an object-store manifest merely because the same function writes both.

For a multi-table data product, there are three honest options:

  1. publish tables independently and design readers to tolerate compatible versions;
  2. publish a higher-level versioned manifest that names the exact table revisions readers should use together; or
  3. use a catalog or storage system that genuinely provides the required multi-table transaction.

A higher-level manifest works only when every reader resolves it; direct table reads bypass the protocol. Catalog transactions should be verified from their actual guarantees, not inferred from the word transactional.

Cleanup comes after truth

Cleanup follows truth resolution. Failed validation can enter quarantine; a known failed commit can become eligible for orphan cleanup after in-flight work and references are checked. Unknown outcomes remain protected. Retention shorter than a legitimate write or reconciliation window can delete active files or erase the evidence needed to resolve a commit.

What I want to hear in a review

Before approving a publish path, I want concise answers to these questions:

  1. What immutable identity names the candidate?
  2. Which base revision and source revision produced it?
  3. Does every validation report name its exact subject and assumptions?
  4. What single reference do readers resolve?
  5. Which atomic operation changes that reference?
  6. Which conflicts are safe to retry, and which require a new candidate?
  7. How is a lost commit response reconciled by identity?
  8. Who may quarantine or delete files, and after which outcome is known?
  9. Is the claimed atomic unit one table, one manifest or a real multi-table transaction?

Validation is essential, but it is evidence—not publication. The protocol becomes safe when the evidence, candidate, base assumptions and reader-visible reference all describe one revision, and when uncertainty is resolved before the workflow acts again.

That is a stronger promise than “checks passed.” It tells us exactly what passed, exactly what readers can see, and exactly what to do when the answer from the commit never arrives.

Further reading