On this page

An order happened on 9 August. The source delivered it on 13 August, three days after the daily revenue table for the 9th had been published.

Two responses are easy to automate. Ignore the order because the partition is “closed,” or rebuild everything from 9 August to today. The first response makes a known fact disappear. The second lets one late record turn into an unbounded amount of work.

I want a third answer: the order still belongs to 9 August, but the way we correct that date depends on how late the record is and what it can change.

Part 1 gave every run an owned interval. Part 2 made writes and side effects converge when that interval is replayed. This article owns the policy between those ideas: when a published interval may be corrected automatically, and what happens after that automatic path closes.

Give each clock one job

A useful late-data record carries at least three times:

  • event_at says when the business event happened;
  • arrived_at says when the platform first observed it durably; and
  • the run timestamp says when a particular attempt processed it.

They answer different questions. For the late order, event_at assigns the fact to the 9 August business interval. arrived_at measures source delay and selects a correction route. The run time is execution metadata.

Putting the order into a 13 August partition because it arrived on the 13th makes operations look tidy, but changes the business answer. Using the run time is worse: the same record can move between dates when replayed. Arrival time should explain lateness, not erase event time.

This assumes the source gives the event a stable identity and a meaningful event timestamp. If the source later corrects that timestamp, the correction needs the same identity plus a source revision or ordering rule. A wall-clock value generated by the pipeline cannot reconstruct that history.

A lookback discovers change; it does not define ownership

Many batch pipelines handle late arrivals by reading the last few days on every run. That can be a perfectly reasonable discovery strategy when the source has no change feed. It is not yet a correction contract.

Suppose today’s job scans four days and finds the missing order. The pipeline still needs to know:

  1. which published interval owns the record;
  2. whether that interval is still eligible for automatic correction;
  3. which downstream outputs include the corrected fact; and
  4. where the record goes if the automatic window has closed.

Without those answers, “look back four days” is just a recurring full scan with a smaller number. It can also hide a gap: a record arriving on day five silently falls outside the query even though the source delivered it successfully.

I separate the policies. Discovery can use CDC, version columns, object manifests or a bounded lookback. Correction starts after a changed record has been mapped back to the interval it owns.

Close automation, not the path to truth

For each source or domain, I use three correction zones:

  1. Open interval. Normal ingestion is still assembling the first published result.
  2. Automatic correction. The interval has been published, but late records may reopen its owned slice and declared dependants without operator approval.
  3. Explicit repair. Automatic mutation is closed. The record is retained with a repair intent, impact summary and reason for review.
Correction horizonStop waiting; keep a repair path
Bounded
Late eventorder-1842Event time09 Aug · 18:42Owned intervalorders · 09 Aug
  1. 01Before first publish
    Open interval

    Normal processing

  2. 02Before policy cutoff
    Correction window

    Automatic repair

  3. 03After policy cutoff
    Repair queue

    Explicit approval

InvariantBusiness ownership stays on 09 AugOnly the correction route changes
03.1 / CORRECTION HORIZON

Arrival time changes how a correction is scheduled, not which business interval owns the event. Closing automatic correction must leave a durable repair path.

The word sealed is useful only if it means “no longer changed automatically.” It must not mean “assumed correct forever.” A source outage, delayed file or upstream correction can still produce older truth. Beyond the horizon, the platform should make that truth visible to an operator instead of silently dropping it.

Batch and streaming systems implement this differently but share the policy. A batch job may use a cutoff and bounded lookback; a stream processor may use a watermark and allowed lateness. Both stop waiting to bound latency and state. Neither boundary proves that no older event exists.

Repair the dependency closure

Correcting the orders / 09 Aug fact slice is only the first step. A daily revenue aggregate for 9 August also changes. A trailing seven-day metric changes for report dates whose windows include the 9th. A trailing 30-day metric has a wider reach.

That does not imply rebuilding every date after 9 August. It means walking the declared dependency closure of the corrected slice.

Correction scopeRepair the dependency closure, not all history
3 outputs
Event
order-1842
Interval
09 Aug
Revision
repair-07
Route
Explicit
Late inputOne orderevent_at · 09 Aug
Owned fact sliceorders / 09 AugRecompute once
Declared dependents
  1. Daily revenue09 Aug
  2. 7-day rollup09–15 Aug
  3. 30-day rollup09 Aug–07 Sep
Outside closureorders · 08 Augtraffic sessionsUnchanged
03.2 / IMPACT CLOSURE

The late order changes one owned fact slice and only the declared aggregates whose windows contain that slice. Unrelated partitions and models stay outside the repair.

The ranges depend on the model. A lifetime cumulative metric may need every later output; an independent daily model needs one partition. Each model should declare which inputs can change one output and which outputs can be changed by one corrected input. The system can then build a bounded correction manifest instead of turning start_date = 09 Aug into “run until now.”

A useful manifest records the event or source revision that triggered repair, the owned input interval, the reviewed model revision and the exact outputs selected for recomputation. Article 4 handles how those results become visible safely. Here the manifest’s job is narrower: keep the blast radius explainable before compute starts.

Choose the horizon from evidence

There is no generally correct three-day or seven-day grace period. The useful number differs by source and by consequence.

I start with four inputs:

Signal What it tells me
Arrival-delay distribution How often records arrive one hour, one day or one week late
Business impact Whether a missed record changes an informal dashboard or a financial result
Correction cost How much state and downstream work the automatic path keeps available
Recovery SLO How quickly an older discrepancy must be reviewed and corrected

A percentile alone is not the policy: a rare late order may also be the most valuable. The horizon should be named, versioned and owned, with its effect on freshness, compute and repair load measured instead of hidden inside one incremental query.

Operate the repair path like a queue

An explicit repair path is not a folder of rejected files that somebody may inspect later. It has state and an owner.

At minimum I want to see:

  • the lateness distribution by source and event type;
  • automatically corrected intervals and their downstream fan-out;
  • the count and oldest age of unresolved repair intents;
  • the reason each intent is waiting, approved or rejected; and
  • reconciliation drift between source totals and published totals.

Alerts should focus on a growing backlog, a shifting delay distribution or material drift. Paging on every late record trains the team to ignore the signal. Quietly counting only accepted records hides the records that matter most.

If almost every repair is approved unchanged, the automatic horizon may be too short. If automatic corrections consume large compute for immaterial changes, it may be too long. The queue turns that decision into evidence instead of anecdote.

What I want to hear in a review

Before approving a late-data path, I want concise answers to these questions:

  1. Which timestamp assigns a record to its business interval?
  2. Where is the first durable arrival timestamp recorded?
  3. How are changed source records discovered?
  4. Which intervals may be corrected automatically, and who owns that policy?
  5. What durable state receives a record outside the horizon?
  6. How is the affected dependency closure calculated?
  7. Which metrics reveal an unhealthy source or repair backlog?

The important boundary is now clear. We stop waiting because compute, state and publication latency must be bounded. We keep accounting because older truth does not stop being true when that boundary passes.

That is the goal: not endless reprocessing, and not artificial finality—a correction policy with a visible end and a repair path beyond it.

Further reading