← All projects

Data platforms · Lakehouse architecture

Mini Lakehouse

One operator can deploy, replay and recover each workload without turning the platform into one tightly coupled release.

Role
Sole designer and engineer
Period
Jul 2026 — Present
Status
Deployed & maintained
Stack
Apache Iceberg · AWS · Airflow · dbt
Hybrid platform mapProvisioned foundation, explicit data layers
Deployed
Terraformprovisions infrastructure
  • OCI runtimeAirflow · PostgreSQL
  • Private edgeCloudflare Tunnel · Tailscale
  • AWS resourcesIAM · Network · S3 · EMR
Airflow orchestrationsubmit · observe · retry
SourcesBatch inputsGitHub · arXiv
ComputeSpark ingestEMR Serverless
01 · S3 / IcebergLandingReplayable windows
02 · S3 / IcebergCuratedCanonical entities
03 · S3 / IcebergAnalyticsdbt-owned marts
Independent servingoutside the data-layer sequence
Document workloadRemote OCR serviceprovider-neutral · resumable
Read-only accessAthena + Inspectoranalytics · document review
On this page

The platform I wanted to operate

I built Mini Lakehouse to answer a practical question: how close could a personal data platform get to production behaviour without inheriting the cost and operational weight of a large organisation?

“Production-like” did not mean adding every service I could find. It meant being able to point to the owner of a dataset, replay a failed day without corrupting current state, replace one runtime without releasing the whole platform, and explain what would happen when a remote job stopped halfway through.

The deployed system ingests GitHub Archive and arXiv metadata, processes documents through remote OCR workers, publishes Iceberg data products and serves analytics through Athena, dbt and small read-only applications. The AWS data plane is separate from a private services host on Oracle Cloud; expensive compute is created only when a workload needs it.

Why the first architecture had to change

The project started as a local stack built around AIStor, Polaris, Trino and Prefect. It was useful for learning Iceberg quickly, but the closer I pushed it towards a real deployment, the more the boundaries blurred. Source code could create or evolve tables, generic ingestion abstractions hid source-specific invariants, and a shared runtime made unrelated changes move together.

The larger problem was that a local stack could not teach the failure modes I cared about. It did not expose the same identity, network and managed-execution constraints as jobs crossing OCI, AWS and a remote GPU provider.

I rebuilt the project around four constraints that genuinely changed the design:

  • The OCI host has no public ingress; operator and browser access must cross private, authenticated paths.
  • AWS workloads use short-lived, workload-specific identity rather than a mounted credential directory.
  • Schedulers and remote providers can retry, so correctness must come from deterministic producers rather than an exactly-once claim.
  • Airflow, Spark, OCR, dbt domains and applications must be releasable and recoverable independently.

Boundaries with operational meaning

The platform is hybrid by intent, not because every component needs a different cloud. OCI hosts the long-running control services I already operate. AWS owns the data plane and ephemeral Spark execution. OCR work can move between remote GPU or CPU providers without importing their SDKs into the Airflow runtime.

Within the data plane, versioned contracts describe the Iceberg structure and validation rules. Spark owns landing and curated publication; OCR workers own document artifacts and their run records; each dbt domain owns its analytics models, tests, runtime identity and release. Airflow submits work, observes it and emits a logical asset only after the producer succeeds. It never becomes the place where business state is stored.

Terraform owns the infrastructure boundary, not the table lifecycle. It provisions the OCI host, Cloudflare Tunnel and Tailscale access, plus AWS networking, buckets, KMS, IAM, ECR, EMR and the containers for runtime secrets. Versioned YAML contracts and PyIceberg own Glue databases and tables, so an infrastructure apply cannot silently change a dataset schema, partition spec or identifier field.

Access is split by workload rather than inherited from the host. The OCI machine has no public ingress: browser traffic crosses Cloudflare Access and operator traffic stays on Tailscale. Airflow may submit and cancel EMR work but cannot read or write the lake on a job’s behalf; the EMR runtime owns source and data permissions. Each dbt domain receives a separate identity scoped to its curated inputs, analytics output and query-result prefix, while CI exchanges short-lived credentials instead of storing AWS keys.

This creates more deployment surfaces than a single-machine stack. I accept that cost because the boundaries line up with the things I need to debug and roll back. If a small deployment did not need independent recovery, I would collapse some of them.

Following one GitHub day through the system

The GitHub Archive path is the clearest example of how those boundaries work. At 07:30, Airflow resolves the previous local day and submits a pinned EMR entrypoint with its locked environment and contract bundle. The Spark driver fetches all 24 hourly archives, validates gzip content and records SHA-256 metadata before publication.

Landing is treated as an authoritative window: replaying the day replaces exactly those hourly partitions. Curated events merge by event_id; actor and repository state use (last_observed_at, last_event_id) so an older observation cannot overwrite a newer one. Only after the remote job succeeds does Airflow publish the curated asset that triggers the Engineering dbt domain.

One bounded data runA GitHub day, replayed safely
Idempotent path
  1. 01
    AirflowResolve window

    Previous local day

  2. 02
    SparkFetch 24 hours

    Checksum every object

  3. 03
    LandingReplace partition

    Authoritative replay

  4. 04
    CuratedMerge state

    Deterministic keys

  5. 05
    AirflowEmit asset

    Success only

  6. 06
    dbtPublish marts

    Domain-owned release

Run invariants
  • No duplicate event_id
  • Older observations cannot win
  • Retry converges
01 / RUN TRACE

Airflow coordinates the run; deterministic producers own replay, merge and publication boundaries.

Failure is part of the data contract

I designed replay before trying to make the happy path look exactly-once. A broken or empty archive never reaches S3 as valid input. A retry can reuse an object only after its checksum is verified. If Spark is interrupted, Airflow cancels the remote job and retries the same bounded window; partition replacement and deterministic merges make that rerun converge.

The OCR path follows a similar pattern with a different recovery boundary. Its request ID is derived from the document revision and configuration. Provider tokens are checkpointed, artifacts and manifests are validated, and the Iceberg run record is marked imported only after the outputs are safe to reuse.

Failure and recoveryRetries return to a known boundary
Explicit recovery
  1. 01Resolve window
  2. 02Produce data
  3. 03Validate output
  4. 04Commit boundary
Failure 01Broken source object

Reject before publish

gzip + SHA-256
Failure 02Interrupted Spark run

Replay the same window

overwrite + MERGE
Failure 03Partial EMR upload

Keep previous pointer

complete checksum set
02 / RECOVERY MAP

Each producer rejects incomplete work and retries from a bounded boundary; Airflow stores no business state.

This design does not pretend that every failure is solved. The development environment favours rebuild and restore over high availability. The OCI host is not HA, backup and restore drills are not yet automated, and S3 data plus catalog metadata must be recovered as one Iceberg system. Those are current operational gaps, not hidden fallbacks.

Delivery without a platform-wide release

CI starts by proving that a revision is internally consistent. Depending on the changed paths, a pull request enters the general make check quality gate, the focused Airflow-bundle gate or both. Together they check frozen lockfiles, data contracts, dbt projects, formatting, linting, types, unit tests, Terraform, Compose definitions and Dockerfile build rules. The split keeps DAG and runtime validation inside their own release boundary without making the general workflow repeat it.

A merge to main does not deploy the monorepo as one unit. Path-filtered workflows select the component that changed: Airflow, either dbt domain, OCR, the Inspector or EMR jobs. Image releases validate the component boundary, assume a narrowly scoped AWS role through GitHub OIDC, build both ARM64 and AMD64 images with provenance and an SBOM, then deploy the exact ECR digest. The runtime never follows a mutable tag.

EMR follows the same rule without a container. CI packages entrypoints, locked dependencies and the contract bundle under the Git commit SHA. The revision becomes visible in SSM only after every artifact and its checksum manifest exist, so a partial upload cannot become the active release.

Deployment and rollback therefore use the same reviewed evidence. A component can move independently, while a rollback must name an existing digest and the matching revision from main; it does not rebuild an old commit and hope for the same bytes. Release jobs use short-lived credentials rather than stored AWS keys, and container deployments reach the private services host through Tailscale instead of SSH keys or a public ingress path.

Where the project stands

The complete topology is deployed and maintained. It currently supports two batch sources, curated Iceberg products, provider-neutral document OCR, two analytics domains and read-only inspection workflows. Its active runtime and data products now have independent validation, release and rollback boundaries.

The next useful work is operational evidence rather than another engine: boundary data-quality rules, actionable freshness alerts, end-to-end lineage and telemetry, cost/SLA dashboards, and automated restore drills with explicit RPO/RTO. After that, I can test permissions and idempotency against disposable infrastructure before deciding whether incremental dbt or deeper governance is justified.

I will add cost and workload numbers only after they come from a repeatable measurement process. The criterion for future scope stays the same: every capability needs a use case, an owner, a source of truth and a recovery story I can test.