Zero-ETL
Storage-managed table materialization removes a separate connector, not background processing or commit latency.
Zero-ETL means the storage system manages the stream-to-table path instead of requiring a separate connector deployment to export the same events. It does not mean that ingestion writes table snapshots directly or that no transformation takes place.
What replaces the connector?
Ursa appends entries into raw WAL objects. Background compaction and materialization read those entries, decode records through the configured serialization and schema support, and write destination files. The lakehouse group-commit pipeline then publishes those files through table metadata.
This is part of the storage lifecycle described in architecture, not an independent consumer application. It still performs reads, decoding, writes, retries, and commits. The same pass also writes the internal compacted objects that serve stream reads; the destination table is a separate representation.
Three different completion points
| Completion point | What it establishes |
|---|---|
| Append acknowledgment | The log write has completed its durable storage and index path. It does not acknowledge a table commit. |
| Materialization task completion | Output files have been produced and task results persisted for subsequent processing. |
| Table catalog commit | The output becomes part of a table snapshot that query engines can read. |
In particular, LakehouseTableMaterializer.commit() flushes file writers; the downstream commit runner performs the catalog commit. A failed or delayed materialization task does not make an earlier append acknowledgment equivalent to table visibility.
What must be configured?
The current API expresses materialization through TableMaterializationPolicy, with namespace defaults and stream overrides. Resolution requires a catalog reference that resolves to a registered catalog; a stream can explicitly opt out. Merely creating a stream does not guarantee a queryable table.
The remaining decisions include:
- Destination — catalog connection and table identifier. Ursa
1.0.0has no table-ownership mode: the destination table's lifecycle belongs to the destination system, and the table does not provide stream read-back. - Schema and decoding — how payloads become columns, which schema source is used, and what evolution the selected writer supports. Not every payload requires the same external schema registry.
- Write behavior — append or supported upsert behavior, primary keys, partitioning, and sorting where the destination implements them.
- Failure handling — how decoding or schema errors are surfaced, retried, or delivered to a supported dead-letter destination. API options are not proof that every writer supports every behavior.
- Retention and cleanup — source-log trimming, retained compacted data, and destination-table lifecycle are related but distinct policies.
What zero-ETL does not promise
- Zero latency: table freshness includes scheduling, conversion, commit, and failure-recovery time.
- Zero copies: the delivered table is a second set of files. What is shared is the WAL read pass and the storage lifecycle, not the bytes.
- Unified authorization: broker, object-store, and table-catalog permissions still need to be configured consistently. The code does not turn them into one access-control system.
- No data modeling: schema compatibility, destination semantics, and business transformations remain application decisions.
- Identical support across protocols and formats: the corresponding decoders, materializers, and readers must exist and be configured.
See stream–table duality for the distinction between internal compacted objects and delivered tables.